# Google DeepMind 推出 Gemini 3.6 Flash 與 Gemini 3.5 Flash-Lite 兩款新模型，大幅縮減執行時間並提升 token 效率

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：Artificial Analysis (@ArtificialAnlys) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥🔥 · 日期：2026-07-22

> 原始來源：https://x.com/ArtificialAnlys/status/2079596244339707956

## 證據與延伸閱讀

- [定價與快取折扣](https://artificialanalysis.ai/)

## 中文摘要

Google DeepMind 推出 Gemini 3.6 Flash 與 Gemini 3.5 Flash-Lite 兩款新模型，大幅縮減執行時間並提升 token 效率。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/9e10b95dc1ea44c3.jpg)
> Google 發布 Gemini 3.6 Flash 與 Gemini 3.5 Flash-Lite，兩模型每任務解碼時間皆較前代減半，其中 Gemini 3.5 Flash-Lite 智力指數大幅提升 11 分至 36 分，而 Gemini 3.6 Flash 智力指數維持 50 分與前代持平。

**Gemini 3.6 Flash 效能表現**
Artificial Analysis 在模型發布前進行了基準測試，整理出 Gemini 3.6 Flash 的核心數據與效能變化：
- 維持與前代相同的智慧水準：在 Artificial Analysis 智慧指數（Intelligence Index）拿下 50 分，與 Gemini 3.5 Flash 相當，略低於 Muse Spark 1.1 的 51 分與 GPT-5.6 Luna 的 51 分。
- 專案表現差異：在 `GDPval-AA v2` 評測中進步至 1421 分（提升 72 分），但在 `HLE` 評測中則微幅下滑 3 個百分點至 38%。
- 執行時間減半：平均每個任務的執行時間為 1.3 分鐘，相較於 Gemini 3.5 Flash 的 2.7 分鐘縮短超過 50%。
- 速度與成本最佳化：在上市前測試中，輸出速度達到每秒 304 個 token。每個任務的成本降低約 18%，從 0.59 美元降至 0.50 美元。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/7f5d0ea59b33513a.jpg)
> Gemini 3.6 Flash 在 GDPval-AA v2 基準測試取得 1421 分（相較 Gemini 3.5 Flash 提升 72 分），同時 Gemini 3.5 Flash-Lite 的表現亦大幅超越 Gemini 3.1 Flash-Lite。

**Gemini 3.5 Flash-Lite 效能表現**
相較於輕量級前代，Gemini 3.5 Flash-Lite 展現顯著的智慧升級與更快的處理速度：
- 智慧水準顯著提升：在智慧指數拿下 36 分，比 Gemini 3.1 Flash-Lite 的 25 分大幅增加 11 分，超越 Mistral Medium 3.5 的 30 分，但略低於 Nemotron 3 Ultra 的 38 分與 DeepSeek V4 Flash（max）的 40 分。
- Agentic 評測大幅進步：在 `GDPval-AA v2` 暴增 498 分來到 1140 分，`TerminalBench v2.1` 提升 22.5 分至 53.6 分，`Tau3-Banking` 則上升 7.8 個百分點至 16.5%。
- 執行時間大幅縮短：平均每個任務執行時間為 0.6 分鐘，較前代的 1.0 分鐘減少將近一半，輸出速度在測試中測得每秒 350 個 token。
- 成本增加：平均每個任務的成本從 0.04 美元倍增至 0.09 美元，主要歸因於新定價策略，儘管平均每個任務消耗的輸出 token 從 2 萬降至 1.3 萬。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/2e38ca24a2801730.jpg)
> 根據 Artificial Analysis 的綜合智力評測，Gemini 3.6 Flash 與 Gemini 3.5 Flash 保持相同的智力表現，而 Gemini 3.5 Flash-Lite 的智力指數相較於 Gemini 3.1 Flash-Lite 顯著提升 11 分。

**模型共通規格與定價細節**
兩款新模型在架構與計費機制上延續了 Google 的基礎設定：
- 視窗大小：兩者皆保留與前代相同的 1M 視窗。
- 多模態支援：皆支援文字、圖片、影片與語音輸入，但輸出僅支援文字。
- 價格與快取優惠：Gemini 3.6 Flash 定價為每百萬輸入／輸出 token 1.50／7.50 美元；Gemini 3.5 Flash-Lite 定價為每百萬輸入／輸出 token 0.30／2.50 美元。兩款模型皆維持對快取輸入 token 提供 90% 的折扣。詳細分析可參閱 [Artificial Analysis 官方網站](https://artificialanalysis.ai/)。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/6ce33a0daf41d16b.png)
> Artificial Analysis 的品牌識別圖像，左側帶有標誌與「Independent analysis of AI」字樣，右側為紫紅相間的條紋圖案。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/d2ac34c5ee7b7681.jpg)
> Google 發布 Gemini 3.6 Flash 與 Gemini 3.5 Flash-Lite，兩者相較前代每任務處理時間均減少一半，其中 Gemini 3.5 Flash-Lite 智力指數大幅提升 11 分，而 Gemini 3.6 Flash 則維持與 3.5 Flash 相同的 50 分智力指數。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/63829f77de71011c.jpg)
> Google 發布 Gemini 3.6 Flash 與 Gemini 3.5 Flash-Lite，其中 Gemini 3.6 Flash 保持與 3.5 Flash 相同的智力得分（50分）但任務成本降至 $0.50，而 Gemini 3.5 Flash-Lite 在智力指數上比 3.1 Flash-Lite 提升 11 分達到 36 分。

## 媒體內容

**Google 發布 Gemini 3.6 Flash 與 Gemini 3.5 Flash-Lite，兩模型每任務解碼時間皆較前代減半，其中 Gemini 3.5 Flash-Lite 智力指數大幅提升 11 分至 36 分，而 Gemini 3.6 Flash 智力指數維持 50 分與前代持平。**

**數據表（1）Artificial Analysis Intelligence Index**

| 項目 | 數值 |
| --- | --- |
| Claude Fable 5 (with fallback) | 60、GPT-5.6 Sol (max)=59、Kimi K3=57、Claude Opus 4.8 (max)=56、GPT-5.6 Terra (max)=55、Grok 4.5 (high)=54、Claude Sonnet 5 (max)=53、GPT-5.6 Luna (max)=51、GLM-5.2 (max)=51、Muse Spark 1.1 (xhigh)=51、Gemini 3.5 Flash=50、Gemini 3.6 Flash=50、Gemini 3.1 Pro Preview=46、Qwen3.7 Max=46、MiniMax-M3=44、DeepSeek V4 Pro (max)=44、MiMo-V2.5-Pro=42、Inkling=41、DeepSeek V4 Flash (max)=40、Nemotron 3 Ultra=38、Gemini 3.5 Flash-Lite=36、Mistral Medium 3.5=30、Claude 4.5 Haiku=30、Gemma 4 31B=29、Gemini 3.1 Flash-Lite=25、gpt-oss-120b (high)=24 |

**數據表（2）Time per Intelligence Index Task**

| 項目 | 數值 |
| --- | --- |
| Gemini 3.5 Flash-Lite | 0.6、Gemini 3.1 Flash-Lite=1.0、Gemini 3.6 Flash=1.3、Qwen3.7 Max=1.6、GPT-5.6 Luna (max)=1.6、Gemini 3.1 Pro Preview=1.6、GPT-5.6 Terra (max)=2.1、gpt-oss-120b (high)=2.2、Claude 4.5 Haiku=2.7、Nemotron 3 Ultra=2.7、Gemini 3.5 Flash=2.7、Muse Spark 1.1 (xhigh)=2.8、Mistral Medium 3.5=3.0、Grok 4.5 (high)=3.3、GLM-5.2 (max)=3.3、GPT-5.5 (xhigh)=3.4、MiniMax-M3=3.9、GPT-5.6 Sol (max)=4.2、Claude Fable 5 (with fallback)=4.9、Gemma 4 31B=5.5、MiMo-V2.5-Pro=5.6、DeepSeek V4 Flash (max)=5.8、Claude Opus 4.8 (max)=6.9、DeepSeek V4 Pro (max)=7.9、Kimi K3=8.5、Claude Sonnet 5 (max)=8.9、Kimi K2.6=10.8 |

**Google 發布 Gemini 3.6 Flash 與 Gemini 3.5 Flash-Lite，兩者相較前代每任務處理時間均減少一半，其中 Gemini 3.5 Flash-Lite 智力指數大幅提升 11 分，而 Gemini 3.6 Flash 則維持與 3.5 Flash 相同的 50 分智力指數。**

**數據表（1）Time per Intelligence Index Task**

| 項目 | 數值 |
| --- | --- |
| Gemini 3.5 Flash-Lite | 0.6 |
| Gemini 3.1 Flash-Lite | 1.0 |
| Gemini 3.6 Flash | 1.3 |
| Qwen3.7 Max | 1.6 |
| GPT-5.6 Luna (max) | 1.6 |
| Gemini 3.1 Pro Preview | 1.6 |
| GPT-5.6 Terra (max) | 2.1 |
| gpt-oss-120b (high) | 2.2 |
| Claude 4.5 Haiku | 2.7 |
| Nemotron 3 Ultra | 2.7 |
| Gemini 3.5 Flash | 2.7 |
| Muse Spark 1.1 (xhigh) | 2.8 |
| Mistral Medium 3.5 | 3.0 |
| Grok 4.5 (high) | 3.3 |
| GLM-5.2 (max) | 3.3 |
| GPT-5.5 (xhigh) | 3.4 |
| MiniMax-M3 | 3.9 |
| GPT-5.6 Sol (max) | 4.2 |
| Claude Fable 5 (with fallback) | 4.9 |
| Gemma 4 31B | 5.5 |
| MiMo-V2.5-Pro | 5.6 |
| DeepSeek V4 Flash (max) | 5.8 |
| Claude Opus 4.8 (max) | 6.9 |
| DeepSeek V4 Pro (max) | 7.9 |
| Kimi K3 | 8.5 |
| Claude Sonnet 5 (max) | 8.9 |
| Kimi K2.6 | 10.8 |

**數據表（2）Intelligence Index vs. Time per Intelligence Index Task**

| 項目 | X | Y |
| --- | --- | --- |
| Gemini 3.5 Flash-Lite | 0.6 | 36 |
| Gemini 3.1 Flash-Lite | 1.0 | 25 |
| Gemini 3.6 Flash | 1.3 | 50 |
| Qwen3.7 Max | 1.6 | 46 |
| Gemini 3.1 Pro Preview | 1.6 | 47 |
| GPT-5.6 Terra (max) | 2.1 | 55 |
| gpt-oss-120b (high) | 2.2 | 24 |
| Claude 4.5 Haiku | 2.7 | 30 |
| Nemotron 3 Ultra | 2.7 | 38 |
| Gemini 3.5 Flash | 2.7 | 50 |
| Muse Spark 1.1 (xhigh) | 2.8 | 51 |
| Mistral Medium 3.5 | 3.0 | 30 |
| Grok 4.5 (high) | 3.3 | 54 |
| GLM-5.2 (max) | 3.3 | 51 |
| GPT-5.5 (xhigh) | 3.4 | 55 |
| MiniMax-M3 | 3.9 | 44 |
| GPT-5.6 Sol (max) | 4.2 | 59 |
| Claude Fable 5 (with fallback) | 4.9 | 60 |
| Gemma 4 31B | 5.5 | 29 |
| MiMo-V2.5-Pro | 5.6 | 42 |
| DeepSeek V4 Flash (max) | 5.8 | 40 |
| Claude Opus 4.8 (max) | 6.9 | 56 |
| DeepSeek V4 Pro (max) | 7.9 | 44 |
| Kimi K3 | 8.5 | 57 |
| Claude Sonnet 5 (max) | 8.9 | 53 |
| Kimi K2.6 | 10.8 | 44 |

**Google 發布 Gemini 3.6 Flash 與 Gemini 3.5 Flash-Lite，其中 Gemini 3.6 Flash 保持與 3.5 Flash 相同的智力得分（50分）但任務成本降至 $0.50，而 Gemini 3.5 Flash-Lite 在智力指數上比 3.1 Flash-Lite 提升 11 分達到 36 分。**

**數據表（1）Cost per Intelligence Index Task**

| 項目 | 數值 |
| --- | --- |
| DeepSeek V4 Flash (max) | $0.02 |
| MiMo-V2.5-Pro | $0.03 |
| Gemini 3.1 Flash-Lite | $0.04 |
| gpt-oss-120b (high) | $0.04 |
| DeepSeek V4 Pro (max) | $0.04 |
| Gemini 3.5 Flash-Lite | $0.09 |
| MiniMax-M3 | $0.12 |
| GPT-5.6 Luna (max) | $0.21 |
| Nemotron 3 Ultra | $0.25 |
| Muse Spark 1.1 (xhigh) | $0.26 |
| Claude 4.5 Haiku | $0.27 |
| Gemini 3.1 Pro Preview | $0.29 |
| Grok 4.5 (high) | $0.31 |
| GLM-5.2 (max) | $0.32 |
| Gemini 3.6 Flash | $0.50 |
| GPT-5.6 Terra (max) | $0.55 |
| Gemini 3.5 Flash | $0.59 |
| Kimi K3 | $0.95 |
| Qwen3.7 Max | $0.98 |
| GPT-5.6 Sol (max) | $1.04 |
| Mistral Medium 3.5 | $1.20 |
| Claude Sonnet 5 (max) | $1.53 |
| Claude Opus 4.8 (max) | $1.80 |
| Claude Fable 5 (with fallback) | $2.75 |

**數據表（2）Intelligence vs. Cost per Intelligence Index Task**

| 項目 | X | Y |
| --- | --- | --- |
| DeepSeek V4 Flash (max) | $0.02 | 40 |
| Gemini 3.1 Flash-Lite | $0.04 | 25 |
| MiMo-V2.5-Pro | $0.03 | 42 |
| DeepSeek V4 Pro (max) | $0.05 | 44 |
| Gemini 3.5 Flash-Lite | $0.09 | 36 |
| MiniMax-M3 | $0.12 | 44 |
| Claude 4.5 Haiku | $0.27 | 29 |
| Nemotron 3 Ultra | $0.25 | 37 |
| Muse Spark 1.1 (xhigh) | $0.26 | 51 |
| GPT-5.6 Luna (max) | $0.21 | 51 |
| Kimi K2.6 | $0.35 | 44 |
| GLM-5.2 (max) | $0.32 | 51 |
| Grok 4.5 (high) | $0.31 | 54 |
| Gemini 3.1 Pro Preview | $0.29 | 46 |
| Gemini 3.6 Flash | $0.50 | 50 |
| Gemini 3.5 Flash | $0.59 | 50 |
| Mistral Medium 3.5 | $0.60 | 30 |
| Qwen3.7 Max | $0.98 | 46 |
| GPT-5.6 Terra (max) | $0.55 | 55 |
| Kimi K3 | $0.95 | 58 |
| GPT-5.6 Sol (max) | $1.04 | 58 |
| GPT-5.5 (xhigh) | $1.00 | 55 |
| Claude Sonnet 5 (max) | $1.53 | 53 |
| Claude Opus 4.8 (max) | $1.80 | 55 |
| Claude Fable 5 (with fallback) | $2.75 | 60 |

**Gemini 3.6 Flash 在 GDPval-AA v2 基準測試取得 1421 分（相較 Gemini 3.5 Flash 提升 72 分），同時 Gemini 3.5 Flash-Lite 的表現亦大幅超越 Gemini 3.1 Flash-Lite。**

**數據表（1）GDPval-AA v2 Leaderboard**

| 項目 | 數值 |
| --- | --- |
| Claude Fable 5 (with fallback) | 1760 |
| GPT-5.6 Sol (max) | 1743 |
| Kimi K3 | 1679 |
| Claude Sonnet 5 (max) | 1607 |
| Claude Opus 4.8 (max) | 1600 |
| GPT-5.6 Luna (max) | 1584 |
| GPT-5.6 Terra (max) | 1581 |
| Grok 4.5 (high) | 1535 |
| GLM-5.2 (max) | 1514 |
| GPT-5.5 (xhigh) | 1493 |
| Gemini 3.6 Flash | 1421 |
| MiniMax-M3 | 1395 |
| Muse Spark 1.1 (xhigh) | 1374 |
| Gemini 3.5 Flash | 1349 |
| DeepSeek V4 Pro (max) | 1307 |
| Qwen3.7 Max | 1273 |
| MiMo-V2.5-Pro | 1265 |
| Inkling | 1239 |
| Kimi K2.6 | 1189 |
| DeepSeek V4 Flash (max) | 1189 |
| Nemotron 3 Ultra | 1164 |
| Gemini 3.5 Flash-Lite | 1140 |
| Gemini 3.1 Pro Preview | 965 |
| Mistral Medium 3.5 | 929 |
| Claude 4.5 Haiku | 907 |
| Gemma 4 31B | 804 |
| gpt-oss-120B (high) | 799 |
| Gemini 3.1 Flash-Lite | 642 |

**數據表（2）AA-Briefcase Elo**

| 項目 | 數值 |
| --- | --- |
| Claude Fable 5 (with fallback) | 1583 |
| Kimi K3 | 1543 |
| GPT-5.6 Sol (max) | 1496 |
| Claude Sonnet 5 (max) | 1388 |
| Claude Opus 4.8 (max) | 1354 |
| Grok 4.5 (high) | 1323 |
| GLM-5.2 (max) | 1260 |
| GPT-5.5 (xhigh) | 1154 |
| MiniMax-M3 | 1110 |
| Gemini 3.6 Flash | 961 |
| DeepSeek V4 Pro (max) | 932 |
| Qwen3.7 Max | 908 |
| MiMo-V2.5-Pro | 873 |
| Nemotron 3 Ultra | 870 |
| Gemini 3.5 Flash | 866 |
| Muse Spark 1.1 (xhigh) | 863 |
| Inkling | 836 |
| DeepSeek V4 Flash (max) | 831 |
| Kimi K2.6 | 816 |
| Gemini 3.5 Flash-Lite | 634 |
| Claude 4.5 Haiku | 603 |
| Mistral Medium 3.5 | 506 |
| Gemini 3.1 Pro Preview | 458 |
| Gemma 4 31B | 364 |
| Gemini 3.1 Flash-Lite | 221 |
| gpt-oss-120B (high) | 0 |

**這張圖彙整 Artificial Analysis 各項獨立評測（GDPval-AA v2、τ³-Banking、Terminal-Bench v2.1、SciCode、Humanity's Last Exam、GPQA Diamond、AA-LCR、MMMU-Pro 等）的逐項分數，橫向比較 Gemini 3.6 Flash、Gemini 3.5 Flash-Lite 與各競品模型在每項基準上的表現。**

**數據表（1）GDPval-AA v2**

| 項目 | 數值 |
| --- | --- |
| Claude Fable 5 (with fallback) | 62% |
| Kimi K3 | 62% |
| GPT-5.6 Sol (max) | 59% |
| GPT-5.6 Luna (max) | 55% |
| Claude 4.8 Opus | 55% |
| Grok 4.5 (high) | 54% |
| GLM-5.2 (max) | 54% |
| GPT-5.5 (high) | 52% |
| Gemini 3.6 Flash | 51% |
| Muse Spark 1.1 (xhigh) | 50% |
| MiniMax-M3 | 46% |
| 120b (high) | 45% |
| Gemini 3.5 Flash | 44% |
| DeepSeek V4 (max) | 42% |
| V4 Pro (max) | 40% |
| Qwen3.7 Max | 39% |
| MiMo-V2.5-Pro | 38% |
| Inkling | 37% |
| DeepSeek V4 Flash (max) | 34% |
| Nemotron 3 Ultra | 33% |
| Gemini 3.1 Flash-Lite | 32% |
| Gemini 4 3IB | 23% |
| Mistral Medium 3.5 | 21% |
| Claude 4.5 Haiku | 20% |
| Gemma 4 31B | 15% |
| gpt-oss-120b (high) | 15% |
| Command A+ | 11% |
| Gemini 3.5 Flash-Lite | 7% |

**數據表（2）τ³-Banking**

| 項目 | 數值 |
| --- | --- |
| Kimi K3 | 33% |
| GPT-5.6 Sol (max) | 33% |
| Grok 4.5 (high) | 33% |
| GPT-5.6 Luna (max) | 32% |
| Claude Fable 5 (with fallback) | 31% |
| Claude 4.8 Opus | 28% |
| GLM-5.2 (max) | 28% |
| DeepSeek V4 (max) | 27% |
| Gemini 3.6 Flash | 27% |
| GPT-5.5 (high) | 27% |
| Gemini 3.5 Flash | 26% |
| Muse Spark 1.1 (xhigh) | 25% |
| MiniMax-M3 | 25% |
| Qwen3.7 Max | 25% |
| Inkling | 24% |
| DeepSeek V4 Flash (max) | 23% |
| Gemini 3.1 Flash-Lite | 16% |
| Gemini 4 3IB | 16% |
| Nemotron 3 Ultra | 15% |
| Mistral Medium 3.5 | 14% |
| Claude 4.5 Haiku | 13% |
| 120b (high) | 12% |
| gpt-oss-120b (high) | 11% |
| Command A+ | 9% |
| Gemma 4 31B | 9% |
| MiMo-V2.5-Pro | 9% |
| Gemini 3.5 Flash-Lite | 6% |

**數據表（3）Terminal-Bench v2.1**

| 項目 | 數值 |
| --- | --- |
| GPT-5.6 Sol (max) | 88% |
| GPT-5.6 Luna (max) | 88% |
| Kimi K3 | 85% |
| Claude 4.8 Opus | 85% |
| Claude Fable 5 (with fallback) | 85% |
| GPT-5.5 (high) | 84% |
| Grok 4.5 (high) | 82% |
| GLM-5.2 (max) | 81% |
| Gemini 3.5 Flash | 81% |
| Muse Spark 1.1 (xhigh) | 79% |
| Gemini 3.6 Flash | 78% |
| Qwen3.7 Max | 78% |
| Gemini 3.1 Flash-Lite | 75% |
| MiMo-V2.5-Pro | 74% |
| DeepSeek V4 Flash (max) | 65% |
| V4 Pro (max) | 65% |
| DeepSeek V4 (max) | 64% |
| Inkling | 62% |
| Gemini 4 3IB | 55% |
| Nemotron 3 Ultra | 54% |
| Gemini 3.5 Flash-Lite | 54% |
| Gemma 4 31B | 51% |
| Claude 4.5 Haiku | 44% |
| Mistral Medium 3.5 | 43% |
| gpt-oss-120b (high) | 31% |
| 120b (high) | 26% |
| Command A+ | 23% |

**數據表（4）SciCode**

| 項目 | 數值 |
| --- | --- |
| Claude Fable 5 (with fallback) | 60% |
| GPT-5.6 Sol (max) | 59% |
| Kimi K3 | 59% |
| Grok 4.5 (high) | 58% |
| GPT-5.6 Luna (max) | 56% |
| GPT-5.5 (high) | 56% |
| Claude 4.8 Opus | 54% |
| Muse Spark 1.1 (xhigh) | 54% |
| Gemini 3.6 Flash | 54% |
| Gemini 3.5 Flash | 53% |
| GLM-5.2 (max) | 53% |
| Qwen3.7 Max | 53% |
| MiMo-V2.5-Pro | 50% |
| V4 Pro (max) | 50% |
| DeepSeek V4 (max) | 50% |
| Inkling | 49% |
| MiniMax-M3 | 46% |
| DeepSeek V4 Flash (max) | 45% |
| Gemma 4 31B | 45% |
| Claude 4.5 Haiku | 43% |
| Gemini 3.1 Flash-Lite | 42% |
| Gemini 3.5 Flash-Lite | 41% |
| Nemotron 3 Ultra | 40% |
| Mistral Medium 3.5 | 40% |
| 120b (high) | 39% |
| Command A+ | 38% |

**數據表（5）Humanity's Last Exam**

| 項目 | 數值 |
| --- | --- |
| Claude Fable 5 (with fallback) | 53% |
| GPT-5.6 Sol (max) | 47% |
| Kimi K3 | 46% |
| GPT-5.6 Luna (max) | 45% |
| Claude 4.8 Opus | 45% |
| Grok 4.5 (high) | 44% |
| GLM-5.2 (max) | 44% |
| Gemini 3.6 Flash | 42% |
| GPT-5.5 (high) | 41% |
| Gemini 3.5 Flash | 40% |
| Muse Spark 1.1 (xhigh) | 40% |
| Qwen3.7 Max | 38% |
| MiniMax-M3 | 38% |
| MiMo-V2.5-Pro | 37% |
| DeepSeek V4 (max) | 37% |
| V4 Pro (max) | 36% |
| DeepSeek V4 Flash (max) | 34% |
| Inkling | 32% |
| Nemotron 3 Ultra | 30% |
| Gemini 4 3IB | 27% |
| Gemma 4 31B | 23% |
| gpt-oss-120b (high) | 18% |
| Gemini 3.1 Flash-Lite | 18% |
| Gemini 3.5 Flash-Lite | 16% |
| Mistral Medium 3.5 | 13% |
| Command A+ | 11% |
| Claude 4.5 Haiku | 10% |

**數據表（6）GPQA Diamond**

| 項目 | 數值 |
| --- | --- |
| Gemini 3.1 Flash-Lite | 94% |
| GPT-5.6 Sol (max) | 94% |
| GPT-5.5 (high) | 94% |
| Kimi K3 | 94% |
| Grok 4.5 (high) | 93% |
| MiniMax-M3 | 93% |
| Claude Fable 5 (with fallback) | 93% |
| GPT-5.6 Luna (max) | 93% |
| Claude 4.8 Opus | 92% |
| Gemini 3.6 Flash | 92% |
| Muse Spark 1.1 (xhigh) | 92% |
| GLM-5.2 (max) | 91% |
| DeepSeek V4 (max) | 90% |
| Gemini 3.5 Flash | 89% |
| DeepSeek V4 Flash (max) | 89% |
| V4 Pro (max) | 89% |
| Inkling | 87% |
| Nemotron 3 Ultra | 87% |
| MiMo-V2.5-Pro | 87% |
| Gemini 4 3IB | 86% |
| Gemini 3.5 Flash-Lite | 84% |
| Gemma 4 31B | 82% |
| Claude 4.5 Haiku | 78% |
| Command A+ | 76% |
| Mistral Medium 3.5 | 75% |
| 120b (high) | 67% |

**數據表（7）CritPT**

| 項目 | 數值 |
| --- | --- |
| GPT-5.6 Sol (max) | 32% |
| GPT-5.5 Pro (max) | 31% |
| GPT-5.6 Luna (max) | 30% |
| Claude Fable 5 (with fallback) | 29% |
| Kimi K3 | 27% |
| GLM-5.2 (max) | 23% |
| Grok 4.5 (high) | 21% |
| Gemini 3.6 Flash | 21% |
| Muse Spark 1.1 (xhigh) | 18% |
| Claude 4.8 Opus | 17% |
| Gemini 3.5 Flash | 15% |
| Qwen3.7 Max | 15% |
| Inkling | 13% |
| MiniMax-M3 | 13% |
| DeepSeek V4 (max) | 11% |
| DeepSeek V4 Flash (max) | 7% |
| MiMo-V2.5-Pro | 5% |
| Nemotron 3 Ultra | 4% |
| Gemini 4 3IB | 4% |
| Gemma 4 31B | 3% |
| Gemini 3.1 Flash-Lite | 1% |
| gpt-oss-120b (high) | 1% |
| Gemini 3.5 Flash-Lite | 1% |
| Command A+ | 0% |
| Claude 4.5 Haiku | 0% |
| Mistral Medium 3.5 | 0% |

**數據表（8）AA-Omniscience Accuracy**

| 項目 | 數值 |
| --- | --- |
| Claude Fable 5 (with fallback) | 61% |
| GPT-5.6 Sol (max) | 59% |
| GPT-5.5 (high) | 57% |
| Kimi K3 | 55% |
| Grok 4.5 (high) | 52% |
| Gemini 3.6 Flash | 52% |
| Gemini 3.5 Flash | 50% |
| Claude 4.8 Opus | 47% |
| GLM-5.2 (max) | 46% |
| Muse Spark 1.1 (xhigh) | 46% |
| Gemini 3.1 Flash-Lite | 43% |
| GPT-5.6 Luna (max) | 42% |
| Inkling | 41% |
| Claude 4.5 Haiku | 40% |
| DeepSeek V4 (max) | 38% |
| DeepSeek V4 Flash (max) | 37% |
| Qwen3.7 Max | 36% |
| Gemini 3.5 Flash-Lite | 30% |
| Mistral Medium 3.5 | 30% |
| GLM-5.2 (max) | 25% |
| MiMo-V2.5-Pro | 25% |
| Nemotron 3 Ultra | 23% |
| gpt-oss-120b (high) | 22% |
| Gemma 4 31B | 22% |
| Claude 4.5 Haiku | 20% |
| Gemini 4 3IB | 17% |
| MiniMax-M3 | 15% |
| Command A+ | 9% |

**數據表（9）AA-Omniscience Non-Hallucination Rate**

| 項目 | 數值 |
| --- | --- |
| Command A+ | 86% |
| MiniMax-M3 | 84% |
| Qwen3.7 Max | 77% |
| MiMo-V2.5-Pro | 75% |
| GLM-5.2 (max) | 74% |
| Claude 4.5 Haiku | 72% |
| Nemotron 3 Ultra | 71% |
| Gemini 3.5 Flash | 66% |
| Gemini 3.6 Flash | 64% |
| Muse Spark 1.1 (xhigh) | 63% |
| Gemini 3.1 Flash-Lite | 62% |
| GPT-5.5 (high) | 50% |
| GPT-5.6 Luna (max) | 49% |
| Grok 4.5 (high) | 46% |
| Claude Fable 5 (with fallback) | 46% |
| Gemini 3.5 Flash-Lite | 45% |
| Kimi K3 | 39% |
| GPT-5.6 Sol (max) | 37% |
| Gemma 4 31B | 18% |
| Gemini 3.1 Flash-Lite | 18% |
| Mistral Medium 3.5 | 18% |
| Claude 4.5 Haiku | 15% |
| GPT-5.6 Sol (max) | 14% |
| Gemini 4 3IB | 11% |
| 120b (high) | 10% |
| DeepSeek V4 (max) | 9% |
| DeepSeek V4 Flash (max) | 6% |
| Command A+ | 4% |

**數據表（10）AA-LCR**

| 項目 | 數值 |
| --- | --- |
| Kimi K3 | 75% |
| GPT-5.5 (high) | 74% |
| MiniMax-M3 | 74% |
| GPT-5.6 Sol (max) | 74% |
| GPT-5.6 Luna (max) | 74% |
| GLM-5.2 (max) | 74% |
| Gemini 3.6 Flash | 73% |
| Gemini 3.5 Flash | 73% |
| Claude Fable 5 (with fallback) | 71% |
| Claude 4.8 Opus | 71% |
| Grok 4.5 (high) | 70% |
| Qwen3.7 Max | 70% |
| Gemini 3.5 Flash-Lite | 70% |
| Inkling | 69% |
| Claude 4.5 Haiku | 69% |
| DeepSeek V4 (max) | 68% |
| DeepSeek V4 Flash (max) | 68% |
| Nemotron 3 Ultra | 67% |
| V4 Pro (max) | 66% |
| Gemma 4 31B | 65% |
| Gemini 3.1 Flash-Lite | 63% |
| DeepSeek V4 Flash (max) | 63% |
| Gemini 4 3IB | 63% |
| gpt-oss-120b (high) | 62% |
| Command A+ | 62% |
| Gemini 3.5 Flash-Lite | 61% |
| Claude 4.5 Haiku | 51% |
| Command A+ | 46% |

**數據表（11）AA-Briefcase**

| 項目 | 數值 |
| --- | --- |
| Claude Fable 5 (with fallback) | 1574 |
| Kimi K3 | 1543 |
| GPT-5.6 Sol (max) | 1496 |
| Claude 4.8 Opus | 1388 |
| Grok 4.5 (high) | 1347 |
| GLM-5.2 (max) | 1323 |
| GPT-5.5 (high) | 1260 |
| Gemini 3.6 Flash | 1154 |
| MiniMax-M3 | 1110 |
| DeepSeek V4 (max) | 961 |
| V4 Pro (max) | 932 |
| Qwen3.7 Max | 908 |
| MiMo-V2.5-Pro | 873 |
| Nemotron 3 Ultra | 870 |
| Gemini 3.5 Flash | 866 |
| Muse Spark 1.1 (xhigh) | 863 |
| Inkling | 836 |
| DeepSeek V4 Flash (max) | 831 |
| Gemini 3.1 Flash-Lite | 634 |
| Mistral Medium 3.5 | 603 |
| Gemini 3.5 Flash-Lite | 506 |
| Gemma 4 31B | 457 |
| Command A+ | 364 |
| Claude 4.5 Haiku | 358 |
| 120b (high) | 221 |

**數據表（12）MMMU-Pro**

| 項目 | 數值 |
| --- | --- |
| Gemini 3.5 Flash | 84% |
| GPT-5.6 Sol (max) | 83% |
| Gemini 3.6 Flash | 83% |
| GPT-5.6 Luna (max) | 82% |
| Kimi K3 | 81% |
| Grok 4.5 (high) | 81% |
| GPT-5.5 (high) | 80% |
| Gemini 3.5 Flash-Lite | 80% |
| MiniMax-M3 | 79% |
| GPT-5.6 Sol (max) | 79% |
| Claude Fable 5 (with fallback) | 79% |
| Claude 4.8 Opus | 77% |
| Gemini 3.1 Flash-Lite | 76% |
| Inkling | 73% |
| Gemma 4 31B | 73% |
| Mistral Medium 3.5 | 65% |
| Command A+ | 63% |
| Claude 4.5 Haiku | 59% |

## 標籤

新產品, 功能更新, LLM, Google, Google DeepMind, Gemini
