# Z AI 發布 GLM-5.3：Artificial Analysis 評 60 分，權重釋出後將並列開放權重模型最高分

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：Artificial Analysis (@ArtificialAnlys) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥🔥🔥🔥🔥 · 日期：2026-08-19

> 原始來源：https://x.com/ArtificialAnlys/status/2089830890709135426

## 證據與延伸閱讀

- [Z AI 發布 GLM-5.3：Artificial Analysis 評 60 分，權重釋出後將並列開放權重模型最高分。](https://x.com/ArtificialAnlys/status/2089830890709135426)
- [Intelligence Index取得60分](https://artificialanalysis.ai/) — 一手來源
- [輸出token數18,700、成本0.68美元](https://artificialanalysis.ai/models/glm-5-3) — 一手來源
- [AA-Omniscience分數從4升至14分](https://x.com/ArtificialAnlys/status/2089830895956181198)

## 中文摘要

Z AI 發布 GLM-5.3：Artificial Analysis 評 60 分，權重釋出後將並列開放權重模型最高分。

這次更新最大的進步在 Agentic 能力與實際知識準確度，但 token 使用量與部署成本也同步上升。

**模型定位** 2026 年 8 月 19 日，Z AI 發布 GLM-5.3，目前可透過 Z AI 的 first-party API 使用，團隊表示預計在一週內釋出權重。Artificial Analysis 指出，GLM-5.3 的總參數量仍為 753B，MoE 架構下的 active parameters 為 40B，與 GLM-5.2 相同；模型具備 1M token 的 context window，採 MIT 授權。若權重如期公開，GLM-5.3 與 Kimi K3 將代表開放權重模型與 proprietary frontier 之間的差距進一步縮小。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/9bf94ac6bee6d9f3.jpg)
> GLM-5.3 在 Artificial Analysis Intelligence Index 取得 60 分，與 Kimi K3 持平，並較前代 GLM-5.2 提升 7 分。

**Agentic 能力** GLM-5.3 在 GDPval-AA v2——用於評估真實世界 Agentic 知識工作的測試——取得最明顯的提升：

- Elo 從 GLM-5.2 的 1524 上升至 1770，成長 246 分。
- 在所有受測模型中排名第二；同系列貼文文字將 Claude Opus 5 記為 1855，圖表則顯示 1845。
- 它超越先前開放權重領先者 Kimi K3 的 1668，差距超過 100 分。
- Artificial Analysis 因此認為，GLM-5.3 的 Agentic 表現已進入 frontier models 的行列，而不只是一般問答能力改善。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/b7eb5bc0bf44bebf.jpg)
> 圖表顯示 GLM-5.3 為 1769 分、Claude Opus 5 為 1845 分；同系列貼文文字則分別記為 1770 與 1855，兩組各相差 1 分與 10 分。

**效率與成本** GLM-5.3 的能力提升伴隨較低的 token 效率。Artificial Analysis Intelligence Index v4.1 顯示，它每項任務約使用 18,700 個輸出 token，高於 GLM-5.2 的 15,700 個，亦比 Kimi K3 的 14,700 個多 27%。這使 GLM-5.3 每項 Intelligence Index 任務的成本為 0.68 美元，是 GLM-5.2 0.44 美元的 1.5 倍；不過在相同智能層級中仍較便宜：

- 比 Kimi K3 的 0.84 美元低 19%。
- 比 GPT-5.6 Sol 的 1.23 美元低 45%。

因此，GLM-5.3 的部署成本上升，部分原因是 token 使用量成長約 20%，但整體單項任務價格仍保有競爭力。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/581dbbc16b21bab6.jpg)
> GLM-5.3 在 GDPval-AA v2 分項取得 63%，為圖中開放權重模型最高；高於 Kimi K3 的 59% 與 GLM-5.2 的 48%。

**知識準確度** 在 AA-Omniscience 測試中，GLM-5.3 從 GLM-5.2 的 4 分提升至 14 分，成為僅次於 Kimi K3（20 分）的第二佳開放權重模型。這項進步不只是更常拒答所造成：準確率由 24% 提升至 34%，嘗試回答的比例也從 46% 上升至 55%。不過，幻覺率同時由 26% 小幅回升至 30%；Artificial Analysis 的另一則貼文則將原始準確率記為 23%，但同樣報告提升至 34%，顯示摘要資料中的起始數字存在一個百分點差異。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/47dc01d27725d9e9.jpg)
> GLM-5.3 在 AA-Omniscience Index 取得 14 分，相較於 GLM-5.2 的 4 分有所提升。

**API 與定價** GLM-5.3 的標準價格如下，快取輸入 token 另有折扣：

- 輸入：每 1M token 收費 1.40 美元。
- 輸出：每 1M token 收費 4.40 美元。
- 快取輸入：套用 81% 的 cache hit 折扣後，每 1M token 收費 0.26 美元。

完整分析可參考 [Artificial Analysis 的 GLM-5.3 模型頁面](https://artificialanalysis.ai/models/glm-5-3)，整體評測資訊則見 [Artificial Analysis](https://artificialanalysis.ai/)。

## 媒體內容

**GLM-5.3 在 Artificial Analysis Intelligence Index 取得 60 分，與 Kimi K3 持平，並較前代 GLM-5.2 提升 7 分。**

**數據表（1）Artificial Analysis Intelligence Index**

| 項目 | 數值 |
| --- | --- |
| Claude Opus 5 (max) | 63 |
| Claude Fable 5 (with fallback) | 62 |
| GPT-5.6 Sol (max) | 61 |
| Kimi K3 (max) | 60 |
| GLM-5.3 (max) | 60 |
| Qwen3.8 Max | 58 |
| Claude Opus 4.8 (max) | 57 |
| GPT-5.6 Terra (max) | 57 |
| Grok 4.5 (high) | 56 |
| Claude Sonnet 5 (max) | 55 |
| Muse Spark 1.1 (xhigh) | 53 |
| GLM-5.2 (max) | 53 |
| GPT-5.6 Luna (max) | 52 |
| DeepSeek V4 Flash 0731 (max) | 52 |
| Gemini 3.6 Flash | 52 |
| Qwen3.7 Max | 47 |
| MiniMax-M3 | 45 |
| MiMo-V2.5-Pro | 43 |
| Inkling | 42 |
| Nemotron 3 Ultra | 38 |
| Gemini 3.5 Flash-Lite | 37 |
| Mistral Medium 3.5 | 30 |
| Claude 4.5 Haiku | 30 |
| Gemma 4 31B | 30 |
| gpt-oss-120b (high) | 24 |
| Command A+ | 23 |

**數據表（2）Intelligence Index vs. Cost per Intelligence Index Task**

| 項目 | X | Y |
| --- | --- | --- |
| MiMo-V2.5-Pro | $0.035 | 43 |
| DeepSeek V4 Flash 0731 (max) | $0.05 | 52 |
| GPT-5.6 Luna (max) | $0.06 | 52 |
| gpt-oss-120b (high) | $0.07 | 24 |
| Gemini 3.5 Flash-Lite | $0.09 | 37 |
| MiniMax-M3 | $0.15 | 45 |
| Claude 4.5 Haiku | $0.22 | 30 |
| Gemini 3.6 Flash | $0.28 | 52 |
| Muse Spark 1.1 (xhigh) | $0.29 | 53 |
| Grok 4.5 (high) | $0.33 | 56 |
| GPT-5.6 Terra (max) | $0.38 | 57 |
| Inkling | $0.39 | 42 |
| Nemotron 3 Ultra | $0.42 | 38 |
| GLM-5.2 (max) | $0.44 | 53 |
| Mistral Medium 3.5 | $0.51 | 30 |
| Qwen3.7 Max | $0.52 | 47 |
| GLM-5.3 (max) | $0.68 | 60 |
| Kimi K3 (max) | $0.84 | 60 |
| Qwen3.8 Max | $1.10 | 58 |
| GPT-5.6 Sol (max) | $1.23 | 61 |
| Claude Sonnet 5 (max) | $1.70 | 55 |
| Claude Fable 5 (with fallback) | $2.00 | 62 |
| Claude Opus 4.8 (max) | $2.30 | 57 |
| Claude Opus 5 (max) | $3.10 | 63 |

**圖表顯示 GLM-5.3 為 1769 分、Claude Opus 5 為 1845 分；同系列貼文文字則分別記為 1770 與 1855，兩組各相差 1 分與 10 分。**

**數據表**

| 項目 | 數值 |
| --- | --- |
| Claude Opus 5 (max) | 1845 |
| GLM-5.3 (max) | 1769 |
| Grok 4.6 (high) | 1747 |
| Claude Fable 5 (with fallback) | 1738 |
| Qwen3.8 Max | 1735 |
| GPT-5.6 Sol (max) | 1723 |
| Kimi K3 (max) | 1681 |
| Muse Spark 1.2 (xhigh) | 1628 |
| DeepSeek V4 Pro 0813 (max) | 1590 |
| GPT-5.6 Luna (max) | 1578 |
| GPT-5.6 Terra (max) | 1576 |
| Qwen3.8 27B | 1546 |
| Gemini 3.7 Flash (high) | 1532 |
| GLM-5.2 (max) | 1505 |
| MiniMax-M3 | 1387 |
| Inkling | 1239 |
| Nemotron 3 Ultra | 1163 |
| Gemini 3.5 Flash-Lite | 1140 |
| Muse Glimmer (high) | 953 |
| Mistral Medium 3.5 | 934 |
| Claude 4.5 Haiku | 913 |
| Nemotron 3.5 Lightning | 824 |
| gpt-oss-120b (high) | 800 |
| Command A+ | 717 |

**GLM-5.3 在 AA-Omniscience Index 取得 14 分，相較於 GLM-5.2 的 4 分有所提升。**

**數據表（1）AA-Omniscience Index**

| 項目 | 數值 |
| --- | --- |
| Claude Fable 5 (with fallback) | 43 |
| Claude Opus 5 (max) | 37 |
| Claude Opus 4.8 (max) | 29 |
| Muse Spark 1.1 (xhigh) | 28 |
| Grok 4.5 (high) | 25 |
| Gemini 3.6 Flash | 22 |
| GPT-5.6 Sol (max) | 22 |
| Kimi K3 (max) | 20 |
| Claude Sonnet 5 (max) | 16 |
| GLM-5.3 (max) | 14 |
| Qwen3.7 Max | 13 |
| Gemini 3.5 Flash-Lite | 5 |
| GLM-5.2 (max) | 4 |
| Qwen3.8 Max | 3 |
| MiMo-V2.5-Pro | 3 |
| Inkling | 2 |
| MiniMax-M3 | 1 |
| GPT-5.6 Terra (max) | 0 |
| Nemotron 3 Ultra | 0 |
| Command A+ | -4 |
| Claude 4.5 Haiku | -4 |
| GPT-5.6 Luna (max) | -10 |
| DeepSeek V4 Flash 0731 (max) | -14 |
| Mistral Medium 3.5 | -37 |
| Gemma 4 31B | -48 |
| gpt-oss-120b (high) | -49 |

**數據表（2）AA-Omniscience Accuracy**

| 項目 | 數值 |
| --- | --- |
| Claude Fable 5 (with fallback) | 65% |
| Claude Opus 5 (max) | 61% |
| GPT-5.6 Sol (max) | 59% |
| Muse Spark 1.1 (xhigh) | 52% |
| Grok 4.5 (high) | 52% |
| Gemini 3.6 Flash | 50% |
| Claude Opus 4.8 (max) | 49% |
| Kimi K3 (max) | 48% |
| GPT-5.6 Terra (max) | 47% |
| GPT-5.6 Luna (max) | 43% |
| Inkling | 42% |
| DeepSeek V4 Flash 0731 (max) | 40% |
| Claude Sonnet 5 (max) | 40% |
| GLM-5.3 (max) | 34% |
| Qwen3.8 Max | 32% |
| Qwen3.7 Max | 31% |
| Gemini 3.5 Flash-Lite | 29% |
| Mistral Medium 3.5 | 25% |
| GLM-5.2 (max) | 24% |
| Nemotron 3 Ultra | 23% |
| MiMo-V2.5-Pro | 22% |
| gpt-oss-120b (high) | 22% |
| Gemma 4 31B | 20% |
| Claude 4.5 Haiku | 18% |
| MiniMax-M3 | 17% |
| Command A+ | 9% |

**數據表（3）AA-Omniscience Hallucination Rate**

| 項目 | 數值 |
| --- | --- |
| Command A+ | 14% |
| MiniMax-M3 | 18% |
| MiMo-V2.5-Pro | 25% |
| Qwen3.7 Max | 26% |
| GLM-5.2 (max) | 26% |
| Claude 4.5 Haiku | 27% |
| GLM-5.3 (max) | 30% |
| Nemotron 3 Ultra | 30% |
| Gemini 3.5 Flash-Lite | 34% |
| Claude Opus 4.8 (max) | 39% |
| Claude Sonnet 5 (max) | 39% |
| Qwen3.8 Max | 42% |
| Muse Spark 1.1 (xhigh) | 50% |
| Kimi K3 (max) | 53% |
| Grok 4.5 (high) | 54% |
| Gemini 3.6 Flash | 56% |
| Claude Opus 5 (max) | 61% |
| Claude Fable 5 (with fallback) | 64% |
| Inkling | 68% |
| Mistral Medium 3.5 | 82% |
| Gemma 4 31B | 85% |
| GPT-5.6 Terra (max) | 88% |
| gpt-oss-120b (high) | 91% |
| DeepSeek V4 Flash 0731 (max) | 92% |
| GPT-5.6 Sol (max) | 92% |
| GPT-5.6 Luna (max) | 93% |

**GLM-5.3 在 GDPval-AA v2 分項取得 63%，為圖中開放權重模型最高；高於 Kimi K3 的 59% 與 GLM-5.2 的 48%。**

**數據表（1）GDPval-AA v2**

| 項目 | 數值 |
| --- | --- |
| Claude Opus 5 | 67% |
| Claude 5 (high) | 66% |
| GLM-5.3 (max) | 63% |
| Grok 4.6 (high) | 62% |
| Claude Fable 5 (with fallback) | 62% |
| Qwen3.8 Max | 62% |
| 5 Opus | 61% |
| Qwen3 85B | 61% |
| 2.4T A95B | 59% |
| Kimi K3 (max) | 59% |
| GPT-5.6 Sol (xhigh) | 56% |
| Claude Opus | 56% |
| Muse 3 (spark) | 56% |
| GPT-5.6 Sol (high) | 55% |
| GPT-5.5 (high) | 55% |
| Claude 4.5 Opus | 54% |
| DeepSeek V4 | 54% |
| Pro 0813 (max) | 54% |
| 4.8 (xhigh) | 53% |
| Grok 4.5 (high) | 52% |
| Luna (max) | 51% |
| Terra (max) | 50% |
| Sol (high) | 50% |
| GPT-5.5 (medium) | 50% |
| Flash (medium) | 49% |
| GLM-5.2 (max) | 48% |
| GPT-5.5 (max) | 44% |
| Muse 1 (spark) | 33% |
| Nemotron 3 Ultra | ...% |

**數據表（2）τ³-Banking**

| 項目 | 數值 |
| --- | --- |
| Qwen3.8 Max | 51% |
| Grok 4.6 (high) | 51% |
| GLM-5.3 (max) | 50% |
| Qwen 1.2T A95B | 49% |
| Kimi K3 (max) | 46% |
| 2.4T A95B | 45% |
| Claude Opus 5 | 44% |
| 5 Opus | 43% |
| GPT-5.6 Sol (max) | 42% |
| Claude Opus | 42% |
| Grok 4.5 (high) | 40% |
| 5 (max) | 40% |
| Pro 0813 (max) | 39% |
| GPT-5.5 (max) | 39% |
| DeepSeek V4 | 38% |
| Claude 5 (high) | 38% |
| Claude Fable 5 (with fallback) | 37% |
| GPT-5.6 Sol (high) | 37% |
| Claude 4.5 Opus | 36% |
| GPT-5.5 (high) | 35% |
| Sol (high) | 35% |
| GPT-5.5 (medium) | 35% |
| Flash (medium) | 35% |
| Muse 3.7 (spark) | 34% |
| 1.2 (spark) | 33% |
| GLM-5.2 (max) | 32% |
| 4.8 (high) | 31% |
| Nemotron 3 Ultra | 14% |

**數據表（3）Terminal-Bench v2.1**

| 項目 | 數值 |
| --- | --- |
| GPT-5.6 Sol (xhigh) | 90% |
| Claude Opus 5 | 89% |
| Grok 4.6 (high) | 88% |
| GLM-5.3 (max) | 88% |
| GPT-5.6 Sol (high) | 88% |
| Claude 5 (high) | 88% |
| 5 Opus | 87% |
| Claude Fable 5 (with fallback) | 86% |
| Claude (medium) | 86% |
| GPT-5.6 Sol (max) | 86% |
| Flash (xhigh) | 85% |
| Kimi K3 (max) | 85% |
| Muse 3 (spark) | 84% |
| Sol (high) | 84% |
| GPT-5.5 (high) | 83% |
| Claude Fable 5 | 82% |
| GPT-5.5 (max) | 82% |
| 4.8 (high) | 81% |
| 4.8 (xhigh) | 81% |
| GLM-5.3 (medium) | 81% |
| 4.8 Opus | 80% |
| Claude Opus | 79% |
| 2.4T A95B | 79% |
| Grok 4.5 (high) | 78% |
| Qwen3.8 Max | 78% |
| Nemotron 3 Ultra | 54% |

**數據表（4）SciCode**

| 項目 | 數值 |
| --- | --- |
| Claude Fable 5 (with fallback) | 60% |
| Kimi K3 (max) | 59% |
| Flash (xhigh) | 58% |
| Muse 3.7 (spark) | 58% |
| GPT-5.6 (medium) | 57% |
| Flash (medium) | 57% |
| GLM-5.3 (max) | 56% |
| 5 Opus | 56% |
| GPT-5.6 Sol (high) | 56% |
| 1.2 (spark) | 56% |
| Muse 3 (spark) | 56% |
| GPT-5.5 (xhigh) | 56% |
| GPT-5.6 Sol (xhigh) | 55% |
| GPT-5.5 (high) | 55% |
| GPT-5.6 Sol (max) | 54% |
| Claude Opus | 54% |
| Claude 5 (high) | 54% |
| Claude Fable 5 | 54% |
| 4.7 Opus | 53% |
| Sol (high) | 53% |
| Grok 4.5 (high) | 53% |
| Terra (max) | 52% |
| Grok 4.6 (high) | 51% |
| Claude 4.5 Opus | 50% |
| Claude Spinnet | 49% |
| 4.8 (high) | 40% |

**數據表（5）Humanity's Last Exam**

| 項目 | 數值 |
| --- | --- |
| Claude Fable 5 (with fallback) | 55% |
| GPT-5.6 Sol (xhigh) | 55% |
| Claude Opus 5 | 54% |
| Claude 5 (high) | 53% |
| GPT-5.6 Sol (max) | 51% |
| Claude Opus | 49% |
| GLM-5.3 (max) | 49% |
| 4.8 (high) | 48% |
| Flash (xhigh) | 47% |
| Muse 3.7 (spark) | 47% |
| 1.1 (spark) | 46% |
| Kimi K3 (max) | 46% |
| GPT-5.6 Sol (high) | 46% |
| GPT-5.5 (xhigh) | 45% |
| Muse 3 (spark) | 45% |
| 1.2 (spark) | 43% |
| GPT-5.5 (high) | 43% |
| 4.8 (high) | 43% |
| Grok 4.6 (high) | 42% |
| Grok 4.5 (high) | 42% |
| 4.1 Opus | 42% |
| Terra (max) | 41% |
| Qwen3.8 Max | 41% |
| Nemotron 3 Ultra | 28% |

**數據表（6）GPQA Diamond**

| 項目 | 數值 |
| --- | --- |
| Grok 4.6 (high) | 95% |
| Muse 3.7 (spark) | 95% |
| GLM-5.3 (max) | 94% |
| GPT-5.6 Sol (xhigh) | 94% |
| Claude Opus 5 | 94% |
| 5 Opus | 94% |
| Flash (xhigh) | 93% |
| Qwen3.8 Max | 93% |
| 2.4T A95B | 93% |
| Kimi K3 (max) | 93% |
| GPT-5.5 (high) | 93% |
| Claude Opus | 93% |
| Claude 5 (high) | 93% |
| Grok 4.5 (high) | 93% |
| DeepSeek V4 | 93% |
| Pro 0813 (max) | 92% |
| GPT-5.6 Sol (high) | 92% |
| Claude Fable 5 (with fallback) | 92% |
| Sol (high) | 92% |
| GPT-5.5 (max) | 91% |
| Claude (medium) | 91% |
| Terra (max) | 90% |
| Flash (medium) | 90% |
| Muse 3 (spark) | 89% |
| Nemotron 3 Ultra | 87% |

**數據表（7）CritPt**

| 項目 | 數值 |
| --- | --- |
| GPT-5.6 Sol (max) | 32% |
| Terra (max) | 30% |
| Claude 5 (high) | 29% |
| Claude Opus 5 | 29% |
| GPT-5.6 Sol (xhigh) | 29% |
| Claude Fable 5 (with fallback) | 28% |
| GLM-5.3 (max) | 28% |
| 5 Opus | 27% |
| GPT-5.5 (high) | 27% |
| Claude Opus | 26% |
| GPT-5.6 Sol (high) | 25% |
| Kimi K3 (max) | 23% |
| GPT-5.5 (high) | 23% |
| Sol (medium) | 21% |
| GLM-5.2 (max) | 21% |
| Claude 4.5 Opus | 21% |
| 4.8 (high) | 20% |
| Luna (max) | 20% |
| Qwen3.8 Max | 19% |
| Grok 4.6 (high) | 18% |
| Nemotron 3 Ultra | 3% |

**數據表（8）AA-Omniscience Accuracy**

| 項目 | 數值 |
| --- | --- |
| Claude Fable 5 (with fallback) | 65% |
| Claude Opus 5 | 61% |
| Claude 5 (high) | 60% |
| Claude Opus | 59% |
| GPT-5.6 Sol (max) | 59% |
| GLM-5.3 (max) | 59% |
| GPT-5.6 Sol (xhigh) | 58% |
| 5 Opus | 58% |
| GPT-5.5 (high) | 58% |
| Sol (medium) | 57% |
| GPT-5.6 Sol (high) | 57% |
| 5 (medium) | 55% |
| GPT-5.5 (max) | 54% |
| Grok 4.6 (high) | 52% |
| Gemini 3.7 | 52% |
| Flash (xhigh) | 49% |
| Flash (medium) | 49% |
| Grok 4.5 (high) | 48% |
| DeepSeek V4 | 48% |
| Pro 0813 (max) | 47% |
| Claude 4.8 Opus | 45% |
| Grok 4.6 (high) | 43% |
| Kimi K3 (max) | 40% |
| Terra (max) | 34% |
| GLM-5.2 (max) | 32% |
| Nemotron 3 Ultra | 23% |

**數據表（9）AA-Omniscience Non-Hallucination Rate**

| 項目 | 數值 |
| --- | --- |
| GLM-5.2 (max) | 74% |
| GLM-5.3 (max) | 70% |
| Nemotron 3 Ultra | 70% |
| Qwen3.8 Max | 67% |
| Grok 4.6 (high) | 66% |
| 1.2 (spark) | 61% |
| 2.4T A95B | 61% |
| 4.8 (high) | 61% |
| Claude Opus | 58% |
| Claude 5 (high) | 58% |
| Qwen3.8 Max | 50% |
| Kimi K3 (max) | 47% |
| Grok 4.5 (high) | 46% |
| Claude Opus | 40% |
| Nemotron 3 Ultra | 7% |

**數據表（10）AA-LCR**

| 項目 | 數值 |
| --- | --- |
| Muse 1 (spark) | 83% |
| Kimi K3 (max) | 83% |
| 1.2 (spark) | 81% |
| Muse 3.7 (spark) | 81% |
| 1.1 (spark) | 80% |
| Flash (medium) | 80% |
| GPT-5.6 Sol (xhigh) | 79% |
| GLM-5.3 (max) | 79% |
| Terra (max) | 79% |
| Flash (xhigh) | 78% |
| GPT-5.5 (high) | 78% |
| Sol (high) | 77% |
| 5 (medium) | 77% |
| Claude Spinnet | 77% |
| Claude Fable 5 (with fallback) | 76% |
| GLM-5.3 (high) | 76% |
| Claude Opus | 76% |
| GPT-5.6 Sol (high) | 76% |
| Claude 5 (high) | 76% |
| 5 Opus | 75% |
| DeepSeek V4 | 75% |
| Pro 0813 (max) | 75% |
| GPT-5.6 Sol (max) | 75% |
| GPT-5.5 (max) | 74% |
| 4.8 (high) | 74% |
| Grok 4.6 (high) | 73% |
| Nemotron 3 Ultra | 71% |

## 標籤

新產品, 功能更新, LLM, Z AI
