# DeepSWE v1.1：GLM-5.3 max 的 Pass@1 為 69%，接近 Claude Fable 5 max 的 70%，單次成本則為 3.99 美元

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：Together AI (@togethercompute) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥 · 日期：2026-08-23

> 原始來源：https://x.com/togethercompute/status/2091316030047941085

## 證據與延伸閱讀

- [DeepSWE v1.1：GLM-5.3 max 的 Pass@1 為 69%，接近 Claude Fable 5 max 的 70%，單次成本則為 3.99 美元。](https://deepswe.datacurve.ai/) — 官方文件 · 最後核對：2026-08-23 · 支持主張：官方 DeepSWE v1.1 leaderboard 補上 113 個 long-horizon engineering tasks、91 個 repositories、更新日期與 confidence intervals，並列出 Claude Fable 5 max 與 GLM-5.3 max 的 Pass@1、平均成本、輸出 tokens 和 steps；@zainhas 再提供 two- and four-attempt 的 rounded figures，但該分析必須按作者歸屬，不能視為獨立重現的 leaderboard 結果。
- [GLM-5.3 與 Claude Fable 5 的 DeepSWE v1.1 benchmark 比較 — @zainhas](https://x.com/zainhas/status/2091297526347677701) — 一手來源 · 最後核對：2026-08-23 · 支持主張：官方 DeepSWE v1.1 leaderboard 補上 113 個 long-horizon engineering tasks、91 個 repositories、更新日期與 confidence intervals，並列出 Claude Fable 5 max 與 GLM-5.3 max 的 Pass@1、平均成本、輸出 tokens 和 steps；@zainhas 再提供 two- and four-attempt 的 rounded figures，但該分析必須按作者歸屬，不能視為獨立重現的 leaderboard 結果。

## 中文摘要

DeepSWE v1.1：GLM-5.3 max 的 Pass@1 為 69%，接近 Claude Fable 5 max 的 70%，單次成本則為 3.99 美元。

**評測範圍** 官方榜單於 2026 年 8 月 20 日更新，涵蓋 91 個儲存庫中的 113 項原創長程軟體工程任務，並提供信賴區間，讓單次比較不只呈現單一百分比。

**單次結果**
- Claude Fable 5 max：Pass@1 為 70%±4%，每項任務平均成本 21.63 美元，輸出 119k tokens，執行 88 個步驟。
- GLM-5.3 max：Pass@1 為 69%±3%，每項任務平均成本 3.99 美元，輸出 80k tokens，執行 124 個步驟。

官方榜單四捨五入後顯示 GLM-5.3 為 69%、Claude Fable 5 為 70%；@zainhas 引用的較精確數字則是 69.0% 與 69.7%。GLM-5.3 每項任務約 3.99 美元，Claude Fable 5 約 21 美元，成本差距明顯。

**多次嘗試** @zainhas 的分析指出，兩次嘗試時 GLM-5.3 為 81.1%、Claude Fable 5 為 77.1%；四次嘗試則為 87.6% 對 84.1%。他認為 GLM-5.3 的優勢來自「差一點成功」的案例比例較高，因此更能從重試中獲益。不過，這是作者自行發布的分析，並非經獨立重現的官方榜單結果。

## 標籤

Benchmark, LLM, 研究論文, Anthropic
