# GPT-6 Astra 列 Terminal-Bench 4.0 第 1，成本約為 Fable 5.1 max 一半

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：Tibo (@thsottiaux) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥🔥🔥🔥 · 日期：2026-09-05

> 原始來源：https://x.com/thsottiaux/status/2095995318471168397

## 證據與延伸閱讀

- [GPT-6 Astra 列 Terminal-Bench 4.0 第 1，成本約為 Fable 5.1 max 一半。](https://openai.com/index/gpt-6-astra) — 官方文件 · 最後核對：2026-09-05 · 支持主張：Adds the raw official leaderboard rows behind the post: both max runs have 330 trials; Astra is rank 1 at 58.18% with a 2.79-point 95% confidence half-width and total_cost_usd 3267.18, while Fable 5.1 is rank 2 at 57.88% with a 3.76-point half-width and total_cost_usd 6243.5. The raw-cost ratio is about 52.3% (3267.18/6243.50), so Tibo’s 50% wording is an approximate summary rather than the exact API ratio. Terminal-Bench 4.0 also changed resources and task membership, making the version and ru…
- [GPT-6 Astra on Terminal-Bench 4.0 with the Codex harness — tbench.ai](https://tbench.ai/leaderboard/terminal-bench/4.0) — 官方文件 · 最後核對：2026-09-05 · 支持主張：Adds the raw official leaderboard rows behind the post: both max runs have 330 trials; Astra is rank 1 at 58.18% with a 2.79-point 95% confidence half-width and total_cost_usd 3267.18, while Fable 5.1 is rank 2 at 57.88% with a 3.76-point half-width and total_cost_usd 6243.5. The raw-cost ratio is about 52.3% (3267.18/6243.50), so Tibo’s 50% wording is an approximate summary rather than the exact API ratio. Terminal-Bench 4.0 also changed resources and task membership, making the version and ru…
- [GPT-6 Astra on Terminal-Bench 4.0 with the Codex harness — tbench.ai](https://tbench.ai/news/terminal-bench-4-0) — 官方文件 · 最後核對：2026-09-05 · 支持主張：Adds the raw official leaderboard rows behind the post: both max runs have 330 trials; Astra is rank 1 at 58.18% with a 2.79-point 95% confidence half-width and total_cost_usd 3267.18, while Fable 5.1 is rank 2 at 57.88% with a 3.76-point half-width and total_cost_usd 6243.5. The raw-cost ratio is about 52.3% (3267.18/6243.50), so Tibo’s 50% wording is an approximate summary rather than the exact API ratio. Terminal-Bench 4.0 also changed resources and task membership, making the version and ru…
- [GPT-6 Astra on Terminal-Bench 4.0 with the Codex harness — ofhuhcpkvzjlejydnvyd.supabase.co](https://ofhuhcpkvzjlejydnvyd.supabase.co/functions/v1/leaderboard-read?package=terminal-bench/terminal-bench&name=4-0-0) — 官方文件 · 最後核對：2026-09-05 · 支持主張：Adds the raw official leaderboard rows behind the post: both max runs have 330 trials; Astra is rank 1 at 58.18% with a 2.79-point 95% confidence half-width and total_cost_usd 3267.18, while Fable 5.1 is rank 2 at 57.88% with a 3.76-point half-width and total_cost_usd 6243.5. The raw-cost ratio is about 52.3% (3267.18/6243.50), so Tibo’s 50% wording is an approximate summary rather than the exact API ratio. Terminal-Bench 4.0 also changed resources and task membership, making the version and ru…
- [Tibo 指出 GPT-6 Astra 拿第一名且成本是第二名 50%](https://x.com/thsottiaux/status/2095995318471168397) — 一手來源

## 中文摘要

GPT-6 Astra 列 Terminal-Bench 4.0 第 1，成本約為 Fable 5.1 max 一半。

**排行榜資料** Tibo 於 9 月 4 日表示，Astra 使用 Codex 執行框架拿下第一名，成本是第二名的「50%」。[官方排行榜 API](https://ofhuhcpkvzjlejydnvyd.supabase.co/functions/v1/leaderboard-read?package=terminal-bench/terminal-bench&name=4-0-0) 的原始列則顯示，兩個 `max` 推理強度執行各完成 330 次試驗：

- GPT‑6 Astra 搭配 Codex：排名第 1，解決率 58.18%，95% 信賴區間半寬 2.79 個百分點，成功 192 次，`total_cost_usd` 為 3267.18。
- Fable 5.1 搭配 Claude Code：排名第 2，解決率 57.88%，95% 信賴區間半寬 3.76 個百分點，成功 191 次，`total_cost_usd` 為 6243.5。

API 另列 Astra 的 `xhigh` 與 `high` 執行並列第 2，成本分別為 2350.51 與 2269.42 美元；本文的成本比較只針對 Astra 與 Fable 5.1 的兩個 `max` 執行。

以 API 的精確成本計算，3267.18 ÷ 6243.50 約為 52.3%，所以貼文的「50%」是約略說法，不能視為原始資料的精確比例。兩個 95% 信賴區間重疊；來源沒有提供顯著性檢定，不能把 58.18% 與 57.88% 的差距宣稱為統計顯著。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/5fffc58cd1095fd2.jpg)
> GPT-6 Astra 搭配 Codex harness 在 Terminal-Bench 4.0 以 58.2% 的解決率名列第 1，Fable 5.1 為 57.9%。成本資料另見 Tibo 貼文與官方 API：Tibo 概括為第 2 名的 50%，API 精確成本比約為 52.3%。

**比較邊界** 這項結果測量的是指定代理程式、模型、`max` 推理強度、執行日期、330 次試驗與 Terminal-Bench 4.0 資源下的終端機任務解決率，不能延伸成整體程式開發品質。Astra 的資料日期是 2026 年 9 月 3 日，Fable 5.1 是 2026 年 9 月 1 日。排行榜提供彙總解決率、信賴區間、token 與成本，但這份資料沒有每個任務的失敗類型，也沒有完整執行框架設定。

**Terminal-Bench 4.0 的變更** [更新說明](https://tbench.ai/news/terminal-bench-4-0)指出，4.0 同時改動了評測資源與任務集合：

- 重新校準任務的時間、CPU 與記憶體，所有任務改用固定 8 小時代理程式逾時。
- 移除 8 個任務：2 個因已飽和、2 個因拒答、2 個因已有公開解答，另有 2 個涉及尚未解決的品質或平台相容性問題。
- 修正 19 個任務，更新指示、環境與驗證器，處理回報的不穩定與規格錯誤。
- 相較 3.0，4.0 的代理程式逾時與錯誤較少；剩餘錯誤主要是模型拒答和輸出 token 超過上限。

Terminal-Bench 現在採用持續更新的評測與語意化版本管理。資源設定和任務集合都改變，官方將此視為破壞性變更，要求重新執行試驗；4.0 的排名不能直接當成 3.0 或更早版本的縱向進步證據。

## 標籤

功能更新, Codex, Claude Code
