# GPT-6 Astra 在 ARC-AGI-3 得分 99.9%，與 Sol 採不同評測框架

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：Tibo (@thsottiaux) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥🔥🔥🔥 · 日期：2026-09-05

> 原始來源：https://x.com/thsottiaux/status/2095601101701820752

## 證據與延伸閱讀

- [GPT-6 Astra 在 ARC-AGI-3 得分 99.9%，與 Sol 採不同評測框架。](https://openai.com/index/gpt-6-astra) — 官方文件 · 最後核對：2026-09-05 · 支持主張：Adds a benchmark artifact and harness context: the image caption says ARC-AGI-3 tests learning on unfamiliar interactive tasks, the average human tester scored 48%, Astra was measured with a Responses API harness, and a separate Responses-harness estimate puts Sol at about 30%. The visible 7.8% Sol bar is therefore a different reported figure whose harness/context must be labeled; the chart methodology and ranking were not independently verified.
- [Tibo 分享 ARC-AGI-3 比較圖](https://x.com/thsottiaux/status/2095601101701820752) — 一手來源

## 中文摘要

GPT-6 Astra 在 ARC-AGI-3 得分 99.9%，與 Sol 採不同評測框架。

**圖表數據** 2026 年 9 月 3 日，Tibo 分享的比較圖標示 GPT‑6 Astra 99.9%、Claude Opus 5 30.2%、GPT‑5.6 Sol 7.8%；圖說另列平均人類測試者 48%。[OpenAI 官方發布頁](https://openai.com/index/gpt-6-astra)也列出 Astra 在 ARC-AGI-3 得分 99.9%，並寫到它在 96% 的關卡超過人類行動效率基準。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/379fbb323422086f.jpg)
> GPT-6 Astra 在 ARC-AGI-3 得分 99.9%；圖中另列 Claude Opus 5（30.2%）與 GPT-5.6 Sol（7.8%）。Astra 使用 Responses API 執行框架，與圖中 Sol 的測法不同；OpenAI 估計 Sol 若採同一框架，得分約為 30%。

**評測脈絡** ARC-AGI-3 測試模型學習陌生的互動任務；Astra 的 99.9% 是透過 Responses API 執行框架測得。圖說另提供同一執行框架下 GPT‑5.6 Sol 約 30% 的估計，因此圖表的 7.8% 是不同脈絡下的另一個數值，不能和約 30% 混用。原始資料沒有 ARC-AGI-3 的執行設定、樣本數或信賴區間，圖表的方法與排名也尚未獲獨立驗證。

**解讀方式** Tibo 問「我們需要另一個 AGI 評測，下一個門檻又會移到哪裡？」這是對評測目標是否移動的評論，不是數字的額外證明。圖表只能說明 Astra 在指定任務和執行框架下的分數；比較時應同時記錄模型分數、人類基準、Responses API 執行框架、任務設計與計分規則，不能據此推論模型已具備可泛化的 AGI 能力。

## 標籤

新產品, 研究論文, GPT-6 Astra, OpenAI
