# SpaceXAI 推出 Grok 4.6：每百萬輸入 token 2 美元、輸出 token 6 美元，提升長期 Agent 任務能力

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：SpaceXAI (@SpaceXAI) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥🔥🔥🔥🔥 · 日期：2026-08-12

> 原始來源：https://x.com/SpaceXAI/status/2087562800982077492

## 證據與延伸閱讀

- [SpaceXAI 推出 Grok 4.6：每百萬輸入 token 2 美元、輸出 token 6 美元，提升長期 Agent 任務能力。](https://x.ai/news/grok-4-6) — 官方文件
- [Grok 4.6 延續 500k context window](https://artificialanalysis.ai/models/grok-4-6)
- [Artificial Analysis 指標得分 61](https://x.com/ArtificialAnlys/status/2087564648325530099)
- [Grok 4.6 的智慧分數與單項任務成本](https://x.com/ArtificialAnlys/status/2087564652884660250)
- [Arena.ai Code Arena 排行第 7](https://x.com/arena/status/2087566422390231534)

## 中文摘要

SpaceXAI 推出 Grok 4.6：每百萬輸入 token 2 美元、輸出 token 6 美元，提升長期 Agent 任務能力。

**可用管道** Grok 4.6 已於 2026 年 8 月 12 日推出，使用者可在 Grok Build、Cursor、Grok Bot 與 API 中使用，另已透過 OpenRouter、Vercel、Cloudflare 等合作夥伴提供。Cognition 也宣布將模型加入 Devin Desktop 與 Devin CLI。Cursor 與 Grok Build 在首週提供兩倍的包含用量；此外，官方提供速度更快、價格為兩倍的 fast variant。

**主要能力** Grok 4.6 延續 Grok 4.5 的 500k token context window，並將重點放在需要多步驟持續工作的情境，包括研究主題、分析資訊、跨程式庫作業，以及把概念轉成可運作的應用程式或工作成果。官方表示，模型在較長的工作軌跡中會做更多自我測試與驗證，先檢查成果再繼續下一步。

- 面對寬泛的產品構想，Grok 4.6 能研究陌生領域、規劃應用程式結構、實作核心互動，並根據多輪回饋持續修正。
- 在視覺與互動專案上，它能先建立應用程式的結構與視覺語言，再透過 loop 逐步迭代。
- Artificial Analysis 指出，Grok 4.6 在知識工作、終端機操作與客戶服務等 Agent 任務上表現強勁。
- Devin 表示，Grok 4.6 特別擅長在修改程式碼前詳盡探索程式庫與分析根因，也能遵循 repo 慣例並嚴格測試。

**評測結果** Artificial Analysis Intelligence Index 是由九項 benchmark 組成的綜合指標，Grok 4.6 得分 61，與 GPT-5.6 Sol 持平，落後 Claude Opus 5 的 63 分與 Claude Fable 5 的 62 分，略高於 Kimi K3。Grok 4.6 比一個多月前推出的 Grok 4.5 高 5 分；與 Grok 4.3 相比則高 23 分。 

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/ffab5398c99c7a1c.png)
> Grok 4.6 在 Artificial Analysis 智慧指數達到 61 分，與 GPT-5.6 Sol 持平；其他 Agent 與程式開發基準則各有高低。

在更具體的 Agent 評測中，Grok 4.6 的 GDPval-AA v2 Elo 為 1753，僅落後 Claude Opus 5，且與 Claude Fable 5 和 Qwen3.8 Max 的信賴區間重疊；𝜏³-Banking 得分 50.7%，接近 Qwen3.8 Max 的 51.3%；Terminal-Bench v2.1 得分 88.4%，與領先模型同一水準。Artificial Analysis 的私有長期 Agent 知識工作 benchmark AA-Briefcase 中，Grok 4.6 得到 Elo 1577，落後 Claude Opus 5 系列，但在評分規準、簡報呈現與分析品質上都維持穩定表現。 

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/be699302c7f64405.jpg)
> 來源：[@ArtificialAnlys](https://x.com/ArtificialAnlys/status/2087564650477244722)（回覆）｜Grok 4.6 (high) 相較於 Grok 4.5 獲得顯著提升，在 GDPval-AA v2、τ³-Banking 及 Terminal-Bench v2.1 各項基準測試中均位居頂尖前列並超越 GPT-5.6 Sol。

**成本與效率** Grok 4.6 的標準價格維持在每 100 萬 input token 2 美元、output token 6 美元，與 Grok 4.5 相同。這比 Claude Opus 5 的 5／25 美元與 GPT-5.6 Sol 的 5／30 美元低 60% 以上；每項任務成本約 0.84 美元，與 Kimi K3 相同，但 Intelligence Index 略高，因此落在「智慧程度與單項任務成本」的 Pareto frontier。快取命中價格為每 100 萬 token 0.5 美元，高於 Grok 4.5 的 0.3 美元。 

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/779872ae5c29159c.jpg)
> 來源：[@ArtificialAnlys](https://x.com/ArtificialAnlys/status/2087564652884660250)（回覆）｜Grok 4.6 在 Artificial Analysis Intelligence Index 獲得 61 分，達到與 GPT-5.6 Sol 相當的前沿水準，同時具備更低的每項任務成本。

在 AA-Briefcase 中，Grok 4.6 平均約以 53 個 turn、5 億個 input token 完成任務；Claude Opus 5（max）則約需 103 個 turn、20 億個 input token，顯示 Grok 4.6 在完成長期工作時具有較高的 turn 效率。 

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/7297059d4e91d85d.jpg)
> 來源：[@ArtificialAnlys](https://x.com/ArtificialAnlys/status/2087564654965125375)（回覆）｜Grok 4.6 (high) 在 AA-Briefcase Elo 取得 1577 分，較前代 Grok 4.5 (high) 的 1313 分大幅提升，超越 GPT-5.6 Sol 與 Kimi K3，僅次於 Claude Opus 5 (max)。

**外部排行** Arena.ai 的 Code Arena: WebDev 排行顯示，Grok 4.6（High）以 1618 分排名第 7，明顯高於 Grok 4.5 的第 13 名與 1553 分；它與 GPT-5.6 Sol（xHigh）的 1622 分、Claude Fable 5 的 1627 分僅差 4 至 9 分，落在第 5 至第 7 名的密集區間。Arena.ai 也提醒，隨著更多投票加入、信賴區間收窄，排名仍可能變得更清楚。 

<video src="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/1786565960624-3f6xp8sg.mp4" poster="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/4cf3540ce9ec5e0f.jpg" controls playsinline preload="metadata" style="max-width:100%;height:auto;display:block;margin:1rem 0"></video>
> 來源：[@arena](https://x.com/arena/status/2087566422390231534)（回覆）｜Code Arena WebDev 排行榜，列出 Grok 4.6 排名第 7 的評比結果長條圖

**訓練與安全** 根據 SpaceXAI 的發布說明，Grok 4.6 採用比 Grok 4.5 更長的補充訓練，加入模型產生且經整理的推理與進階技術概念資料、高品質工程資料，以及改良的 optimizer 與訓練配方。團隊也使用 Grok 4.5 重新產生不同推理投入程度、Agent harness 與 STEM、軟體工程、知識工作領域的 SFT 軌跡，再以模型檢查排除問題軌跡；後續則以知識工作、一般程式開發、核心最佳化、Web 開發與電腦輔助設計等環境進行 Agentic RL 訓練。 

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/3206ede9095f2dae.jpg)
> 來源：[@cognition](https://x.com/cognition/status/2087579582492987881)（回覆）｜Grok 4.6 在 FrontierCode 1.1 Extended 取得 61.3 分，相較 Grok 4.5 大幅提升並超越 GPT-5.6 Sol，僅次於 Claude Opus 5 與 Claude Fable 5

SpaceXAI 表示，Grok 4.6 的安全防護已配合能力提升重新校準，並完成迄今範圍最廣的上市前能力與防護測試，也做了上市後及第三方測試；但這些仍屬發布方對訓練與安全性的說明，實際效果仍需持續觀察。 

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/17b95ce96497e17b.jpg)
> Grok 4.6 的深灰色背景標題畫面，中央以白字顯示版本名稱「Grok 4.6」，背景帶有細緻的顆粒質感與右側泛光的流動曲線光影。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/63d19784270de0fc.jpg)
> 來源：[@ArtificialAnlys](https://x.com/ArtificialAnlys/status/2087564648325530099)（回覆）｜Grok 4.6 在 Artificial Analysis Intelligence Index 取得 61 分，較前代 Grok 4.5 提升 5 分並追平 GPT-5.6 Sol。

## 媒體內容

**Grok 4.6 在 Artificial Analysis 智慧指數達到 61 分，與 GPT-5.6 Sol 持平；其他 Agent 與程式開發基準則各有高低。**

**數據表**

|   | Grok 4.6 High | Grok 4.5 High | GPT-5.6 Sol Max | Fable 5 Max |
| --- | --- | --- | --- | --- |
| AA Intelligence Index | 61 | 56 | 61 | 62 |
| GDPVal-AA v2 | 1753 | 1526 | 1728 | 1741 |
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% |
| DeepSWE v1.1 | 65.9% | 54% | 73% | 70% |
| FrontierCode v1.1 (Extended) | 61.3% | 56.6% | 60.6% | 64.9% |
| APEX-Agents | 57.5% | 47.1% | 56.7% | 59.2% |
| Terminal-Bench v3.0 | 26% | 15.7% | 34.6% | 34.1% |
| APEX-SWE | 56.4% | 53.6% | – | 58.8% |
| AA-Briefcase | 1577 | 1313 | 1502 | 1574 |
| Harvey LAB (Vals) | 15.8% | 12.9% | 2.5% | 11.3% |

**Grok 4.6 在 FrontierCode 1.1 Extended 取得 61.3 分，相較 Grok 4.5 大幅提升並超越 GPT-5.6 Sol，僅次於 Claude Opus 5 與 Claude Fable 5**

**數據表**

| 項目 | 數值 |
| --- | --- |
| SWE-1.7 | 54.3 |
| GPT-5.6 Terra | 55.8 |
| Claude Sonnet 5 | 56.2 |
| Grok 4.5 | 56.5 |
| GPT-5.5 | 56.7 |
| Kimi K3 | 58.2 |
| Claude Opus 4.8 | 59.6 |
| GPT-5.6 Sol | 60.6 |
| Grok 4.6 | 61.3 |
| Claude Opus 5 | 63.6 |
| Claude Fable 5 | 64.9 |

**Grok 4.6 在 Artificial Analysis Intelligence Index 取得 61 分，較前代 Grok 4.5 提升 5 分並追平 GPT-5.6 Sol。**

**數據表**

| 項目 | 數值 |
| --- | --- |
| Claude Opus 5 (max) | 63 |
| Claude Fable 5 (with fallback) | 62 |
| GPT-5.6 Sol (max) | 61 |
| Grok 4.6 (high) | 61 |
| Kimi K3 (max) | 60 |
| Qwen3.8 Max | 58 |
| Muse Spark 1.2 (xhigh) | 57 |
| GPT-5.6 Terra (max) | 57 |
| Grok 4.5 (high) | 56 |
| Claude Sonnet 5 (max) | 55 |
| GLM-5.2 (max) | 53 |
| GPT-5.6 Luna (max) | 52 |
| DeepSeek V4 Flash 0731 (max) | 52 |
| Gemini 3.6 Flash | 52 |
| MiniMax-M3 | 45 |
| MiMo-V2.5-Pro | 43 |
| Inkling | 42 |
| Nemotron 3 Ultra | 38 |
| Gemini 3.5 Flash-Lite | 37 |
| Muse Glimmer (high) | 35 |
| Mistral Medium 3.5 | 30 |
| Gemma 4 31B | 30 |

**Grok 4.6 (high) 相較於 Grok 4.5 獲得顯著提升，在 GDPval-AA v2、τ³-Banking 及 Terminal-Bench v2.1 各項基準測試中均位居頂尖前列並超越 GPT-5.6 Sol。**

**數據表（1）GDPval-AA v2 Leaderboard**

| 項目 | 數值 |
| --- | --- |
| Claude Opus 5 (max) | 1849 |
| Grok 4.6 (high) | 1753 |
| Claude Fable 5 (with fallback) | 1741 |
| Qwen3.8 Max | 1737 |
| GPT-5.6 Sol (max) | 1728 |
| Kimi K3 (max) | 1682 |
| Muse Spark 1.2 (xhigh) | 1628 |
| Claude Sonnet 5 (max) | 1598 |
| GPT-5.6 Luna (max) | 1581 |
| GPT-5.6 Terra (max) | 1578 |
| DeepSeek V4 Flash 0731 (max) | 1558 |
| Grok 4.5 (high) | 1526 |
| GLM-5.2 (max) | 1506 |
| Gemini 3.6 Flash | 1422 |
| MiniMax-M3 | 1389 |
| MiMo-V2.5-Pro | 1266 |
| Inkling | 1240 |
| Nemotron 3 Ultra | 1163 |
| Gemini 3.5 Flash-Lite | 1140 |
| Muse Glimmer (high) | 953 |
| Mistral Medium 3.5 | 933 |
| Gemma 4 31B | 811 |

**數據表（2）τ³-Banking: Score**

| 項目 | 數值 |
| --- | --- |
| Qwen3.8 Max | 51.3% |
| Grok 4.6 (high) | 50.7% |
| Kimi K3 (max) | 46.0% |
| GPT-5.6 Sol (max) | 44.3% |
| Claude Opus 5 (max) | 42.1% |
| Grok 4.5 (high) | 42.1% |
| GPT-5.6 Terra (max) | 40.2% |
| DeepSeek V4 Flash 0731 (max) | 39.4% |
| Claude Fable 5 (with fallback) | 38.1% |
| Claude Sonnet 5 (max) | 37.3% |
| Muse Spark 1.2 (xhigh) | 34.8% |
| GLM-5.2 (max) | 34.6% |
| GPT-5.6 Luna (max) | 31.1% |
| Gemini 3.6 Flash | 29.9% |
| Inkling | 29.1% |
| Muse Glimmer (high) | 23.5% |
| Gemini 3.5 Flash-Lite | 17.5% |
| MiniMax-M3 | 15.3% |
| Mistral Medium 3.5 | 15.1% |
| Gemma 4 31B | 14.8% |
| Nemotron 3 Ultra | 14.2% |
| MiMo-V2.5-Pro | 9.9% |

**數據表（3）Terminal-Bench v2.1: Score**

| 項目 | 數值 |
| --- | --- |
| Claude Opus 5 (max) | 89.1% |
| Grok 4.6 (high) | 88.4% |
| GPT-5.6 Terra (max) | 88.0% |
| GPT-5.6 Sol (max) | 88.0% |
| Kimi K3 (max) | 85.0% |
| Claude Fable 5 (with fallback) | 84.6% |
| Grok 4.5 (high) | 81.6% |
| Qwen3.8 Max | 81.3% |
| GPT-5.6 Luna (max) | 80.9% |
| Claude Sonnet 5 (max) | 80.5% |
| Muse Spark 1.2 (xhigh) | 80.1% |
| DeepSeek V4 Flash 0731 (max) | 78.7% |
| GLM-5.2 (max) | 77.9% |
| Gemini 3.6 Flash | 77.5% |
| MiniMax-M3 | 65.2% |
| MiMo-V2.5-Pro | 65.2% |
| Inkling | 55.1% |
| Nemotron 3 Ultra | 53.9% |
| Gemini 3.5 Flash-Lite | 53.6% |
| Muse Glimmer (high) | 51.7% |
| Mistral Medium 3.5 | 50.6% |
| Gemma 4 31B | 43.4% |

**Grok 4.6 在 Artificial Analysis Intelligence Index 獲得 61 分，達到與 GPT-5.6 Sol 相當的前沿水準，同時具備更低的每項任務成本。**

**數據表**

| 項目 | X | Y |
| --- | --- | --- |
| DeepSeek V4 Flash 0731 (max) | $0.027 | 52 |
| MiMo-V2.5-Pro | $0.035 | 43 |
| GPT-5.6 Luna (max) | $0.05 | 52.5 |
| Gemini 3.5 Flash-Lite | $0.09 | 37.5 |
| MiniMax-M3 | $0.15 | 45.5 |
| Grok 4.5 (high) | $0.31 | 56 |
| GLM-5.2 (max) | $0.31 | 53 |
| Muse Spark 1.2 (xhigh) | $0.33 | 57 |
| Inkling | $0.34 | 42 |
| Nemotron 3 Ultra | $0.38 | 38.5 |
| Mistral Medium 3.5 | $0.45 | 30.5 |
| GPT-5.6 Terra (max) | $0.5 | 56.5 |
| Gemini 3.6 Flash | $0.55 | 52 |
| Kimi K3 (max) | $0.65 | 60.5 |
| Grok 4.6 (high) | $0.7 | 61 |
| Qwen3.8 Max | $0.95 | 58 |
| GPT-5.6 Sol (max) | $1.2 | 61 |
| Claude Sonnet 5 (max) | $1.7 | 55 |
| Claude Opus 5 (max) | $2.2 | 63 |
| Claude Fable 5 (with fallback) | $2.3 | 62 |

**Grok 4.6 (high) 在 AA-Briefcase Elo 取得 1577 分，較前代 Grok 4.5 (high) 的 1313 分大幅提升，超越 GPT-5.6 Sol 與 Kimi K3，僅次於 Claude Opus 5 (max)。**

**數據表**

| 項目 | 數值 |
| --- | --- |
| Claude Opus 5 (max) | 1715、Grok 4.6 (high)=1577、Claude Fable 5 (with fallback)=1574、Kimi K3 (max)=1541、GPT-5.6 Sol (max)=1502、Qwen3.8 Max=1420、Claude Sonnet 5 (max)=1383、Muse Spark 1.2 (xhigh)=1358、Grok 4.5 (high)=1313、DeepSeek V4 Flash 0731 (max)=1286、GLM-5.2 (max)=1252、MiniMax-M3=1107、Gemini 3.6 Flash=963、MiMo-V2.5-Pro=880、Nemotron 3 Ultra=874、Inkling=842、Gemini 3.5 Flash-Lite=635、Mistral Medium 3.5=517、Gemma 4 31B=374 |

## 標籤

新產品, 功能更新, Agent, SpaceXAI, Grok
