# Anthropic 發布 Claude Opus 5，價格與 Opus 4.8 相同，智慧接近價格兩倍的 Fable 5

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：Claude (@claudeai) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥🔥🔥🔥🔥 · 日期：2026-07-24

> 原始來源：https://x.com/claudeai/status/2080699495453528290

## 中文摘要

Anthropic 發布 Claude Opus 5，價格與 Opus 4.8 相同，智慧接近價格兩倍的 Fable 5。

Anthropic 正式推出 [Claude Opus 5](https://www.anthropic.com/news/claude-opus-5)：定價維持與前一代 Opus 4.8 相同（每百萬輸入 token 5 美元、每百萬輸出 token 25 美元），卻以 Fable 5 一半的價格提供接近其前沿水準的智慧。它預設啟用思考、強化了推理與 Agentic 程式開發，上下文視窗 1M token 既是預設也是上限（沒有更小的變體），單次最多輸出 128k token。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/6f34c7b3ee360c53.jpg)
> Claude Opus 5 的標題畫面

**模型效能與基準表現**
- 在 Frontier-Bench、GDPval-AA 等程式開發與知識工作評測中創下全新的 state-of-the-art，且以較低成本達到前一代 Opus 4.8 超過兩倍的表現；官方同時註明，它在網路安全任務上仍落後 Mythos 5。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/44e0c935d936c3b2.webp)
> Opus 5 在多項編程與知識工作評測上達到前沿 SOTA，並在 ARC-AGI-3 創新問題解決測試中以 30.2% 分數遠超其他模型

- 在 Frontier-Bench v0.1 上，各個 effort 等級的成本與得分都壓過同級對手。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/554d79e081111013.webp)
> Opus 5、Fable 5、Opus 4.8 與 GPT-5.6 Sol 在 Frontier-Bench v0.1 依努力程度劃分的代理編程成本與得分比較。

- 在 ARC-AGI-3 評測中，其分數達到次佳模型的三倍以上。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/372157ec2ca6567a.webp)
> Opus 5 (high) 在 ARC-AGI-3 評測中取得約 30% 得分，達到次佳模型 GPT-5.6 Sol 的三倍以上。

- 在 CursorBench 3.2 的 max effort 下，表現與 Fable 5 的 peak score 相差不到 0.5%，但單一任務成本僅有一半。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/8895de5efc0304d3.webp)
> Opus 5、Fable 5、Opus 4.8 與 GPT-5.6 Sol 在 CursorBench 上不同 effort level 下的成本與分數比較

- 在電腦操作評測 OSWorld 2.0 上，以約三分之一的成本超越 Fable 5 的最佳成績。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/aa4681cd18d401aa.webp)
> Opus 5、Fable 5、Opus 4.8 與 GPT-5.6 Sol 在 OSWorld 2.0 評測中不同 effort level 下的任務成本與分數比較。

- 在 Zapier 的 AutomationBench 上，相同單任務成本下的通過率約為次佳模型的 1.5 倍；即使在最低 effort 設定，通過的任務數也多於其他所有模型。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/6e807829b99620db.png)
> Opus 5、Fable 5、Opus 4.8 與 GPT-5.6 Sol 在 AutomationBench 評測中各工作量設定下的任務成本與通過率比較

- 在真實知識工作評測 GDPval-AA v2 與 Humanity's Last Exam 上，以相近或更低的單次成本換到更高的正確率。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/7df9a0337b33c8fe.webp)
> Opus 5、Fable 5、Opus 4.8 與 GPT-5.6 Sol 在 GDPval-AA v2 基準測試中，不同 effort level 下運算成本與 Elo 分數的比較

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/8934ca13c0288ea2.webp)
> Opus 5 在 Humanity's Last Exam (with tools) 基準測試中，以類似或更低的每任務成本展現優於 Fable 5 與 Opus 4.8 的解題正確率

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/93016b8e11ec6dfb.webp)
> Opus 5、Fable 5、Opus 4.8 與 GPT-5.6 Sol 在不同 Effort Level 下的 Artificial Analysis Coding Agent Index 表現與成本比較

- 在生命科學評測中全面超越 Opus 4.8，例如有機化學內部評測高出 10.2 個百分點，蛋白質序列變異功能預測高出 7.7 個百分點，並能視覺化氣流流過空氣動力學物件的情況。

**官方公布的三個實例**
- FreeCAD 任務中模型看不到圖面，於是自己寫了一條電腦視覺 pipeline 從原始像素抽出幾何，重建整個機械零件；相同設定下沒有其他競品模型能在五次嘗試內解出。
- 面對一個開源套件管理器的真實 bug，Opus 5 找出根本原因，補掉社群 patch 漏掉的邊界情況；競品模型只修掉表面症狀就回報問題已解決。
- 一家交易公司的工程師在單一 session 內建出新的交易所行情 feed；找不到 live feed 可對照時，Opus 5 自己建了一套測試 harness 驗證解析是否正確。

**核心架構與行為變更**
- 預設啟用思考功能（thinking），模型會自行決定每個回合的思考時機與深度，開發者可透過 [effort 參數](https://platform.claude.com/docs/en/build-with-claude/effort)（支援 `low`、`medium`、`high`、`xhigh`、`max`）控制思考深度；在 Claude Code 與 Claude Platform 上預設為 high。
- 在 `xhigh` 或 `max` 的 effort 等級下，設定 `thinking: {"type": "disabled"}` 會回傳 400 錯誤，這是一項重要的行為變更。官方也提醒關閉思考有已知瑕疵：模型偶爾會把工具呼叫寫進純文字回覆裡，而不產生結構化的 `tool_use` 區塊。
- 支援 Fast mode，執行速度約為預設速度的 2.5 倍，定價為基礎價格的兩倍；此功能仍在研究預覽階段，目前只在 Claude Platform 與 Claude Code 提供，Amazon Bedrock、Google Cloud 與 Microsoft Foundry 尚未支援。
- 最低可快取 prompt 長度降至 512 tokens，過往因過短而無法快取的 prompt 現在無需修改程式碼即可建立快取條目。
- 行為方面，其預設回應與書面交付內容更長，在代理式對話中更常向使用者敘述進度，並會主動驗證自身工作——也因此官方建議刪掉沿用自舊模型的驗證指令，那些指令會造成過度驗證、白白消耗 token。

**API 與整合設定**
- 支援完整的 effort 階梯，執行高階運算時需設定較大的 `max_tokens`（例如 64000）以提供模型思考空間；要注意 `max_tokens` 是「思考＋回覆文字」的總量硬上限：
  ```bash cURL
  curl https://api.anthropic.com/v1/messages \
    -H "x-api-key: $ANTHROPIC_API_KEY" \
    -H "anthropic-version: 2023-06-01" \
    -H "content-type: application/json" \
    -d '{
      "model": "claude-opus-5",
      "max_tokens": 64000,
      "stream": true,
      "output_config": {
        "effort": "max"
      },
      "messages": [
        {
          "role": "user",
          "content": "Explain why the sum of two even numbers is always even."
        }
      ]
    }'
  ```

- 新增對話中途工具變更（Mid-conversation tool changes）beta 功能，允許在對話回合之間新增或移除工具同時保留 prompt 快取，請求時須帶入 `mid-conversation-tool-changes-2026-07-01` beta header。
- 支援伺服器端預設 fallbacks 模式，依拒絕類別套用 Anthropic 建議的備用模型，須帶入 `server-side-fallback-2026-07-01` beta header。
- 遷移至新版本時，開發者需將程式碼中的模型 ID 更新為 `claude-opus-5`（詳見[遷移指南](https://platform.claude.com/docs/en/about-claude/models/migration-guide#migrating-from-claude-opus-4-8-to-claude-opus-5)）。

**誰能用、在哪裡能用**
- 所有付費方案與 Claude API 當日開放：在 Claude Max 上是新的預設模型，在 Claude Pro 上是最強的可選模型。
- 三大雲平台同步供應，Amazon Bedrock 的模型 ID 為 `anthropic.claude-opus-5`，Google Cloud 與 Microsoft Foundry 亦可使用；Opus 4.8 在這些平台上仍然保留。
- Claude Code 需升級到 v2.1.219 以上才選得到 Opus 5（執行 `claude update`）；Max、Team Premium、Enterprise 隨用隨付與 Anthropic API 預設即為 Opus 5，Pro、Team Standard 與 Enterprise 訂閱席次則預設 Sonnet 5。
- 企業客戶要留意一項限制：[Priority Tier 不支援 Opus 5](https://platform.claude.com/docs/en/about-claude/models/migration-guide)，Opus 4.8 才保留這項承諾，容量需要另外規劃。
- Claude Code 的 fast mode 用 `/fast` 切換（VS Code 擴充套件不支援），訂閱方案上只走 usage credits、不計入方案的額度；而且在對話進行到一半才開啟時，整段既有 context 會以 fast 模式的未快取價重算一次，越晚開越貴。

**安全與對齊**
- 在自動化行為審查中，Opus 5 的整體失準行為得分為 2.3，是近期模型中最低的一個。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/616887dae9fc5807.png)
> Opus 5 在自動化行為審查中取得最低的 2.30 分失準行為得分，展現最高程度的模型對齊。

- 網路安全方面，Opus 5 找出漏洞的能力已逼近 Mythos 5，但把漏洞轉成實際威脅（exploit 開發）的能力仍大幅落後。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/da9a45c75c6b48e6.png)
> Opus 5 在網路安全任務上強於 Opus 4.8，但在漏洞利用開發方面仍大幅落後 Mythos 5。

- 安全防護也隨之調整：cyber classifier 的介入頻率預期比 Fable 5 少約 85%，允許在原始碼中尋找漏洞，但仍擋下 binary 層的漏洞掃描、滲透測試與 exploit 生成；已加入 Cyber Verification Program 的企業與研究者，可取得限制較少的版本。
- 被標記的請求預設回退到 Opus 4.8；原本在 Fable 5 上被擋下的生物領域請求，現在改由 Opus 5 承接（先前是 Opus 4.8）。在 Claude Code 裡，Opus 5 的資安類請求會改跑 Opus 4.8，生物類請求則直接以拒絕收場——因為 Opus 5 自己執行生物 classifier，沒有備援模型。
- 依 Anthropic 的 System Card，Opus 5 被判定具備 CB-1（非新型武器合成）但不具 CB-2 能力，因此沿用與 Opus 4.8 相同的 ASL-3 防護，也未跨過 RSP 設定的自動化 AI 研發能力門檻。
- 英國 AI Security Institute（UK AISI）取得早期檢查點做外部測試，結論是 Opus 5 在其網路安全評測上「表現與 Mythos 5 及 Mythos Preview 相當」：在企業網路攻擊模擬靶場「The Last Ones」十次嘗試中八次端到端攻破，在防禦更完整的新靶場「Doing Life」推進到目前觀測到最遠的第 22 步（全長 23 步），但仍未攻破。

**生態系與各大平台支援**
- 在 Claude Code 中，執行 `/model claude-opus-5` 即可切換；若要將現有工作負載升級，可執行 `/claude-api migrate` 指令，內建的 claude-api skill 會先確認要改動的範圍，再更新模型字串並建議針對該模型微調的提示詞。
- Cognition 宣布 Claude Opus 5 已在 Devin Desktop 與 Devin CLI 上線，並將納入 Devin Cloud 的模式組合；在自家 FrontierCode 1.1 基準測試中拿到 63.6% 分數、Extended 通過率 69.6%，以半價逼近 Fable 5，特別擅長困難的除錯與根本原因分析，修 bug 時偏好精準的原地修正而非大規模重構（詳見 [Devin 公告](https://devin.ai/blog/claude-opus-5)）。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/906605275865a567.jpg)
> Claude Opus 5 在 FrontierCode 1.1 Extended 評測中獲得 63.6% 的分數，超越 GPT-5.6 Sol 與 Claude Opus 4.8，僅次於 Claude Fable 5

- Cursor 已上架該模型，在預設 effort 下 CursorBench 得分 66.7 對 Fable 5 的 66.5、價格只要一半；而且與 Fable 5 不同的是，它相容於 Zero Data Retention 政策（評測比較見 [Cursor 評測頁面](http://cursor.com/evals)）。
- GitHub 宣布 Claude Opus 5 已開始在 GitHub Copilot 分批推出（官方註明推出是漸進式的），開放給 Copilot Pro+、Max、Business 與 Enterprise 方案，可在 Visual Studio Code、Visual Studio、Copilot CLI、cloud agent 與行動平台的模型選單中選用；Business 與 Enterprise 需由管理員先在 Copilot 設定中啟用 Claude Opus 5 政策（詳見 [GitHub 變更日誌](https://github.blog/changelog/2026-07-24-claude-opus-5-is-now-available-in-github-copilot/)）。

<video src="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/1784914970570-uacwnpq9.mp4" poster="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/511b4a26b220facd.jpg" controls playsinline preload="metadata" style="max-width:100%;height:auto;display:block;margin:1rem 0"></video>
> GitHub Copilot 整合 Claude Opus 5 模型介面與功能展示

- Lovable 也已引入該模型，在完整應用程式的建構測試中，成品品質是其測得過的最高水準、每次執行之間的變異明顯較小，最嚴苛的程式任務較 Claude Opus 4.7 提升 22%。

<video src="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/1784914839108-279zi0ln.mp4" poster="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/3e6dba2eeb9ebb7b.jpg" controls playsinline preload="metadata" style="max-width:100%;height:auto;display:block;margin:1rem 0"></video>
> 各種斑點鳥蛋從上方依序落下並排列成數字「5」形狀，最後顯示「Opus 5」字樣。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/2996edc310736831.webp)
> 由多顆不同花紋與尺寸的鳥蛋排列組合成數字「5」的復古風格插圖

---

*此篇由 Claude Opus 5 模型自動撰寫*


## 媒體內容

**Opus 5 在自動化行為審查中取得最低的 2.30 分失準行為得分，展現最高程度的模型對齊。**

**數據表**

| 項目 | 數值 |
| --- | --- |
| Opus 4.8 | 2.85 |
| Mythos 5 | 2.81 |
| Sonnet 5 | 3.35 |
| Opus 5 | 2.30 |

**Opus 5 在網路安全任務上強於 Opus 4.8，但在漏洞利用開發方面仍大幅落後 Mythos 5。**

**數據表（1）Vulnerability identification (grade > 0)**

|   | Pass@1 (% of challenges) |
| --- | --- |
| Opus 4.8 | 61.5% |
| Mythos 5 | 80.0% |
| Opus 5 | 79.4% |

**數據表（2）Exploitation success (grade 1.0)**

|   | Challenges (count) |
| --- | --- |
| Opus 4.8 | 0 |
| Mythos 5 | 13 |
| Opus 5 | 4 |

**Claude Opus 5 在 FrontierCode 1.1 Extended 評測中獲得 63.6% 的分數，超越 GPT-5.6 Sol 與 Claude Opus 4.8，僅次於 Claude Fable 5**

**數據表**

| 項目 | 數值 |
| --- | --- |
| Claude Sonnet 4.6 | 40.0 |
| SWE-1.7 | 54.6 |
| GPT-5.6 Terra | 55.8 |
| Claude Sonnet 5 | 56.2 |
| GPT-5.5 | 56.7 |
| Claude Opus 4.8 | 59.6 |
| GPT-5.6 Sol | 60.6 |
| Claude Opus 5 | 63.6 |
| Claude Fable 5 | 64.9 |

**GitHub Copilot 整合 Claude Opus 5 模型介面與功能展示**

**影片中的 Prompt 與操作**

操作步驟：

1. （00:00）顯示 GitHub Copilot 介面與終端機日誌
2. （00:03）點擊模型選擇器並切換至 Claude Opus 5
3. （00:40）在模型選單中調整強度設定

**逐字稿**

- `00:00` Anthropic 最新的 Opus 模型 Claude Opus 5 現已在 GitHub Copilot 中推出。（Claude Opus 5, Anthropic's newest Opus model, is now available in GitHub Copilot.）
- `00:04` 它專為需要縝密推理、（It's designed for complex, long-running coding tasks that require careful reasoning,）
- `00:09` 有效工具使用以及在多個步驟中可靠執行的複雜、長時間執行的程式開發任務而設計。（effective tool use, and reliable execution across multiple steps.）
- `00:13` 在我們早期的測試中，Opus 5 在 Agentic 程式開發工作流程中展現了強大的效能，（In our early testing, Opus 5 showed strong performance on agentic coding workflows,）
- `00:18` 包含自主程式碼變更、迴歸驗證，以及需要協調多個工具的任務。（including autonomous code changes, regression verification, and tasks that require coordinating）
- `00:23` 該模型在進行針對性變更、驗證工作，（multiple tools. The model was especially effective at making targeted changes, validating work,）
- `00:29` 以及減少複雜任務上不必要的執行負荷時特別有效。（and reducing unnecessary execution overhead on complex tasks.）
- `00:33` Claude Opus 5 將開放給 Copilot Pro+、Max、Business 和 Enterprise 使用者使用。（Claude Opus 5 will be available to Copilot Pro+, Max, Business, and Enterprise users.）
- `00:39` 你將能夠在各個 GitHub Copilot 服務的模型選擇器中選取該模型，（You'll be able to select the model in the model picker across GitHub Copilot services,）
- `00:44` 包含 Visual Studio Code、Visual Studio、GitHub Copilot CLI、GitHub Copilot Cloud Agent、（including Visual Studio Code, Visual Studio, the GitHub Copilot CLI, GitHub Copilot Cloud Agent,）
- `00:50` GitHub.com 以及 GitHub Copilot 應用程式。（GitHub.com, and the GitHub Copilot app.）

**Opus 5 在多項編程與知識工作評測上達到前沿 SOTA，並在 ARC-AGI-3 創新問題解決測試中以 30.2% 分數遠超其他模型**

**數據表**

|   | Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol |
| --- | --- | --- | --- | --- |
| Agentic terminal coding (Frontier-Bench v0.1) | 43.3% | 33.7% | 21.1% | 34.4% |
| Knowledge work (GDPval-AA v2) | 1861 | 1747 | 1593 | 1736 |
| Novel problem-solving (ARC-AGI-3) | 30.2% | — | 1.5% | 7.8% |
| Agentic search (BrowseComp) | 90.8% | 87.4% | 84.3% | 90.4% |
| Multidisciplinary reasoning (Humanity's Last Exam, no tools) | 56.3% | 56.5% | 49.8% | — |
| Multidisciplinary reasoning (Humanity's Last Exam, with tools) | 64.7% | 63.9% | 57.9% | — |
| Computer use (OSWorld 2.0) | 70.6% | 66.1% | 55.7% | 62.6% |
| Agentic coding (DeepSWE v1.1) | 68.8% | 69.7% | 59.0% | 72.7% |
| Agentic coding (FrontierCode v1.1, Main) | 53.4% | 53.5% | 46.5% | 47.5% |
| Business workflows (AutomationBench) | 26.0% | 17.4% | 17.0% | 18.1% |
| Legal (Legal Agent Benchmark, Held-out) | 11.7% | 13.3% | 10.4% | 2.5% |
| Health (HealthBench Professional) | 59.8% | 66.0% (Mythos 5) | 57.4% | 60.5% |
| Biology (BioMysteryBench, hard) | 49.4% | 46.5% | 42.4% | — |
| Biology (BioMysteryBench, human solved) | 90.1% | 89.0% (Mythos 5) | 88.5% | — |

**Opus 5 (high) 在 ARC-AGI-3 評測中取得約 30% 得分，達到次佳模型 GPT-5.6 Sol 的三倍以上。**

**數據表**

| 項目 | X | Y |
| --- | --- | --- |
| Opus 5 (high) | $20,500 | 30.2% |
| GPT-5.6 Sol | $24,000 | 7.8% |
| Opus 4.8 (high) | $13,000 | 1.5% |

**Opus 5 在 Humanity's Last Exam (with tools) 基準測試中，以類似或更低的每任務成本展現優於 Fable 5 與 Opus 4.8 的解題正確率**

**數據表**

| 模型 | low effort (成本, 正確率) | medium effort (成本, 正確率) | high effort (成本, 正確率) | xhigh effort (成本, 正確率) | max effort (成本, 正確率) |
| --- | --- | --- | --- | --- | --- |
| Opus 5 | ($0.24, 56.1%) | ($0.69, 61.3%) | ($1.30, 63.2%) | ($1.99, 64.8%) | ($2.70, 64.7%) |
| Fable 5 | ($0.90, 58.2%) | ($1.40, 61.4%) | ($1.90, 61.8%) | ($2.60, 63.1%) | ($4.30, 63.9%) |
| Opus 4.8 | ($0.33, 50.2%) | ($0.52, 55.2%) | ($0.62, 55.7%) | ($1.12, 57.6%) | ($1.54, 58.0%) |

## 策展筆記

數據校正：Anthropic 官方 System Card 在 OSS-Fuzz 這段自己打架——內文寫 Opus 4.8「scored on 38.5% of targets」，同頁圖表卻標 61.5%，兩者剛好互補。對照 Claude Sonnet 5 System Card 的原始寫法可確認圖表為準：那份寫 Opus 4.8「failed to score on 38.5%」、Sonnet 5 失敗 45.5%、Mythos 5 失敗 20.0%，換算成得分率正是 61.5%、54.5%、80.0%，與圖表三根長條完全吻合。也就是 Opus 5 System Card 內文把「未得分比例」誤寫成了「得分比例」。

## 標籤

新產品, 功能更新, LLM, Anthropic, Claude, Claude Code, Cursor, Copilot, Cognition, Devin, Lovable
