# DeepSeek 推出 DeepSeek-V4-Pro；Arena.ai 以每百萬 token 輸入 0.435 美元、輸出 0.87 美元評估其 Agent 開發性價比

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：DeepSeek (@deepseek_ai) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥🔥🔥🔥🔥 · 日期：2026-08-13

> 原始來源：https://x.com/deepseek_ai/status/2087864585504305397

## 證據與延伸閱讀

- [DeepSeek 推出 DeepSeek-V4-Pro；Arena.ai 以每百萬 token 輸入 0.435 美元、輸出 0.87 美元評估其 Agent 開發性價比。](https://x.com/deepseek_ai/status/2087864585504305397)
- [API定價區分尖離峰時段](https://x.com/deepseek_ai/status/2087864589895798968)
- [Vals AI 指出 V4 Pro 提升 11 分](https://x.com/ValsAI/status/2087697657301279220)
- [Arena.ai Code/Text Arena 評測結果](https://x.com/arena/status/2087767198974533648)
- [Arena.ai 價格擊敗 Opus 與 GLM](https://x.com/arena/status/2087784211642192332)
- [Cline 指出 V4-Pro Terminal Bench 表現](https://x.com/cline/status/2087602193205694891)
- [ClinePass 訂閱服務與指令](https://cline.bot/cline-pass)

## 中文摘要

DeepSeek 推出 DeepSeek-V4-Pro；Arena.ai 以每百萬 token 輸入 0.435 美元、輸出 0.87 美元評估其 Agent 開發性價比。

**產品更新** DeepSeek 表示，DeepSeek-V4-Pro 帶來多項面向生產環境的 Agent 升級，並讓使用者依任務複雜度調整 reasoning effort：

- `low` 適合簡單任務。
- `high` 適合日常 Agent 工作流程。
- `max` 適合複雜任務。

V4-Pro 與 V4-Flash 都支援這項設定。DeepSeek-V4-Pro 已在 app 與 web 版提供，使用者可透過「Expert Mode」啟用；API 也已開放，模型名稱維持不變，設定方式則需參考 API 文件。產品同時原生支援 OpenAI Responses API，並針對 Codex 提供一鍵設定。

**API 定價** 隨著 V4 系列推出，DeepSeek 將 API 定價改為區分尖峰與離峰時段。離峰價格比尖峰低 50%，官方認為這能讓使用者更彈性地安排工作負載。新價格將於 2026 年 8 月 16 日 16:00 UTC 起生效；原始公告未列出各模型的新單價，因此實際費率仍須以 API 文件為準。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/ab88a04cd8aee7ec.jpg)
> DeepSeek-V4 系列 API 推出全新計費標準，引入尖峰與離峰差異化費率，其中離峰時段價格比尖峰時段降低 50%。

**第三方評測** Vals AI 表示，DeepSeek V4 Pro 0813 在 Vals Index 上提升 11 分，成為排名第 2 的開放權重模型；每項任務成本約為 0.14 美元，價格約是 Kimi K3 的 1/17，而 Kimi K3 是唯一排名高於它的開放權重模型。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/e9432915be664f8e.jpg)
> 來源：[@ValsAI](https://x.com/ValsAI/status/2087697657301279220)（回覆）｜DeepSeek V4 Pro 0813 在 Vals Index 以 66.25% 的準確率位居開放權重模型第二，每項測試成本僅 0.14 美元，成本約為唯一領先它的 Kimi K3 的 1/17。

Arena.ai 的 Code Arena: WebDev 初步結果則顯示：

- DeepSeek-V4-Pro（Max）得分 1607，整體約第 8 名，開放模型中排名第 2。
- 它落後 GPT-5.6 Sol（xHigh）的 1622 分，並低於 Kimi K3（Max）的 1674 分。
- 在 Text Arena 中，它以 1465 分約排名開放模型第 5，與 GLM-5.1 的 1467 分、GPT-5.6 Terra（xHigh）與 Grok 4.6（High）的 1464 分相當。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/d9ae11c73ef176b2.jpg)
> 來源：[@arena](https://x.com/arena/status/2087767202971918784)（回覆）｜DeepSeek-V4-Pro (Max) 在 Text Arena 開放模型排行榜中獲得 1,465 分（AutoEval）位居第 5，與 GLM-5.1（1,467 分）等模型表現持平。

Arena.ai 特別提醒，以上是早期 AutoEval 分數，由使用 Arena 人類偏好資料訓練的 Reward Model 自動投票，並非即時真人票選；隨著更多真人票數加入，排名仍可能收斂或變動。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/9236968ad3eb4038.jpg)
> 來源：[@arena](https://x.com/arena/status/2087767198974533648)（回覆）｜DeepSeek-V4-Pro (Max) 在 Code Arena: WebDev 中以 1607 分（AutoEval）排名整體約第 8 名（開放模型第 2 名），僅次於 GPT-5.6 Sol (xHigh) 的 1622 分。

**價格效益** Arena.ai 指出，DeepSeek-V4-Pro（Max）目前在 Code Arena: WebDev 以每百萬 token 輸入 0.435 美元、輸出 0.87 美元的價格，擊敗部分更高價模型，包括輸入／輸出價格為 5／25 美元的 Opus 4.8，以及 1.4／4.4 美元的 GLM-5.2。Arena.ai 認為，這款即將推出的開放權重模型可能改變 WebDev 評測中的效能與價格邊界。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/e2cddd5d8f530247.jpg)
> 來源：[@arena](https://x.com/arena/status/2087784211642192332)（回覆）｜DeepSeek-V4-Pro (Max) 與其他模型在 Code Arena 的價格與 Arena Score 效率前緣比較。

**Agent 程式開發** Cline 表示，DeepSeek 悄然釋出 V4-Pro 0813；相較 4 月的 Preview 模型，它在 Terminal Bench 提升 15.8%，並以約為 Fable 5 的 1/57 成本達到相近表現。Cline 列出的模型規模為 1.6T 個參數、49B 個 active parameters，以及 1M context。Cline 進一步稱它是目前市場上價格效能最佳的模型，但這仍屬 Cline 的評價，不能視為涵蓋所有任務與評測的普遍結論。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/86aaba4d6bcbc8e2.jpg)
> DeepSeek-V4-Pro-0813 與 GLM-5.2、Kimi-K3、Opus-4.8 及 Fable 5 等模型在多項基準測試的表現各有高低，其中 DeepSeek-V4-Pro-0813 在 Cybergym 與 AutomationBench (Public) 取得最高分。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/b4c759b494e4fb0e.png)
> 來源：[@cline](https://x.com/cline/status/2087602193205694891)（回覆）｜V4-Pro 0813 在 Terminal-Bench 2.1 以 87.9 分居次，表現逼近 Fable 5（88.0 分），且價格僅為每百萬 token 輸入/輸出 $0.435 / $0.87。

**ClinePass 方案** DeepSeek-V4-Pro 現已可透過 ClinePass 使用。Cline 將 ClinePass 定位為以約五分之一價格提供開放權重模型的訂閱服務，並稱在目前價格下幾乎可不限量使用；首月促銷價為 4.99 美元，之後每月 9.99 美元。Cline 提供的安裝指令如下，執行前仍應由使用者自行核對套件來源與權限：

```bash
npm i -g cline
```

方案頁面：[ClinePass](https://cline.bot/cline-pass)。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/ed45b99e9b87e611.png)
> 深色網格背景與紫色漸層光暈上的白色文字與圖示，中央偏左為帶有兩道垂直長條的機器人頭像造型 logo，右側為白色粗體文字 cline.bot。

## 媒體內容

**DeepSeek-V4-Pro-0813 與 GLM-5.2、Kimi-K3、Opus-4.8 及 Fable 5 等模型在多項基準測試的表現各有高低，其中 DeepSeek-V4-Pro-0813 在 Cybergym 與 AutomationBench (Public) 取得最高分。**

**數據表**

| 項目 | 數值 |
| --- | --- |
| HLE (wo/w tools) | DeepSeek-V4-Pro-0813 42.7/60.0 · DeepSeek-V4-Flash-0731 37.8/51.5 · DeepSeek-V4-Pro-Preview 37.7/48.2 · DeepSeek-V4-Flash-Preview 34.8/45.1 · GLM-5.2 40.5/54.7 · Kimi-K3 43.5/56.0 · Opus-4.8 49.8/57.9 · Fable 5 (w/ fallback) 53.3/63.0 |
| Terminal Bench 2.1 | DeepSeek-V4-Pro-0813 87.9 · DeepSeek-V4-Flash-0731 82.7 · DeepSeek-V4-Pro-Preview 72.1 · DeepSeek-V4-Flash-Preview 61.8 · GLM-5.2 81.0 · Kimi-K3 88.3 · Opus-4.8 85.0 · Fable 5 (w/ fallback) 88.0 |
| NL2Repo | DeepSeek-V4-Pro-0813 61.5 · DeepSeek-V4-Flash-0731 54.2 · DeepSeek-V4-Pro-Preview 38.5 · DeepSeek-V4-Flash-Preview 39.4 · GLM-5.2 48.9 · Kimi-K3 - · Opus-4.8 69.7 · Fable 5 (w/ fallback) - |
| Cybergym | DeepSeek-V4-Pro-0813 83.3 · DeepSeek-V4-Flash-0731 76.7 · DeepSeek-V4-Pro-Preview 52.7 · DeepSeek-V4-Flash-Preview 38.7 · GLM-5.2 - · Kimi-K3 80.0 · Opus-4.8 78.3 · Fable 5 (w/ fallback) 83.1 |
| DeepSWE | DeepSeek-V4-Pro-0813 62.7 · DeepSeek-V4-Flash-0731 54.4 · DeepSeek-V4-Pro-Preview 12.8 · DeepSeek-V4-Flash-Preview 7.3 · GLM-5.2 46.2 · Kimi-K3 67.5 · Opus-4.8 58.0 · Fable 5 (w/ fallback) 70.0 |
| Toolathlon-Verified | DeepSeek-V4-Pro-0813 74.1 · DeepSeek-V4-Flash-0731 70.3 · DeepSeek-V4-Pro-Preview 55.9 · DeepSeek-V4-Flash-Preview 49.7 · GLM-5.2 59.9 · Kimi-K3 76.5 · Opus-4.8 76.2 · Fable 5 (w/ fallback) 77.9 |
| Agents' Last Exam | DeepSeek-V4-Pro-0813 25.7 · DeepSeek-V4-Flash-0731 25.2 · DeepSeek-V4-Pro-Preview 16.5 · DeepSeek-V4-Flash-Preview 15.8 · GLM-5.2 23.8 · Kimi-K3 27.6 · Opus-4.8 25.7 · Fable 5 (w/ fallback) - |
| AutomationBench (Public) | DeepSeek-V4-Pro-0813 31.8 · DeepSeek-V4-Flash-0731 25.1 · DeepSeek-V4-Pro-Preview 12.8 · DeepSeek-V4-Flash-Preview 10.8 · GLM-5.2 12.9 · Kimi-K3 30.8 · Opus-4.8 27.2 · Fable 5 (w/ fallback) 29.1 |
| DSBench-FullStack | DeepSeek-V4-Pro-0813 71.1 · DeepSeek-V4-Flash-0731 68.7 · DeepSeek-V4-Pro-Preview 41.8 · DeepSeek-V4-Flash-Preview 37.0 · GLM-5.2 61.8 · Kimi-K3 73.7 · Opus-4.8 71.6 · Fable 5 (w/ fallback) 77.2 |
| DSBench-Hard | DeepSeek-V4-Pro-0813 67.2 · DeepSeek-V4-Flash-0731 59.6 · DeepSeek-V4-Pro-Preview 31.1 · DeepSeek-V4-Flash-Preview 25.8 · GLM-5.2 54.5 · Kimi-K3 63.0 · Opus-4.8 71.7 · Fable 5 (w/ fallback) 68.3 |
| 註腳：* For public Code Agent tasks, V4-Pro-0813 was tested using our upcoming DeepSeek Harness (minimal mode) framework. Settings: max tier, topp=0.95, temperature=1.0. Results may vary slightly with other frameworks. |  |

**DeepSeek-V4 系列 API 推出全新計費標準，引入尖峰與離峰差異化費率，其中離峰時段價格比尖峰時段降低 50%。**

**數據表**

|   | Input (cache hit) | Input (cache miss) | Output |
| --- | --- | --- | --- |
| DeepSeek-V4-Flash (Off-Peak) | $ 0.007 | $ 0.22 | $ 0.66 |
| DeepSeek-V4-Flash (Peak hours) | $ 0.014 | $ 0.44 | $ 1.32 |
| DeepSeek-V4-Pro (Off-Peak) | $ 0.022 | $ 0.66 | $ 1.98 |
| DeepSeek-V4-Pro (Peak hours) | $ 0.044 | $ 1.32 | $ 3.96 (Peak Hours: 01:00–04:00 and 06:00–10:00 UTC, Effective from: 16:00, August 16, 2026 UTC) |

**DeepSeek V4 Pro 0813 在 Vals Index 以 66.25% 的準確率位居開放權重模型第二，每項測試成本僅 0.14 美元，成本約為唯一領先它的 Kimi K3 的 1/17。**

**數據表**

|   | ACCURACY | COST/TEST | LATENCY | ACCURACY | COST/TEST | LATENCY | ACCURACY | COST/TEST | LATENCY | ACCURACY | COST/TEST | LATENCY | ACCURACY | COST/TEST | LATENCY | ACCURACY | COST/TEST | LATENCY | ACCURACY | COST/TEST | LATENCY | ACCURACY | COST/TEST | LATENCY | ACCURACY | COST/TEST | LATENCY | ACCURACY | COST/TEST | LATENCY |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 Kimi K3 | 51.57% | $0.04 | 585s | 51.57% | $0.04 | 585s | 51.57% | $0.04 | 585s | 51.57% | $0.04 | 585s | 51.57% | $0.04 | 585s | 51.57% | $0.04 | 585s | 51.57% | $0.04 | 585s | 51.57% | $0.04 | 585s | 51.57% | $0.04 | 585s | 51.57% | $0.04 | 585s |

**DeepSeek-V4-Pro (Max) 在 Code Arena: WebDev 中以 1607 分（AutoEval）排名整體約第 8 名（開放模型第 2 名），僅次於 GPT-5.6 Sol (xHigh) 的 1622 分。**

**數據表**

| 項目 | 數值 |
| --- | --- |
| 1. Claude Opus 5 (Max) | 1 |
| 691、2. Kimi K3 (Max) | 1 |
| 674、3. Qwen-3.8 Max | 1 |
| 669、4. Claude Opus 5 (High) | 1 |
| 664、5. Grok-4.6 (High) | 1 |
| 630、6. Claude Fable 5 | 1 |
| 627、7. GPT-5.6 Sol (xHigh) | 1 |
| 622、DeepSeek-V4-Pro (Max) | 1 |
| 607 (AutoEval)、8. GLM-5.2 (Max) | 1 |
| 587、9. DeepSeek V4 Flash (High) | 1 |
| 582、10. Claude Opus 4.8 (High) | 1 |
| 564、11. Claude Opus 4.7 | 1 |
| 558、12. Claude Opus 4.7 (High) | 1 |
| 557、13. Grok-4.5 | 1 |
| 554、14. Claude Opus 4.6 (High) | 1 |
| 545、15. Claude Sonnet 5 (High) | 1,541 |

**DeepSeek-V4-Pro (Max) 在 Text Arena 開放模型排行榜中獲得 1,465 分（AutoEval）位居第 5，與 GLM-5.1（1,467 分）等模型表現持平。**

**數據表**

| 項目 | 數值 |
| --- | --- |
| Kimi-K3 Max | 1,489 |
| GLM-5.2 (Max) | 1,471 |
| MiMo V2.5 Pro | 1,468 |
| GLM-5.1 | 1,467 |
| DeepSeek-V4-Pro (Max) | 1,465 |
| Kimi-K2.6 | 1,461 |
| DeepSeek-V4 Pro | 1,458 |
| GLM-5 | 1,457 |
| HY3 | 1,457 |
| DeepSeek-V4 Pro (High) | 1,456 |
| Gemma 4 31B | 1,451 |
| Kimi-K2.5 (Thinking) | 1,450 |
| MiniMax-M3 | 1,444 |
| Inkling | 1,442 |
| Qwen-3.5 397B A17B | 1,442 |
| GLM-4.7 | 1,442 |

**V4-Pro 0813 在 Terminal-Bench 2.1 以 87.9 分居次，表現逼近 Fable 5（88.0 分），且價格僅為每百萬 token 輸入/輸出 $0.435 / $0.87。**

**數據表**

|   | 得分 | 價格 (in / out) |
| --- | --- | --- |
| Fable 5 | 88.0 | $10 / $50 |
| V4-Pro 0813 | 87.9 | $0.435 / $0.87 |
| Opus 4.8 | 85.0 | $5 / $25 |
| V4-Flash | 82.7 | $0.14 / $0.28 |
| V4-Pro Preview | 72.1 | $0.435 / $0.87 |

## 標籤

新產品, 功能更新, Agent, DeepSeek
