# DeepSeek 推出 DeepSeek-V4-Flash-Vision-Exp，讓多模態 Agent 結合視覺理解與工具使用

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：DeepSeek (@deepseek_ai) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥🔥🔥🔥🔥 · 日期：2026-08-21

> 原始來源：https://x.com/deepseek_ai/status/2090730032574631962

## 證據與延伸閱讀

- [DeepSeek 推出 DeepSeek-V4-Flash-Vision-Exp，讓多模態 Agent 結合視覺理解與工具使用。](https://x.com/deepseek_ai/status/2090730032574631962)
- [多模態 Agent benchmark 表現接近 Opus 4.8](https://x.com/OpenRouter/status/2090781624015393010)
- [模型識別名稱與 ZenMux 提供免費模型](https://x.com/ZenMuxAI/status/2090738450748342658)
- [API 影像輸入支援與檔案限制](https://api-docs.deepseek.com/guides/vision)
- [Files API 免費使用與儲存限制](https://api-docs.deepseek.com/guides/files_api)
- [Harness 與開發流程指令](https://x.com/opencode/status/2090796489094111622)
- [透過 Vercel AI Gateway 使用與設定](https://vercel.com/changelog/deepseek-v4-flash-with-vision-now-available-on-ai-gateway) — 官方文件
- [deepseek.com](https://api-docs.deepseek.com/guides/vision/)
- [deepseek.com](https://api-docs.deepseek.com/guides/files_api/)
- [Vercel與ZenMux支援與宣傳](https://x.com/vercel_dev/status/2090847674803364232)
- [DeepSeek Harness核心概念與安裝](https://video.twimg.com/tweet_video/HQPDJwZaIAAW-ZS.mp4)

## 中文摘要

DeepSeek 推出 DeepSeek-V4-Flash-Vision-Exp，讓多模態 Agent 結合視覺理解與工具使用。

**發布重點** DeepSeek 在 2026 年 8 月 21 日宣布，DeepSeek-V4-Flash-Vision-Exp 已可透過 DeepSeek API Platform 使用。官方表示，這個實驗性模型在文字任務上具備與 DeepSeek-V4-Flash 相當的 agents、推理與世界知識能力；在多模態 agent benchmark 上則較 DeepSeek-V4-Flash 有明顯躍升，表現接近 Opus 4.8。DeepSeek Harness 0.1.1 也在同日發布，加入對新模型的開箱即用支援。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/c8e58fd3ca1cddfe.jpg)
> DeepSeek-V4-Flash-Vision-Exp 在文字 Agent 任務媲美 DeepSeek-V4-Flash，並於多模態 Agent benchmark 大幅超越 V4-Flash 且表現接近 Opus-4.8。

**Agent 工作流程** 多模態輸入讓 Agent 能把視覺理解與工具使用結合，處理更多實際工作。DeepSeek 表示，DeepSeek-V4-Flash-Vision-Exp 可在各種 Agent 框架中運作，適合讀取畫面、分析圖表，以及由影像觸發後續工具流程。OpenRouter 的公告進一步宣稱，該模型在文字 agent benchmark 上匹配 DeepSeek-V4-Flash 0731，並在 Agents' Last Exam 與 ZeroBench 擊敗 Opus 4.8；這些數據屬於發布方與平台的說法，摘要未將其視為獨立第三方驗證結果。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/8b98b6afe5657da7.jpg)
> 來源：[@ZenMuxAI](https://x.com/ZenMuxAI/status/2090738450748342658)（回覆）｜以藍色漸層背景為基底的簡報畫面，上方中央偏左印有帶章魚圖示的 ZenMux 品牌 logo，右側為帶鯨魚圖示的 deepseek logo，中間以垂直分隔線隔開；下方中央置有帶鯨魚圖示的 deepseek logo，右側以白色大字顯示「DeepSeek V4 Flash Vision Exp」產品版本名稱。

**模型規格與供應管道** API 使用的模型識別字串是 `deepseek-v4-flash-vision-exp`。DeepSeek 表示，圖片會轉換為 token 計費，每張最多 384 tokens，價格依照 V4-Flash。模型支援 Chat Completions、Messages 與 Responses API，也接受文字與圖片混合輸入；圖片可透過 base64、外部 URL 或 Files API 提供。

目前除了 DeepSeek API Platform，也可在 [OpenRouter](https://openrouter.ai/deepseek/deepseek-v4-flash-vision-exp)、Vercel AI Gateway、OpenCode Go 與 [ZenMux](https://zenmux.ai/deepseek/deepseek-v4-flash-vision-exp-free) 使用。Vercel 的公告稱此版本具備 1M context window，工具使用、推理與快取行為維持原有方式；Vercel AI Gateway 以供應商定價提供服務，不加收推論平台費用，包括 BYOK 請求。Vercel 同時提醒，`-exp` 代表實驗性版本，行為可能變動，若用於 production path，應設定 fallback model。ZenMux 則宣布提供一週免費使用；其頁面宣傳的架構為 284B MoE、13B active，但這是平台貼文中的資訊，並非 DeepSeek API 文件列出的規格。

<video src="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/1787333485264-a0jg7ijm.mp4" poster="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/ed6f0eec2f1e4742.jpg" autoplay loop muted playsinline preload="metadata" style="max-width:100%;height:auto;display:block;margin:1rem 0"></video>
> DeepSeek Harness 專案介紹首頁與核心功能展示

**DeepSeek Harness** 動圖展示的 DeepSeek Harness 首頁以「Everything is a plugin」為核心概念，並標示專案為開放原始碼、developer preview。畫面提供以下取得與啟動方式：

```bash
git clone https://github.com/deepseek-ai/deepseek-harness
npx deepseek-ai/st-ui web
```

畫面也列出 `View on GitHub`、`Developer docs` 與 `Community plugins` 等入口。其設計說明將模型視為 Agent 的靈魂，harness 則負責讓 Agent 理解環境、使用工具，並在真實世界設定中持續工作。DeepSeek Harness 將模型、工具、技能、工作階段、沙盒、儲存、規則、排程與 UI 等能力做成可替換或可重組的外掛，並強調每次執行都可追蹤，也支援多種 runtime mode。

展示中的 Plugins 設定包含 `anthropic`、`cursor`、`deepseek`、`editor-base`、`filesystem`、`git`、`lark-base`、`linear`、`openai`、`sandbox` 與 `vector-db-drag`，皆顯示為 enabled。這些畫面屬於功能展示，不能直接視為 DeepSeek API 的正式支援清單。

**圖片輸入方式** 官方文件說明，模型可處理 JPEG、PNG、GIF 與 WebP，格式會依實際檔案內容判定，不會只依賴檔名或宣告的 MIME type。圖片只能放在 user message；放入 system 或 assistant message 會回傳 400 錯誤，其他不支援視覺能力的模型也會回傳 400 錯誤。主要方式如下：

- 直接以 base64 data URL 嵌入圖片，適合本機檔案，但編碼後內容會計入 48 MiB request body 上限。
- 傳入公開可存取的 `http(s)` 圖片 URL；URL 最長 8192 字元，圖片最多 32 MiB，下載必須在 60 秒內完成。
- 先透過 Files API 上傳，再以回傳的 `file_id` 引用；這適合重複使用同一張圖片，或處理超過 inline 限制的檔案。

例如，OpenAI-compatible Chat Completions 可使用下列模型欄位與圖片區塊：

```python
response = client.chat.completions.create(
    model="deepseek-v4-flash-vision-exp",
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "What is in this image?"},
                {
                    "type": "image_url",
                    "image_url": {"url": "https://example.com/image.jpg"},
                },
            ],
        }
    ],
)
print(response.choices[0].message.content)
```

若圖片是透過 Files API 上傳，請在內容中引用回傳的 `file_id`：

```python
{
    "role": "user",
    "content": [
        {"type": "text", "text": "What is in this image?"},
        {"type": "file", "file_id": "file-api-xxxxxxxxxxxxxxxx"},
    ],
}
```

**Files API 與限制** Files API 可免費使用；同一張圖片上傳一次後，後續請求可重複引用，不必重新上傳，也能降低 request bandwidth。單一上傳檔案最多 64 MiB，儲存空間上限為每位使用者 25 GiB、最多 10,000 個檔案。檔案可設定保存時間為 1 小時至 30 天，或省略 `expires_after` 欄位以永久保存，因此部署時仍需人工確認資料保留政策與檔案生命週期。

圖片經調整大小後，每張最多計入 384 tokens；多張圖片仍逐張計算。每次請求最多 600 張圖片，未使用 `file_id` 時總圖片大小最多 64 MiB，包含 `file_id` 圖片時最多 200 MiB；單邊解析度上限為 8192 像素，若請求包含 15 張以上圖片，單邊上限降為 4096 像素。`detail` 可設定為 `low`、`high`、`original` 或 `auto`；`low` 會先縮小至 512×512，適合不需要細節的任務，其他設定目前等同保留原始影像。

**相容介面與實際影響** 除了 OpenAI-compatible 介面，DeepSeek 也提供 Anthropic-compatible `/messages` 與 Responses API。Anthropic 介面以 `source.type` 區分 `base64`、`url` 和 `file`，Files API 引用則需要標頭 `anthropic-beta: files-api-2025-04-14`；Responses API 使用 `input_image` content part，並沿用相同的圖片大小與數量限制。整體來看，DeepSeek 這次不是只增加圖片辨識，而是把視覺輸入、工具呼叫、Agent framework 與可重複引用的檔案機制接在同一條 API 工作流程中；但由於模型仍是 experimental preview，正式環境採用前應保留 fallback、重新驗證 benchmark，並確認圖片上傳與保存設定。

## 媒體內容

**DeepSeek-V4-Flash-Vision-Exp 在文字 Agent 任務媲美 DeepSeek-V4-Flash，並於多模態 Agent benchmark 大幅超越 V4-Flash 且表現接近 Opus-4.8。**

**數據表（1）Text-Based Agent Evaluation**

|   | DeepSeek V4-Flash-Vision-Exp | DeepSeek V4-Flash- 0731 | Opus-4.8 |
| --- | --- | --- | --- |
| Terminal Bench 2.1 | 83.9 | 82.7 | 85.0 |
| NL2Repo | 57.7 | 54.2 | 69.7 |
| Cybergym | 75.3 | 76.7 | 78.3 |
| DeepSWE | 59.3 | 54.4 | 58.0 |
| Toolathlon-Verified | 75.9 | 70.3 | 76.2 |
| DSBench-Hard | 63.6 | 59.6 | 71.7 |
| AutomationBench (Public) | 25.7 | 25.1 | 27.2 |

**數據表（2）Multimodal Agent Evaluation**

|   | DeepSeek V4-Flash-Vision-Exp | DeepSeek V4-Flash- 0731 | Opus-4.8 |
| --- | --- | --- | --- |
| ApexBench (Pass@1) | 36.5 | 26.2** | 39.4 |
| Agents' Last Exam | 27.3 | 25.2** | 25.7 |
| Chartography | 64.3 | - | 65.0 |
| ZeroBench (Pass@5) | 35.0 | - | 34.0 |

## 標籤

新產品, 功能更新, VLM, Agent, DeepSeek
