# Arena.ai：Code Arena 各領域專有與開源模型差距縮至約 10 至 40 分

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：Arena.ai (@arena) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥🔥 · 日期：2026-08-12

> 原始來源：https://x.com/arena/status/2086844868454662210

## 證據與延伸閱讀

- [Arena.ai：Code Arena 各領域專有與開源模型差距縮至約 10 至 40 分](https://x.com/arena/status/2086844868454662210) — 一手來源
- [Frontend 賽道目前差距約 40 分](https://x.com/arena/status/2052455466220089630)
- [Expert 賽道目前差距約 40 分](https://x.com/arena/status/2052455468736672046)

## 中文摘要

Arena.ai：Code Arena 各領域專有與開源模型差距縮至約 10 至 40 分。

**WebDev 領域** 截至 2026 年 8 月 10 日，Code Arena: WebDev 的前沿差距已從 2025 年底約 150 分，壓縮至目前約 10 分。Muse Spark 1.2 預計釋出模型權重，Arena.ai 將持續觀察它與 Glimmer 等其他開源模型對前沿競爭格局的影響，官方分數則將陸續公布。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/872db063dc87f9dd.png)
> 在 Code Arena: WebDev 中，專有模型與開源模型的前沿差距已從 2025 年底的約 150 分大幅縮小至目前的約 10 分。

**Frontend 領域** 2026 年 5 月 8 日的資料顯示，Code Arena: Frontend 的競爭更為接近。由於該領域歷史較短，差距變化也更快：專有模型的領先幅度一度上升 100 分，之後在 2026 年春季大幅收斂，目前約為 40 分。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/e42adfa2896d4a48.jpg)
> Proprietary 模型在 Code Arena: Frontend 的領先差距已收窄至約 40 分。

**Expert 領域** 專家級 prompt 仍是開源模型最難突破的挑戰。DeepSeek R1 曾在 2025 年初短暫超越專有模型，使差距從「幾乎追平」轉為開源模型領先，但這項優勢未能維持，專有模型很快重新取得第一名；目前差距約 40 分，約等於排名第 1 與第 8 的距離。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/30a40cfe918129d2.jpg)
> 專有模型在 Expert prompts 評測中相較於開源模型持續保持優勢，目前兩者差距約為 40 分。

- 開源模型曾在高難度 prompt 登上第一名，但只維持很短時間。
- 專有模型在 Expert 領域持續守住第一名的能力較穩定。
- Expert 是差距曾經縮小、反轉，最後又重新拉開的 Arena。

## 媒體內容

**在 Code Arena: WebDev 中，專有模型與開源模型的前沿差距已從 2025 年底的約 150 分大幅縮小至目前的約 10 分。**

**數據表（1）Arena Score**

|   | 起始 | 最佳 | 結束 | 起始 | 最佳 | 結束 |
| --- | --- | --- | --- | --- | --- | --- |
| Proprietary | 1408 | 1675 | 1675 | 1408 | 1675 | 1675 |

**數據表（2）Gap**

|   | 起始 | 最大 | 最小 | 結束 |
| --- | --- | --- | --- | --- |
| Gap | 47 | 158 | 8 | 10 |

**Proprietary 模型在 Code Arena: Frontend 的領先差距已收窄至約 40 分。**

**數據表（1）Arena Score**

|   | 起始 | 最佳 | 結束 | 起始 | 最佳 | 結束 |
| --- | --- | --- | --- | --- | --- | --- |
| Proprietary | 1408 | 1538 | 1531 | 1408 | 1538 | 1531 |

**數據表（2）Gap**

|   | 起始 | 最佳 | 結束 |
| --- | --- | --- | --- |
| Gap | 46 | 157 | 38 |

**專有模型在 Expert prompts 評測中相較於開源模型持續保持優勢，目前兩者差距約為 40 分。**

**數據表（1）Arena Score**

|   | 起始 | 最佳 | 結束 | 起始 | 最佳 | 結束 |
| --- | --- | --- | --- | --- | --- | --- |
| Proprietary | 1230 | 1510 | 1508 | 1230 | 1510 | 1508 |

**數據表（2）Gap**

|   | 起始 | 最佳 | 結束 |
| --- | --- | --- | --- |
| Gap | 116 | 0 | 38 |

## 標籤

Benchmark, LLM, 開源專案
