# Artificial Analysis 推出 Speech Agent Arena，以真人任務同時評估語音對話偏好與工具呼叫成功率

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：Artificial Analysis (@ArtificialAnlys) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥🔥🔥 · 日期：2026-08-21

> 原始來源：https://x.com/ArtificialAnlys/status/2090806900631994528

## 證據與延伸閱讀

- [Artificial Analysis 推出 Speech Agent Arena，以真人任務同時評估語音對話偏好與工具呼叫成功率。](https://x.com/ArtificialAnlys/status/2090806900631994528) — 官方文件
- [偏好度隨 TTFA 縮短而上升](https://x.com/ArtificialAnlys/status/2090806904079647195) — 官方文件
- [模型價格與表現差異明顯](https://x.com/ArtificialAnlys/status/2090806907179323426) — 官方文件
- [Speech to Speech Index 採四個面向](https://artificialanalysis.ai/methodology/speech-to-speech-benchmarking) — 官方文件

## 中文摘要

Artificial Analysis 推出 Speech Agent Arena，以真人任務同時評估語音對話偏好與工具呼叫成功率。

**核心評測** Artificial Analysis 表示，現有 Speech to Speech benchmark 多半測試推理、模擬 agent 任務，以及輪流發言、處理打斷等對話互動；Speech Agent Arena 則進一步觀察模型在真人執行任務時的實際表現。參與者會在相同情境下，分別與兩個隱藏模型進行即時語音對話，再選出偏好的對話。評測涵蓋 35 個情境：

- 15 個 agentic 情境，需要透過工具採取行動，例如訂購外送或預約牙科。
- 20 個 non-agentic 情境，不需要工具呼叫，例如詢問營業時間、課程費用與可提供的用品。

偏好結果用 Preference Elo 表示；agentic 情境另計算 Task Success Rate，也就是排除參與者偏離情境或無法驗證的對話後，模型是否以正確參數完成所有必要的最終工具呼叫。

**排行榜結果** 在整體對話偏好方面，Gemini 3.1 Flash Live Preview - Minimal 以 1,046 Elo 領先，其後依序為 Gemini 3.1 Flash Live Preview - High 的 1,014、GPT-Realtime-1.5 的 1,000、GPT Realtime (Aug '25) 的 944，以及 ElevenLabs Agents Default Cascaded System 的 937。後者由 Scribe v2 Realtime、GPT-4o Mini 與 Eleven v3 組成，並使用預先註冊的工具 schema。

人工檢視對話後，Artificial Analysis 發現，偏好度較高的模型通常回應較快、聲音更自然，也較少出現不自然聲音或音訊瑕疵。不過，對話聽起來令人滿意，不代表任務一定完成：Gemini 3.1 Flash Live Preview - Minimal 雖取得最高 Preference Elo，Task Success Rate 卻是 74.6%。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/3fc7c0be1b6301c3.jpg)
> Gemini 3.1 Flash Live Preview - Minimal 在 Arena Preference Elo 指標中以 1,046 Elo 取得領先，Gemini 3.1 Flash Live Preview - High 以 1,014 Elo 居次，GPT-Realtime-1.5 則以 1,000 Elo 排名第三。

**任務成功率** Grok Voice Think Fast 2.0 High 以 94.7% 的 Task Success Rate 居首，GPT-Realtime-2.1 High 為 91.5%，ElevenLabs Agents Default Cascaded System 為 90.5%，GPT-Realtime-2 (High) 為 89.8%；GPT Realtime (Aug '25) 與 GPT-Realtime-2.1 Minimal 則同為 89.4%。這項差異顯示，模型可能用自然且肯定的語氣，讓使用者以為訂單或預約已完成，但實際上並未成功執行必要的最終工具呼叫。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/833fcf19398f1bfc.jpg)
> Gemini 3.1 Flash Live Preview - Minimal 以 1,046 Elo 在 Arena Preference Elo 領先，Gemini 3.1 Flash Live Preview - High 以 1,014 Elo 居次。

**速度與成本** 偏好度通常會隨 Time to First Audio（TTFA，首次音訊回應時間）縮短而上升。Gemini 3.1 Flash Live Preview - Minimal 的 TTFA 為 0.96 秒、Preference Elo 為 1,046；GPT-Realtime-2 (High) 為 1.14 秒與 914 Elo；Qwen Audio 3.0 Realtime Plus 則為 1.54 秒與 699 Elo，顯示回應速度是影響語音互動體驗的重要因素之一。

不同模型的價格與表現也有明顯差異。Gemini 3.1 Flash Live Preview - Minimal 每小時輸入音訊成本為 1.50 美元，雖在偏好度領先，但 Task Success Rate 為 74.6%；Grok Voice Think Fast 2.0 High 每小時 4.80 美元，任務成功率最高達 94.7%；GPT-Realtime-2.1 High 每小時 10.75 美元，任務成功率為 91.5%。因此，偏好度、任務可靠性與成本並未形成單一排名。

**評測限制** 除了「New Patient Dental Booking」範例外，Arena 目前未公開情境 prompt、工具 schema 與參與者指示，以降低模型針對評測過度調整的可能性。評測也使用經篩選、付費的第三方參與者，因此結果反映的是特定真人樣本與受控情境，不等同於所有使用者的實際體驗。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/faeba6ef44acd299.jpg)
> Gemini 3.1 Flash Live Preview - Minimal 在 Speech Agent Arena 的 Preference Elo 指標中以 1,046 分居首，其次為 Gemini 3.1 Flash Live Preview - High（1,014 分）與 GPT-Realtime-1.5（1,000 分）。

牙科預約範例要求模型先查詢會員資料，再確認週二上午 9:30 是否有空，若無法預約則改訂最早可用的上午時段，最後還要讀回預約日期與時間。評估流程會先確認參與者是否依照情境操作，再檢查模型是否完成每個必要的最終工具呼叫、使用正確參數，且沒有執行非預期的最終動作。

**完整評測脈絡** Artificial Analysis 同時以四個等權重面向建立 Speech to Speech Index：Speech Reasoning、Agentic Performance、Arena Preference 與 Task Success Rate，各占 25%。其中 Speech Reasoning 使用包含 1,000 個音訊問題的 Big Bench Audio；Agentic Performance 則依據 𝜏-Voice，測試航空、零售與電信等共 278 個情境，評估模型能否端到端解決客戶問題。

Artificial Analysis 將持續擴充原生音訊與 cascaded Speech to Speech 系統、模型供應商及情境。讀者可查看[Speech to Speech benchmark](https://artificialanalysis.ai/speech-to-speech)、[Speech Agent Arena 排行榜](https://artificialanalysis.ai/speech-to-speech/arena)、[評測概覽與對話範例](https://artificialanalysis.ai/speech-to-speech/arena/overview)，以及[完整評測方法](https://artificialanalysis.ai/methodology/speech-to-speech-benchmarking)。

## 媒體內容

**Gemini 3.1 Flash Live Preview - Minimal 在 Arena Preference Elo 指標中以 1,046 Elo 取得領先，Gemini 3.1 Flash Live Preview - High 以 1,014 Elo 居次，GPT-Realtime-1.5 則以 1,000 Elo 排名第三。**

**數據表**

| 模型 | Arena Preference Elo | Task Success Rate (%) |
| --- | --- | --- |
| Gemini 3.1 Flash Live Preview - Minimal | 1046 | 74.6% |
| Gemini 3.1 Flash Live Preview - High | 1014 | 71.8% |
| GPT-Realtime-1.5 | 1000 | 85.1% |
| GPT Realtime (Aug '25) | 944 | 89.4% |
| ElevenLabs Agents (Default Cascaded System) | 937 | 90.5% |
| Nova 2.0 Sonic (Mar 2026) | 917 | 57.1% |
| GPT-Realtime-2 (High) | 914 | 89.8% |
| GPT Realtime Mini (Oct '25) | 912 | 79.6% |
| Grok Voice Think Fast 2.0 High | 908 | 94.7% |
| GPT-Realtime-2.1 Minimal | 896 | 89.4% |
| GPT-Realtime-2.1 High | 892 | 91.5% |
| Deepgram Voice Agent (Default Cascaded System) | 882 | 73.7% |
| GPT-Realtime-2 (Minimal) | 881 | 84.7% |
| Nova Sonic | 852 | 57.1% |
| Grok Voice Think Fast 1.0 | 839 | 80.7% |
| Qwen3.5 Omni Plus Realtime | 799 | 62.0% |
| Qwen3.5 Omni Flash Realtime | 797 | 29.1% |
| GPT-Realtime-2.1 Mini Minimal | 792 | 76.7% |
| Inworld Realtime (Default Cascaded System) | 785 | 69.9% |
| Qwen Audio 3.0 Realtime Flash | 752 | 81.7% |
| Qwen Audio 3.0 Realtime Plus | 699 | 77.8% |

**Gemini 3.1 Flash Live Preview - Minimal 以 1,046 Elo 在 Arena Preference Elo 領先，Gemini 3.1 Flash Live Preview - High 以 1,014 Elo 居次。**

**數據表（1）Arena Preference Elo vs. Speed**

| 項目 | X | Y |
| --- | --- | --- |
| Gemini 3.1 Flash Live Preview - Minimal | 0.96s | 1046 |
| Gemini 3.1 Flash Live Preview - High | 2.98s | 1014 |
| GPT-Realtime-1.5 | 0.80s | 1000 |
| GPT Realtime (Aug '25) | 0.98s | 944 |
| Grok Voice Think Fast 2.0 High | 0.43s | 915 |
| GPT-Realtime-2.1 Minimal | 0.83s | 915 |
| GPT-Realtime-2 (High) | 1.18s | 915 |
| Nova 2.0 Sonic (Mar 2026) | 1.11s | 895 |
| GPT Realtime Mini (Oct '25) | 0.64s | 890 |
| GPT-Realtime-2.1 High | 1.23s | 890 |
| GPT-Realtime-2 (Minimal) | 1.11s | 885 |
| Grok Voice Think Fast 1.0 | 1.25s | 840 |
| Qwen3.5 Omni Plus Realtime | 2.65s | 798 |
| Qwen3.5 Omni Flash Realtime | 0.78s | 795 |
| GPT-Realtime-2.1 Mini Minimal | 0.85s | 790 |
| Qwen Audio 3.0 Realtime Flash | 1.53s | 750 |
| Qwen Audio 3.0 Realtime Plus | 1.53s | 700 |

**數據表（2）Task Success Rate vs. Speed**

| 項目 | X | Y |
| --- | --- | --- |
| Grok Voice Think Fast 2.0 High | 0.43s | 95.0% |
| GPT-Realtime-2 (High) | 1.18s | 91.5% |
| GPT-Realtime-2.1 High | 1.23s | 91.5% |
| GPT Realtime (Aug '25) | 0.95s | 90.0% |
| GPT-Realtime-2.1 Minimal | 0.83s | 87.0% |
| GPT-Realtime-1.5 | 0.80s | 85.0% |
| GPT Realtime Mini (Oct '25) | 0.64s | 85.0% |
| GPT-Realtime-2 (Minimal) | 1.11s | 85.0% |
| Qwen Audio 3.0 Realtime Flash | 1.55s | 82.0% |
| Grok Voice Think Fast 1.0 | 1.25s | 81.0% |
| GPT-Realtime-2.1 Mini Minimal | 0.85s | 79.5% |
| Qwen Audio 3.0 Realtime Plus | 1.55s | 77.5% |
| Gemini 3.1 Flash Live Preview - Minimal | 0.96s | 75.0% |
| Gemini 3.1 Flash Live Preview - High | 2.98s | 71.5% |
| Qwen3.5 Omni Plus Realtime | 2.65s | 62.0% |
| Nova 2.0 Sonic (Mar 2026) | 1.11s | 57.0% |
| Qwen3.5 Omni Flash Realtime | 0.78s | 28.0% |

**Gemini 3.1 Flash Live Preview - Minimal 在 Speech Agent Arena 的 Preference Elo 指標中以 1,046 分居首，其次為 Gemini 3.1 Flash Live Preview - High（1,014 分）與 GPT-Realtime-1.5（1,000 分）。**

**數據表（1）Arena Preference Elo vs. Cost per Hour of Input Audio**

| 項目 | X | Y |
| --- | --- | --- |
| Gemini 3.1 Flash Live Preview - Minimal | $1.5 | 1046 |
| Gemini 3.1 Flash Live Preview - High | $1.8 | 1014 |
| GPT Realtime Mini (Oct '25) | $3.0 | 913 |
| GPT-Realtime-2 (High) | $4.1 | 915 |
| GPT-Realtime-2 (Minimal) | $3.1 | 882 |
| Grok Voice Think Fast 2.0 High | $4.8 | 911 |
| Grok Voice Think Fast 1.0 | $3.0 | 838 |
| GPT-Realtime-2.1 Mini Minimal | $4.6 | 792 |
| Qwen Audio 3.0 Realtime Flash | $4.8 | 752 |
| Qwen Audio 3.0 Realtime Plus | $4.4 | 699 |
| GPT-Realtime-2.1 High | $10.8 | 890 |
| GPT Realtime (Aug '25) | $11.1 | 944 |
| GPT-Realtime-2.1 Minimal | $11.3 | 892 |
| GPT-Realtime-1.5 | $11.4 | 1000 |

**數據表（2）Task Success Rate vs. Cost per Hour of Input Audio**

| 項目 | X | Y |
| --- | --- | --- |
| Gemini 3.1 Flash Live Preview - Minimal | $1.5 | 74.6% |
| Gemini 3.1 Flash Live Preview - High | $1.8 | 71.7% |
| GPT Realtime Mini (Oct '25) | $3.0 | 79.5% |
| Grok Voice Think Fast 1.0 | $3.0 | 80.5% |
| GPT-Realtime-2 (Minimal) | $3.1 | 84.7% |
| GPT-Realtime-2 (High) | $4.1 | 89.8% |
| Qwen Audio 3.0 Realtime Plus | $4.4 | 78.0% |
| GPT-Realtime-2.1 Mini Minimal | $4.6 | 76.8% |
| Qwen Audio 3.0 Realtime Flash | $4.8 | 81.7% |
| Grok Voice Think Fast 2.0 High | $4.8 | 94.8% |
| GPT-Realtime-2.1 High | $10.8 | 91.6% |
| GPT Realtime (Aug '25) | $11.1 | 89.2% |
| GPT-Realtime-2.1 Minimal | $11.3 | 89.2% |
| GPT-Realtime-1.5 | $11.4 | 85.0% |

## 標籤

Benchmark, Voice, 新產品, Artificial Analysis
