# Google 發布 Gemini 3.5 Transcribe：支援 85 種以上語言與即時語音理解

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：Google AI (@GoogleAI) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥🔥🔥 · 日期：2026-08-27

> 原始來源：https://x.com/GoogleAI/status/2092660092135199039

## 證據與延伸閱讀

- [Google 發布 Gemini 3.5 Transcribe：支援 85 種以上語言與即時語音理解。](https://x.com/OfficialLoganK/status/2092660925509890397) — 一手來源 · 最後核對：2026-08-27 · 支持主張：The post specifies function calling, realtime streaming and lower-WER transcription among the new model's capabilities, alongside custom vocabulary, multi-speaker identification and over 85 languages.
- [Gemini 3.5 Transcribe — blog.google](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe) — 官方文件 · 最後核對：2026-08-27 · 支持主張：Adds a named Gemini transcription model with streaming and prerecorded APIs, custom vocabulary, multi-speaker attribution and public-preview availability.
- [Gemini 3.5 Transcribe — @sundarpichai](https://x.com/sundarpichai/status/2092659467284517088) — 一手來源 · 最後核對：2026-08-27 · 支持主張：Adds a named Gemini transcription model with streaming and prerecorded APIs, custom vocabulary, multi-speaker attribution and public-preview availability.
- [Gemini 3.5 Transcribe — @GoogleAI](https://x.com/GoogleAI/status/2092660089509314735) — 官方文件 · 最後核對：2026-08-27 · 支持主張：Adds context-aware and multimodal behavior beyond basic transcription: filler removal, formatting, screen-context voice commands, and polished email drafts from messy voice input or local files.
- [Gemini 3.5 Transcribe — @GoogleDeepMind](https://x.com/GoogleDeepMind/status/2092659223591039374) — 官方文件 · 最後核對：2026-08-27 · 支持主張：Details handling of complex phone numbers, postal codes, and order IDs in noisy environments, filler-word removal, auto-formatting, custom vocabulary, and speech transcription in 85+ languages.
- [Gemini 3.5 Transcribe — @GoogleDeepMind](https://x.com/GoogleDeepMind/status/2092659221477077101) — 官方文件 · 最後核對：2026-08-27 · 支持主張：Introduces Gemini 3.5 Transcribe as a new speech-to-text model; the accompanying media shows a transcription interface with custom-vocabulary control.

## 中文摘要

Google 發布 Gemini 3.5 Transcribe：支援 85 種以上語言與即時語音理解。

**核心能力** Gemini 3.5 Transcribe 主打更精確的轉錄品質與較低 WER（詞錯誤率），並支援以下功能：

- 透過 Live API 提供即時串流，並透過 Interactions API 處理預錄音訊。
- 自動辨識 85 種以上語言，並能識別多位講者及其意圖。
- 讓使用者加入自訂 vocabulary，以適應專有名詞、獨特人名與產品名稱。
- 支援 function calling，能把語音理解結果銜接到應用程式操作。
- 在雜訊環境中更準確處理複雜電話號碼、郵遞區號與訂單編號。
- 自動移除「嗯」「呃」等填充詞，並替未整理的口語內容套用格式。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/53b180528752cbb0.jpg)
> 來源：[@OfficialLoganK](https://x.com/OfficialLoganK/status/2092660925509890397)｜Gemini 3.5 Transcribe Live 在圖示 25 個 top locales、五款服務的串流語音辨識比較中，以 5.50% 取得最低詞錯誤率（越低越好）。

**超越聽寫** Google AI 將它定位為能理解情境的智慧聽寫工具，不只是把聲音轉成文字。模型可結合螢幕內容執行語音指令，也能把凌亂的語音輸入與本機檔案整理成潤飾後的電子郵件草稿；官方展示的 Gmail 示範中，語音內容被整理成包含段落、摘要與條列重點的郵件。這些描述仍屬 Google 的產品宣傳與示範，並非代表所有應用程式都具備相同整合程度。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/d82e5191fa274354.png)
> 來源：[@GoogleDeepMind](https://x.com/GoogleDeepMind/status/2092659223591039374)｜Gemini 3.5 Transcribe 在 FLEURS (top locales*) 非即時語音辨識評測中以 5.04% 詞錯誤率（WER）領先 Google Cloud Chirp 3（5.66%）、ElevenLabs Scribe v2（7.12%）、OpenAI GPT Live Transcribe（7.47%）與 Deepgram Nova-3（11.35%）。

**可用性與示範** 官方文件將 Gemini 3.5 Transcribe 列為公開預覽；API 已可在 Google AI Studio 與 Gemini Enterprise 使用，消費端則可在 macOS 上的 Gemini app 試用，Google DeepMind 另提到 Android 上的 Gboard；開發者可從 [Google AI Studio 與 Google Antigravity](https://goo.gle/4gzP1K8) 開始建置。示範畫面顯示使用者先加入「Gemini 3.5 Transcribe」、「NeuroWave」與「Projekt Jiskra」等自訂詞彙，再開始轉錄；畫面同時標示流程經過縮短與模擬，實際相容性及可用性可能有所不同。

<video src="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/1787804330636-7rad42si.mp4" poster="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/9014a9f4c90e2967.jpg" controls playsinline preload="metadata" style="max-width:100%;height:auto;display:block;margin:1rem 0"></video>
> 來源：[@sundarpichai](https://x.com/sundarpichai/status/2092659467284517088)｜Gemini 3.5 Transcribe 語音轉文字工具的自訂詞彙輸入介面與語音轉錄示範畫面

<video src="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/1787804366276-88ij611y.mp4" poster="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/b44bf600e0d29eb3.jpg" controls playsinline preload="metadata" style="max-width:100%;height:auto;display:block;margin:1rem 0"></video>
> 來源：[@GoogleAI](https://x.com/GoogleAI/status/2092660089509314735)｜macOS 上的 Gemini app 語音功能與郵件草擬介面

<video src="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/1787804383975-w3k49fe7.mp4" poster="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/71beda1224e4a6c2.jpg" controls playsinline preload="metadata" style="max-width:100%;height:auto;display:block;margin:1rem 0"></video>
> 來源：[@GoogleDeepMind](https://x.com/GoogleDeepMind/status/2092659221477077101)｜Gemini 3.5 Transcribe 示範介面正在轉錄語音並即時適應自訂辭彙

## 媒體內容

**Gemini 3.5 Transcribe 語音轉文字工具的自訂詞彙輸入介面與語音轉錄示範畫面**

**影片中的 Prompt 與操作**

Prompt（00:00）：

```
Projekt Jiskra
```

操作步驟：

1. （00:00）自訂字詞欄位顯示輸入內容與已加入的詞彙
2. （00:04）介面顯示 Start transcription 按鈕，接著進入轉錄狀態

**逐字稿**

- `00:05` 歡迎來到 Gemini 3.5 的轉錄示範。（Welcome to the Gemini 3.5 Transcribe demo.）
- `00:08` 即使有這些背景噪音，它也能輕鬆（Even with this background noise, it easily）
- `00:11` 辨識出 NeuroWave 這類自訂詞，以及像伺服器 ID B742X 這類英數混合代碼。看它在我切換成捷克文時（catches custom words like NeuroWave and alphanumeric codes like server ID B742X. Watch it adapt）
- `00:20` 立即適應。（instantly as I switch to Czech.）
- `00:23` 這真的運作得非常好，也能精準記錄我們的（To funguje opravdu dobře a přesně to zapíše i náš）
- `00:27` Jiskra 專案，接著切回英文來做總結。（projekt Jiskra before dropping back to English to wrap it up.）

**macOS 上的 Gemini app 語音功能與郵件草擬介面**

**影片中的 Prompt 與操作**

操作步驟：

1. （00:10）畫面顯示以功能鍵口述郵件的過程
2. （01:10）三張白板圖片顯示為選取狀態，並出現將內容整理成表格與重點摘要的語音要求
3. （01:34）郵件視窗顯示生成後的表格摘要

**逐字稿**

- `00:00` 這是 macOS 上 Gemini 新推出的語音功能。（Here's a new Gemini app voice feature on macOS.）
- `00:02` 最近我和團隊一起參加了一場工作坊，（So I was at a workshop with my team recently,）
- `00:05` 我想寫一封簡短的感謝總結 email。（and I just want to write a brief thank you summary email.）
- `00:08` 所以我會按住這裡的（So I'm going to hold down the）
- `00:09` 功能鍵，然後口述我的 email。（function key here and dictate my email.）
- `00:12` 來吧。（Here we go.）
- `00:13` 嗨，大家真的很感謝星期（Hey guys, really appreciate the workshop on）
- `00:16` 一的工作坊。（Monday.）
- `00:16` 不，抱歉。（No, sorry.）
- `00:17` 等等，是，不對，是星期二。（Wait, it was, no, it was Tuesday.）
- `00:19` 星期二。（Tuesday.）
- `00:20` 天啊，已經星期五了。（Gosh, it's already Friday.）
- `00:22` 總之，感覺效率超高。（So anyway, felt super productive.）
- `00:24` 真的很感謝這一點。（Really appreciated that.）
- `00:26` 真的很感謝（Really appreciated that）
- `00:27` 大家都把想法帶到桌面上來討論。（everyone brought ideas to the table.）
- `00:30` 整體感覺就是氣氛非常好。（It just felt like a really good energy.）
- `00:33` 好，總結一下。（And okay, summary.）
- `00:34` 所以，不過你知道，我們還是按照計畫進行。（So, but you know, we're still on track.）
- `00:36` 上線日期一如往常，更多細節很快就會公布。（Ship date, same as always, and more details to come shortly.）
- `00:39` 可以幫我把這段整理一下，讓它聽起來簡潔、專業，（And can you just clean that up and make it sound crisp and professional,）
- `00:43` 但也要充滿熱情，好嗎？（but enthusiastic, please?）
- `00:45` 我放開功能鍵。（I released the function key.）
- `00:47` 它會把這些喋喋不休的內容整理並思考一遍。（It's going to take that blabber and think about it.）
- `00:50` 來吧。（And here we go.）
- `00:52` 它把內容整理成一封乾淨俐落、可以直接寄出的 email。（It summarized it into a nice clean email that is ready to send.）
- `00:56` 現在我還可以，（And what I can do now,）
- `00:57` 在這個視窗後面做一件小小的額外操作。（a little extra thing is behind this window.）
- `01:00` 我其實用手機拍了三張（I actually have three images that I snapped）
- `01:03` 我們使用的白板照片。（with my phone of the whiteboard we had.）
- `01:05` 所以我可以在這裡選取那些圖片，然後再次呼叫（So I can actually select those images here and I can invoke）
- `01:09` Gemini，說些像是：嗨，（Gemini again and say something like, Hey,）
- `01:12` 可以把這裡的三張圖片整理成條列（can you take these three images here and put them into bullet）
- `01:16` 嗎？（points?）
- `01:16` 呃，不對，等等，不要條列。（Uh, no, wait, not bullet points.）
- `01:18` 我其實想要的是表格。（I actually want a table.）
- `01:20` 你可以幫我想一下，應該要怎麼（And can you just like, how should we）
- `01:22` 整理嗎？（organize it?）
- `01:23` 嗯，按照日期整理。（Um, put it, make it by day.）
- `01:25` 也就是按照日期整理成表格，並列出每天的（So like table organized by day and the key takeaways for each）
- `01:29` 重點收穫。（day.）
- `01:29` 所以我把游標移到這裡，就放開了功能鍵。（So I released that just moving my cursor over here.）
- `01:35` 來吧。（And here we go.）
- `01:37` 我在這裡得到了一份很棒的摘要，（I've got a nice summary here）
- `01:40` 來自那些圖片。（from those images.）
- `01:41` Gemini Mic 最實用的一點是，首先，（And this is kind of what's really neat about Gemini Mic is that first of all,）
- `01:46` 它能把我這種沒有結構的思緒串流，整理得更清楚、（it's able to take my kind of unstructured stream of thought and kind of helped me make it clear and）
- `01:51` 更有條理，方便用於書面表達，（cogent for written purposes,）
- `01:53` 而且我也能在這台本機電腦的情境中工作。（but I'm also able to work in the context of my local computer here.）
- `01:57` 所以（So）
- `01:57` 我可以參照檔案，實際上也能直接操作資料、參照內容，不必（I'm able to reference files and really just manipulate data and reference things without）
- `02:03` 切換到另一個應用程式，（needing to switch to another app,）
- `02:04` 整個過程感覺非常順暢，也真的很有幫助。（which just feels really seamless and really helpful.）

**Gemini 3.5 Transcribe 示範介面正在轉錄語音並即時適應自訂辭彙**

**影片中的 Prompt 與操作**

操作步驟：

1. （00:04）介面顯示「Start transcription」按鈕
2. （00:05）介面進入錄音狀態，顯示「Stop recording」與刪除按鈕

**逐字稿**

- `00:05` 歡迎來到 Gemini 3.5 的轉錄示範。（Welcome to the Gemini 3.5 Transcribe demo.）
- `00:08` 即使有這些背景噪音，它也能輕鬆辨識自訂（Even with this background noise it easily catches custom）
- `00:12` 詞彙，像是 NeuroWave，以及像伺服器 ID B742X 這樣的英數字元代碼。看著它在我（words like NeuroWave and alphanumeric codes like server ID B742X. Watch it adapt instantly as I）
- `00:21` 切換成捷克語時即時適應。（switch to Czech.）
- `00:23` 這真的運作得非常好，也精準地記錄下我們的 Jiskra 專案。（To funguje opravdu dobře a přesně to zapíše i náš projekt Jiskra.）
- `00:28` 接著（Before）
- `00:28` 切回英文，完成示範。（dropping back to English to wrap it up.）

## 標籤

Gemini, 新產品, 功能更新, STT, Google
