# Cerebras 與 AMD 推出全新分離式推論解決方案，結合 AMD Helios 系統與 Cerebras Wafer-Scale Engine，為 Agentic 程式開發提供超低延遲與高產出能耐

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：Cerebras (@cerebras) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥🔥🔥🔥 · 日期：2026-07-24

> 原始來源：https://x.com/cerebras/status/2080349251318530263

## 證據與延伸閱讀

- [Cerebras與AMD推出分離式推論解決方案](https://www.cerebras.ai/press-release/amd-and-cerebras-announce-industry-leading-ultra-low-latency-and-high-throughput-ai-inference)
- [2026年7月24日共同揭曉合作方案](https://x.com/cerebras/status/2080349251318530263)

## 中文摘要

Cerebras 與 AMD 推出全新分離式推論解決方案，結合 AMD Helios 系統與 Cerebras Wafer-Scale Engine，為 Agentic 程式開發提供超低延遲與高產出能耐。

Cerebras 於 2026 年 7 月 24 日透過官方帳號發文宣布，這項由 AMD 與 Cerebras 攜手打造的合作方案，在 Advancing AI 2026 大會上由 AMD 執行長 Lisa Su 與 Cerebras 執行長 Andrew Feldman 共同揭曉。該解決方案專為應對 Agentic AI 與即時應用的需求而設計，旨在將正確的運算引擎指派給推論管線的每個階段，實現大規模生產環境中的極速表現。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/82bfafb502110180.jpg)
> 兩位講者站在大型藍色背板前，背板上展示著「AMD INSTINCT」晶片及其效能數據，標題為「The Most Powerful Solution For Ultra Low Latency Inference」。

**核心效能特色**
- 提供生產環境中最快速的 token 生成速度。
- 具備 5 倍之高的產能與能源效率（每瓦 token 數）。
- 支援規模達 1 兆以上參數的尖端大型語言模型。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/fa4e7a1a328ae028.png)
> AMD 與 Cerebras 合作展示的一種結合超高吞吐量與超低延遲的新型解構式推論架構。

**架構與運作方式**
- 採用工作負載最佳化的分離式推論（disaggregated inference）架構，將推論管線的兩個主要階段獨立最佳化。
- AMD Helios 提供高產出的提示引擎與大規模擴充能力，負責處理提示詞與大型記憶空間。
- Cerebras Wafer-Scale Engine 則專責記憶體頻寬密集的 token 生成階段，提供超低延遲的解碼效能。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/9b6233d72dddd8b1.png)
> AMD 與 Cerebras 合作推出的新型解構式推論架構，結合了高吞吐量的伺服器系統與低延遲的運算設備。

**上市規劃**
- Cerebras 計劃在其資料中心部署 AMD Helios 系統。
- 此聯合解決方案預計於 2026 年下半年透過 Cerebras Inference Cloud（詳細資訊可參閱 [Cerebras 官方新聞稿](https://www.cerebras.ai/press-release/amd-and-cerebras-announce-industry-leading-ultra-low-latency-and-high-throughput-ai-inference)）首次推出。

## 標籤

硬體, 新產品, Agent, 產業趨勢, Cerebras, AMD
