# Marin 公開訓練 535B-A23B：以即時紀錄提高大型模型訓練透明度

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：Percy Liang (@percyliang) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥🔥🔥🔥 · 日期：2026-08-23

> 原始來源：https://x.com/percyliang/status/2090918065634684997

## 證據與延伸閱讀

- [Marin 公開訓練 535B-A23B：以即時紀錄提高大型模型訓練透明度。](https://x.com/m_newhaus/status/2091203537925492880) — 一手來源 · 最後核對：2026-08-22 · 支持主張：素材具體列出 23.11T token dataset、40 semantic domains、JAX/GPU Expert Parallel、1% compute cost scaling ladder、100-day run、context extension 與長上下文 RL 實驗設計；media snapshot 再提供可追到的 Marin issue／程式檔名與 loss、gradient、MoE、tokens-per-second、MFU 圖表。
- [Marin 535B-A23B 公開訓練與 scaling ladder 實作 — @percyliang](https://x.com/percyliang/status/2090918067987689486) — 一手來源 · 最後核對：2026-08-22 · 支持主張：第一方素材新增了可檢查的公開訓練實作：Harrier K40 overview 報告 23.11T tokens、40 個 semantic domains、五個 quality buckets，並標明 94.6% aligned sample 的 source makeup 只是 audit estimate；W&B report 提供 535B-A23B 18T Hero Run、五階 scaling ladder 與 MFU、tokens/sec、MoE drop fraction、parameter／gradient norm、cross-entropy 圖表；launch_scaling_ladder.py 公開五種 width 的 hero-shape recipe（384 routed experts、top-8、pooled-wave、dropless held-out evaluation），issue #8435 則記錄約 100 天、1% total compute 的 performance tracking，以及 535.3B total／22.76B ac…
- [Marin 535B-A23B 公開訓練與 scaling ladder 實作 — @percyliang](https://x.com/percyliang/status/2090918070617624811) — 一手來源 · 最後核對：2026-08-22 · 支持主張：第一方素材新增了可檢查的公開訓練實作：Harrier K40 overview 報告 23.11T tokens、40 個 semantic domains、五個 quality buckets，並標明 94.6% aligned sample 的 source makeup 只是 audit estimate；W&B report 提供 535B-A23B 18T Hero Run、五階 scaling ladder 與 MFU、tokens/sec、MoE drop fraction、parameter／gradient norm、cross-entropy 圖表；launch_scaling_ladder.py 公開五種 width 的 hero-shape recipe（384 routed experts、top-8、pooled-wave、dropless held-out evaluation），issue #8435 則記錄約 100 天、1% total compute 的 performance tracking，以及 535.3B total／22.76B ac…
- [Marin 535B-A23B 公開訓練與 scaling ladder 實作 — github.com](https://github.com/marin-community/marin/issues/8435) — 官方文件 · 最後核對：2026-08-22 · 支持主張：第一方素材新增了可檢查的公開訓練實作：Harrier K40 overview 報告 23.11T tokens、40 個 semantic domains、五個 quality buckets，並標明 94.6% aligned sample 的 source makeup 只是 audit estimate；W&B report 提供 535B-A23B 18T Hero Run、五階 scaling ladder 與 MFU、tokens/sec、MoE drop fraction、parameter／gradient norm、cross-entropy 圖表；launch_scaling_ladder.py 公開五種 width 的 hero-shape recipe（384 routed experts、top-8、pooled-wave、dropless held-out evaluation），issue #8435 則記錄約 100 天、1% total compute 的 performance tracking，以及 535.3B total／22.76B ac…
- [Marin 535B-A23B 公開訓練與 scaling ladder 實作 — storage.googleapis.com](https://storage.googleapis.com/marin-public/held/harrier-k40-cluster-overview/2026.08.18/index.html?revision=uniform-sampling) — 官方文件 · 最後核對：2026-08-22 · 支持主張：第一方素材新增了可檢查的公開訓練實作：Harrier K40 overview 報告 23.11T tokens、40 個 semantic domains、五個 quality buckets，並標明 94.6% aligned sample 的 source makeup 只是 audit estimate；W&B report 提供 535B-A23B 18T Hero Run、五階 scaling ladder 與 MFU、tokens/sec、MoE drop fraction、parameter／gradient norm、cross-entropy 圖表；launch_scaling_ladder.py 公開五種 width 的 hero-shape recipe（384 routed experts、top-8、pooled-wave、dropless held-out evaluation），issue #8435 則記錄約 100 天、1% total compute 的 performance tracking，以及 535.3B total／22.76B ac…
- [Marin 535B-A23B 公開訓練與 scaling ladder 實作 — wandb.ai](https://wandb.ai/marin-community/marin_moe/reports/535B-A23B-18T-Token-Hero-Run-Scaling-Ladder--VmlldzoxNzc2MDM5Ng) — 官方文件 · 最後核對：2026-08-22 · 支持主張：第一方素材新增了可檢查的公開訓練實作：Harrier K40 overview 報告 23.11T tokens、40 個 semantic domains、五個 quality buckets，並標明 94.6% aligned sample 的 source makeup 只是 audit estimate；W&B report 提供 535B-A23B 18T Hero Run、五階 scaling ladder 與 MFU、tokens/sec、MoE drop fraction、parameter／gradient norm、cross-entropy 圖表；launch_scaling_ladder.py 公開五種 width 的 hero-shape recipe（384 routed experts、top-8、pooled-wave、dropless held-out evaluation），issue #8435 則記錄約 100 天、1% total compute 的 performance tracking，以及 535.3B total／22.76B ac…
- [Marin 535B-A23B 公開訓練與 scaling ladder 實作 — github.com](https://github.com/marin-community/marin/blob/main/experiments/grug/moe_hero_ep/launch_scaling_ladder.py) — 官方文件 · 最後核對：2026-08-22 · 支持主張：第一方素材新增了可檢查的公開訓練實作：Harrier K40 overview 報告 23.11T tokens、40 個 semantic domains、五個 quality buckets，並標明 94.6% aligned sample 的 source makeup 只是 audit estimate；W&B report 提供 535B-A23B 18T Hero Run、五階 scaling ladder 與 MFU、tokens/sec、MoE drop fraction、parameter／gradient norm、cross-entropy 圖表；launch_scaling_ladder.py 公開五種 width 的 hero-shape recipe（384 routed experts、top-8、pooled-wave、dropless held-out evaluation），issue #8435 則記錄約 100 天、1% total compute 的 performance tracking，以及 535.3B total／22.76B ac…
- [Percy Liang 說明 535B-A23B 訓練計畫](https://x.com/percyliang/status/2090918065634684997)

## 中文摘要

Marin 公開訓練 535B-A23B：以即時紀錄提高大型模型訓練透明度。

**公開訓練計畫**  
Percy Liang 表示，535B-A23B 已於 2026 年 8 月啟動訓練。整體行程規劃將 80% 用於 pretraining、20% 用於 midtraining，在 11 座 GB200 NVL72 上運算約 3 個月，預計處理 18.75T token、消耗約 2.7e24 FLOPs，post-training 之後再進行。Mary Newhauser 認為，這可能是目前最開放、最透明的模型訓練之一，但也指出 OLMo 3 的公開程度同樣令人印象深刻。

**資料與可稽核性**  
公開資料集包含 23.11T token，分布於 40 個語意領域與五個品質分級。Harrier K40 overview 以涵蓋資料庫 94.6% token 的確定性對齊樣本估算來源組成；各欄位比例可精確計算，但來源組成仍只是稽核估算，不是完整的來源歸屬。

**Scaling ladder 與觀測**  
正式 Hero Run 前，團隊先用四階 scaling ladder，從 1.6B-A61M、48B token 測到 27.7B-A1.2B、926B token，用來除錯、預測最終損失，也檢查中間 checkpoint 是否符合進度。公開的 W&B 報告提供五次 run 與即時圖表，包括 MFU、每秒 token 數、MoE 丟棄比例、參數範數、梯度範數與交叉熵損失；整套效能追蹤約占總運算量的 1%，預計持續約 100 天。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/2b62d187e0199911.jpg)
> 圖表比較 Hero Run 與四個 scaling ladder 模型的 cross-entropy loss 曲線。

**模型與實作細節**  
公開的 `launch_scaling_ladder.py` 以五種模型寬度實作同一套 hero-shape recipe：

- 每個階段使用 384 個 routed experts 與 top-8 routing。
- 採用 hidden/2-wide experts、pooled-wave transport，以及 Harrier 2026.08.18 two-phase mixture。
- 評估使用 dropless held-out evaluation。
- 規模從 d768、48B token、61M active／1.6B total parameters，擴展到 d6144 Hero Run。Hero Run 的公開規格為 18T token、23B active／535B total parameters，約需 2.7e24 FLOPs。
- step budget 以每個 active parameter 791 token 設定，梯度範數與參數範數每 10 steps 記錄一次。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/ce85f1996760716a.jpg)
> 圖表比較 Hero Run 與四個 scaling ladder 模型在訓練進度中的 dropless paloma macro-loss 曲線。

**工程與不確定性**  
issue #8435 指出，完整方案是 EP64 MoE pretrain，規格為 535.3B total、22.76B active parameters 與 18.0T token，並在 GB200 上使用手工打造的 pooled-wave expert-parallel all-to-all 實作。團隊同時公開 JAX/GPU Expert Parallel、硬體規格、context 從 4k 延伸至 8k、65k、262k 的訓練指標，以及長上下文 RL 的 cooldown 實驗設計；Percy Liang 也坦言這是團隊迄今最大規模的 run，因此過程中仍可能遇到意料之外的問題。可查看 [Harrier K40 資料報告](https://storage.googleapis.com/marin-public/held/harrier-k40-cluster-overview/2026.08.18/index.html?revision=uniform-sampling)、[W&B 即時報告](https://wandb.ai/marin-community/marin_moe/reports/535B-A23B-18T-Token-Hero-Run-Scaling-Ladder--VmlldzoxNzc2MDM5Ng) 與 [Marin issue #8435](https://github.com/marin-community/marin/issues/8435)。

<video src="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/1787464626750-ics3h3em.mp4" poster="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/fd6abd235a51a88a.jpg" controls playsinline preload="metadata" style="max-width:100%;height:auto;display:block;margin:1rem 0"></video>
> 來源：[@m_newhaus](https://x.com/m_newhaus/status/2091203537925492880)｜瀏覽器介面呈現 535B-A23B 18T Token Hero Run 與 Scaling Ladder 專案實驗的多項訓練損失與效能評估圖表與執行列表

## 媒體內容

**瀏覽器介面呈現 535B-A23B 18T Token Hero Run 與 Scaling Ladder 專案實驗的多項訓練損失與效能評估圖表與執行列表**

**影片中的 Prompt 與操作**

操作步驟：

1. @00:03捲動頁面檢視 Scaling Ladder 圖表與執行狀態列表
2. @00:05游標懸停在圖表數據點上檢視詳細數值
3. @00:08向下捲動瀏覽更多訓練監控圖表

## 標籤

開源專案, 研究論文, LLM, NVIDIA
