# Marin 啟動 535B-A23B 公開訓練，以四階縮放實驗預測全程 loss

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：Percy Liang (@percyliang) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥 · 日期：2026-08-22

> 原始來源：https://x.com/percyliang/status/2090918065634684997

## 證據與延伸閱讀

- [Marin 啟動 535B-A23B 公開訓練，以四階縮放實驗預測全程 loss](https://x.com/m_newhaus/status/2091203537925492880) — 最後核對：2026-08-22 · 支持主張：素材具體列出 23.11T token dataset、40 semantic domains、JAX/GPU Expert Parallel、1% compute cost scaling ladder、100-day run、context extension 與長上下文 RL 實驗設計；media snapshot 再提供可追到的 Marin issue／程式檔名與 loss、gradient、MoE、tokens-per-second、MFU 圖表。
- [Marin 535B-A23B 公開訓練與 scaling ladder 實作 — @percyliang](https://x.com/percyliang/status/2090918067987689486) — 最後核對：2026-08-22 · 支持主張：第一方素材新增了可檢查的公開訓練實作：Harrier K40 overview 報告 23.11T tokens、40 個 semantic domains、五個 quality buckets，並標明 94.6% aligned sample 的 source makeup 只是 audit estimate；W&B report 提供 535B-A23B 18T Hero Run、五階 scaling ladder 與 MFU、tokens/sec、MoE drop fraction、parameter／gradient norm、cross-entropy 圖表；launch_scaling_ladder.py 公開五種 width 的 hero-shape recipe（384 routed experts、top-8、pooled-wave、dropless held-out evaluation），issue #8435 則記錄約 100 天、1% total compute 的 performance tracking，以及 535.3B total／22.76B ac…
- [Marin 535B-A23B 公開訓練與 scaling ladder 實作 — @percyliang](https://x.com/percyliang/status/2090918070617624811) — 最後核對：2026-08-22 · 支持主張：第一方素材新增了可檢查的公開訓練實作：Harrier K40 overview 報告 23.11T tokens、40 個 semantic domains、五個 quality buckets，並標明 94.6% aligned sample 的 source makeup 只是 audit estimate；W&B report 提供 535B-A23B 18T Hero Run、五階 scaling ladder 與 MFU、tokens/sec、MoE drop fraction、parameter／gradient norm、cross-entropy 圖表；launch_scaling_ladder.py 公開五種 width 的 hero-shape recipe（384 routed experts、top-8、pooled-wave、dropless held-out evaluation），issue #8435 則記錄約 100 天、1% total compute 的 performance tracking，以及 535.3B total／22.76B ac…
- [Marin 535B-A23B 公開訓練與 scaling ladder 實作 — github.com](https://github.com/marin-community/marin/issues/8435) — 官方 Repository · 最後核對：2026-08-22 · 支持主張：第一方素材新增了可檢查的公開訓練實作：Harrier K40 overview 報告 23.11T tokens、40 個 semantic domains、五個 quality buckets，並標明 94.6% aligned sample 的 source makeup 只是 audit estimate；W&B report 提供 535B-A23B 18T Hero Run、五階 scaling ladder 與 MFU、tokens/sec、MoE drop fraction、parameter／gradient norm、cross-entropy 圖表；launch_scaling_ladder.py 公開五種 width 的 hero-shape recipe（384 routed experts、top-8、pooled-wave、dropless held-out evaluation），issue #8435 則記錄約 100 天、1% total compute 的 performance tracking，以及 535.3B total／22.76B ac…
- [Marin 535B-A23B 公開訓練與 scaling ladder 實作 — storage.googleapis.com](https://storage.googleapis.com/marin-public/held/harrier-k40-cluster-overview/2026.08.18/index.html?revision=uniform-sampling) — 官方文件 · 最後核對：2026-08-22 · 支持主張：第一方素材新增了可檢查的公開訓練實作：Harrier K40 overview 報告 23.11T tokens、40 個 semantic domains、五個 quality buckets，並標明 94.6% aligned sample 的 source makeup 只是 audit estimate；W&B report 提供 535B-A23B 18T Hero Run、五階 scaling ladder 與 MFU、tokens/sec、MoE drop fraction、parameter／gradient norm、cross-entropy 圖表；launch_scaling_ladder.py 公開五種 width 的 hero-shape recipe（384 routed experts、top-8、pooled-wave、dropless held-out evaluation），issue #8435 則記錄約 100 天、1% total compute 的 performance tracking，以及 535.3B total／22.76B ac…
- [Marin 535B-A23B 公開訓練與 scaling ladder 實作 — wandb.ai](https://wandb.ai/marin-community/marin_moe/reports/535B-A23B-18T-Token-Hero-Run-Scaling-Ladder--VmlldzoxNzc2MDM5Ng) — 官方文件 · 最後核對：2026-08-22 · 支持主張：第一方素材新增了可檢查的公開訓練實作：Harrier K40 overview 報告 23.11T tokens、40 個 semantic domains、五個 quality buckets，並標明 94.6% aligned sample 的 source makeup 只是 audit estimate；W&B report 提供 535B-A23B 18T Hero Run、五階 scaling ladder 與 MFU、tokens/sec、MoE drop fraction、parameter／gradient norm、cross-entropy 圖表；launch_scaling_ladder.py 公開五種 width 的 hero-shape recipe（384 routed experts、top-8、pooled-wave、dropless held-out evaluation），issue #8435 則記錄約 100 天、1% total compute 的 performance tracking，以及 535.3B total／22.76B ac…
- [Marin 535B-A23B 公開訓練與 scaling ladder 實作 — github.com](https://github.com/marin-community/marin/blob/main/experiments/grug/moe_hero_ep/launch_scaling_ladder.py) — 官方 Repository · 最後核對：2026-08-22 · 支持主張：第一方素材新增了可檢查的公開訓練實作：Harrier K40 overview 報告 23.11T tokens、40 個 semantic domains、五個 quality buckets，並標明 94.6% aligned sample 的 source makeup 只是 audit estimate；W&B report 提供 535B-A23B 18T Hero Run、五階 scaling ladder 與 MFU、tokens/sec、MoE drop fraction、parameter／gradient norm、cross-entropy 圖表；launch_scaling_ladder.py 公開五種 width 的 hero-shape recipe（384 routed experts、top-8、pooled-wave、dropless held-out evaluation），issue #8435 則記錄約 100 天、1% total compute 的 performance tracking，以及 535.3B total／22.76B ac…
- [Marin 公開 535B-A23B 模型訓練全流程與資料](https://x.com/percyliang/status/2090918065634684997)
- [IMG_1 圖片：不同訓練進度的 cross-entropy 曲線](https://pbs.twimg.com/media/HQRrhYTaIAIH3YN.jpg?name=orig)
- [IMG_2 圖片：535B-A23B Scaling Ladder 的量測與預測曲線](https://pbs.twimg.com/media/HQRsmQua4AAb6Oy.jpg?name=orig)

## 中文摘要

Marin 啟動 535B-A23B 公開訓練，以四階縮放實驗預測全程 loss

**公開訓練計畫** Marin 的 535B-A23B 本週開訓。Percy Liang 公開的計畫是先完成 80% 預訓練，再進行 20% 中期訓練；全程使用 18.75T tokens，在 11 組 GB200 NVL72 上執行約 3 個月，訓練量約 2.7e24 FLOPs，之後才進入後訓練。正式啟動前，團隊先跑四階 scaling ladder，規模從 1.6B-A61M、48B tokens 到 27.7B-A1.2B、926B tokens，用來除錯並預測大型 hero run 的 loss。

**資料與可追溯性** 公開的 [Harrier K40 資料領域與品質總覽](https://storage.googleapis.com/marin-public/held/harrier-k40-cluster-overview/2026.08.18/index.html?revision=uniform-sampling) 顯示，候選資料庫共有 23.11T tokens，分成 40 個語意領域與 5 個校準後的品質分桶。各欄位的占比是候選資料庫的精確值；來源組成則是根據涵蓋 94.6% tokens 的確定性對齊樣本估算，不代表每筆資料的精確歸屬。

**Scaling ladder 與實作** Marin 的 Larry Dial 在 [W&B 報告](https://wandb.ai/marin-community/marin_moe/reports/535B-A23B-18T-Token-Hero-Run-Scaling-Ladder--VmlldzoxNzc2MDM5Ng)公開 hero run 與四個梯級測試的即時圖表。`launch_scaling_ladder.py` 則列出五種寬度共用的配方：

- 每個梯級使用 384 個 routed experts 與 top-8 routing。
- Expert 寬度與 latent 維度都是 hidden size 的一半；傳輸採 pooled-wave，也就是手寫的 expert-parallel all-to-all 實作。
- 資料混合使用 Harrier 2026.08.18 的 two-phase mixture，並跑不丟 token 的留出集評估（dropless held-out evaluation）。
- 規模從 d768、48B tokens、61M active／1.6B total parameters，一路到 d6144 hero、18T tokens、23B active／535B total parameters；最大規模約需 2.7e24 FLOPs。
- 每個 active parameter 配 791 tokens，gradient norm 與 parameter norm 每 10 steps 記錄一次。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/2b62d187e0199911.jpg)
> 圖中的 rav-ladder-d2048-v3 在多條其他曲線之前停止；各曲線停止進度不同，不能直接當成同一終點的排名。

**監測與透明度** Marin 在 [issue #8435](https://github.com/marin-community/marin/issues/8435) 公開設定與追蹤方式。Scaling ladder 約使用總計算量的 1%，用來預測模型表現、比較既有配方，並在約 100 天的訓練期間監控梯度範數、token 丟棄與各階段評估。Percy Liang 公開的完整航程計畫使用 18.75T tokens；issue 與程式配方則把 hero run 標為 18.0T tokens。模型規格為 535.3B total、22.76B active parameters，採 64 路 Expert Parallel（EP64）的 MoE 預訓練，以及針對 GB200 手寫的 pooled-wave all-to-all。

<video src="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/1787428401342-wsvlvspf.mp4" poster="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/fd6abd235a51a88a.jpg" controls playsinline preload="metadata" style="max-width:100%;height:auto;display:block;margin:1rem 0"></video>
> 來源：[@m_newhaus](https://x.com/m_newhaus/status/2091203537925492880)（回覆）｜影片逐段展示 W&B 報告頁面、訓練圖表與 run list。

**作者觀點** Mary Newhauser 認為，這可能是目前最開放、透明的模型訓練之一，但她也指出 OLMo 3 的公開資料同樣完整。她整理的內容還包括 JAX/GPU Expert Parallel、4k→8k→65k→262k 脈絡長度延伸、模型浮點運算使用率（MFU）、loss、每秒 token 數、expert balancing，以及長上下文 RL 的硬體規格與 cooldown 實驗設計。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/ce85f1996760716a.jpg)
> 圖中實線實心點代表量測，虛線空心點代表外推／預測；hero 曲線以虛線星形標記外推到訓練進度 100%。

## 媒體內容

**影片逐段展示 W&B 報告頁面、訓練圖表與 run list。**

**影片中的 Prompt 與操作**

操作步驟：

1. （00:03）捲動頁面檢視實驗執行列表與 Scaling Ladder 圖表
2. （00:06）將游標移至 Scaling Ladder 的 `train/cross_entropy_loss` 圖表上檢視各執行的詳細數值
3. （00:08）繼續向下捲動檢視更多指標折線圖與下方執行清單

**圖中實線實心點代表量測，虛線空心點代表外推／預測；hero 曲線以虛線星形標記外推到訓練進度 100%。**

**數據表**

| 項目 | 數值 |
| --- | --- |
| 圖表標題：`535B-A23B 18T tok Scaling Ladder`,X 軸：`training progress (percentile %)`,範圍 0–100。,Y 軸：`dropless paloma macro-loss`。 |  |
| 圖例註明：`solid + filled | measured` |
| `dashed + hollow | extrapolated / predicted`。,四個梯級模型以實線實心點呈現已量測區段,27.7B-A1.2B 曲線末段另有虛線空心點的外推值。,hero · 535B-A23B · 18T tok 使用虛線空心星形標記,外推到訓練進度 100%。 |

## 標籤

研究論文, 開源專案, LLM, Marin
