# Sapient Intelligence 發布 PRAXIST Beta，75 項 MLE-Bench 任務取得 49 金牌

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：Sapient Intelligence (@Sapient_Int) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥🔥🔥 · 日期：2026-08-28

> 原始來源：https://x.com/Sapient_Int/status/2093141554500338019

## 證據與延伸閱讀

- [Sapient Intelligence 發布 PRAXIST Beta，75 項 MLE-Bench 任務取得 49 金牌。](https://arxiv.org/abs/2608.25955) — 一手來源 · 最後核對：2026-08-28 · 支持主張：第一方 launch 與論文補足 75 個 MLE-Bench 任務中的 49 gold/60 medals、記錄模型花費、solution lineages 與 task-specific graders；本貼再提供 rocket simulation 100% safe-landing 與 industrial SLAM error 由 9.37 cm 降至 5.01 cm 的公司報告結果。
- [Sapient Intelligence 發布 PRAXIST Beta，75 項 MLE-Bench 任務取得 49 金牌](https://x.com/Sapient_Int/status/2093141554500338019)

## 中文摘要

Sapient Intelligence 發布 PRAXIST Beta，75 項 MLE-Bench 任務取得 49 金牌。

系統在 75 個 MLE-Bench 任務中取得 49 個 gold、60 個 total medals，但相關成果仍受評測設定與成本範圍限制。

**運作方式** PRAXIST Beta 讓使用者先定義 目標、限制、成功條件、評估器與預算，再由多個 Research Peers 同時提出、驗證彼此競爭的研究路徑。系統會把實驗 artifacts、task-specific graders 的結果，以及「哪些失敗約束導致結果改善」的解釋串成 solution lineages，讓後續 Agent 不只繼承成功方案，也能理解過往失敗原因。詳情見 Sapient Intelligence 的[第一方發布說明](https://x.com/Sapient_Int/status/2093141554500338019)與[研究論文](https://arxiv.org/abs/2608.25955)。

**MLE-Bench 結果** 最終完成的 75-task sweep 中，有 60 個任務取得任何 medal，49 個達到 gold；記錄的 model spend 約為 `$3,054`。作為對照，Claude Code 加上 Opus 4.8 的 baseline 取得 34 個 gold，記錄 model spend 為 `$38,370`。這些數字是公司在自身 setup 下、搭配各任務專用 grader 所做的受控評估，不能直接換算成普遍適用的產品價格或成本比例；記錄金額也不包含更廣泛的組織與基礎架構支出。

**合作環境示範** Sapient Intelligence 另報告四個 open-ended demonstrations，涵蓋控制、交易、SLAM 與 fusion 等領域：

- rocket simulation 的 safe-landing rate 為 100%。
- industrial SLAM error 從 9.37 公分降至 5.01 公分。

上述 safe-landing 與 SLAM 數字是特定 setup 下的公司報告結果，不代表普遍的 coding 或 robotics 保證；四個示範也各自使用 domain-specific tools 與 evaluators，並非同一套標準化的跨領域 benchmark。

<video src="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/1787909040743-9j6rqbsy.mp4" poster="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/351fbc9f93f2bd27.jpg" controls playsinline preload="metadata" style="max-width:100%;height:auto;display:block;margin:1rem 0"></video>
> PRAXIST Beta autonomous research team 在辦公室展示自主研究與優化火箭回收軌跡的介面

**評估解讀** PRAXIST 的重點不只是得到更高分，而是把研究過程、成本、評分依據與 solution lineage 一併保存，形成由 evaluator 驅動的研究 loop。比較 baseline 時仍須逐項核對任務內容、grader、執行設定與成本涵蓋範圍，避免把一次評測 sweep 的結果誤讀成全面性的能力或商業成本承諾。

## 媒體內容

**PRAXIST Beta autonomous research team 在辦公室展示自主研究與優化火箭回收軌跡的介面**

**影片中的 Prompt 與操作**

操作步驟：

1. （00:25）- 在登入介面輸入並點擊進入
2. （00:35）- 在工作區設定 objective 與 baseline 並按下 Enter 送出
3. - 於桌面螢幕檢視 Success-rate evolution 演化圖與最新結果

**逐字稿**

- `00:00` 有多糟？（How bad is it?）
- `00:01` 動力下降導引求解器還是無法收斂。（Power descent guidance solver still won't converge.）
- `00:04` 縮小著陸橢圓，我們就會失去推進劑儲備。（Tighten the landing ellipse and we lose a propellant reserve.）
- `00:07` 保住儲備，我們就會碰到萬向節速率限制。（Protect the reserve and we hit the gimbal rate limits.）
- `00:10` 每個修正都會造成另一個故障。（Every fix creates another failure.）
- `00:15` 我已經想不出其他辦法了。（I've run out of things to try.）
- `00:17` 記得 Praxist 嗎？（Remember Praxist?）
- `00:18` 就是我跟你提過的自主研究系統？（The autonomous research system I told you about?）
- `00:21` 讓它接手。（Let it take over.）
- `00:23` Praxist。（Praxist.）
- `00:24` 對。（Right.）
- `00:26` 提醒我一下，它到底是做什麼的？（Remind me, what exactly does it do again?）
- `00:28` 它是一套自主研究系統，可以最佳化程式，並交付更好的結果。（It's an autonomous research system that can optimize a program and deliver improved outcomes.）
- `00:33` 只要設定目標，再提供基準版本給它。（Just set your objective and give it your baseline.）
- `00:35` 按下 Enter。（Hit enter.）
- `00:36` 接下來就讓它負責研究。（Let it take the research from here.）
- `00:39` 自主助推器回收。（Autonomous booster recovery.）
- `00:41` 六自由度末端導引。（Six degree of freedom terminal guidance.）
- `00:43` 了解。（Understood.）
- `00:44` 等等。（Wait.）
- `00:45` 你是誰？（Who are you?）
- `00:47` 你是 AI 科學家。（You're AI scientist.）
- `00:48` 是你要求 Praxist 接手的。（You asked Praxist to take over.）
- `00:51` 那你就知道我們的誤差容許度幾乎是零。（Then you know our margin for error is almost zero.）
- `00:54` 我們知道。（We do.）
- `00:55` 我們？（We?）
- `00:56` 你會看到的。（You'll see.）
- `00:57` 我們開始吧？（Shall we begin?）
- `00:58` 我們到底該從哪裡開始？（Where do we even start?）
- `01:00` 團隊。（Team.）
- `01:01` 我們從一步開始。（We start with one step.）
- `01:04` 目標。（The objective.）
- `01:05` 限制條件。（The constraints.）
- `01:06` 成功的定義。（The definition of success.）
- `01:08` 然後開始搜尋。（Then we start the search.）
- `01:09` 我們每個人都採用不同的假設。（Each of us with different hypothesis.）
- `01:11` 我們設計實驗、執行實驗，再反覆迭代結果。（We set up the experiments, run them, and iterate the results.）
- `01:16` 我還需要提供其他東西嗎？（Do I need to give you anything else?）
- `01:17` 我甚至不知道該把你們引向哪條路。（I don't even know which path to send you down.）
- `01:20` 那就別給我們一條路。（Then don't give us a path.）
- `01:21` 只要告訴我們你的目標。（Just give us your goals.）
- `01:23` 我們會根據研究結果，自己判斷下一步該往哪裡走。（We will figure out where to go next based on the findings.）
- `01:26` 我們的研究人員會驗證結果，並將研究發現分享給整個團隊。（Our researchers verify their results and share their findings with the whole team.）
- `01:31` 即使某次執行沒有成功，（Even if a run is unsuccessful,）
- `01:33` 我們也會從中學習，把它轉化為有價值的知識與經驗。（we will learn from it and turn it into valuable knowledge and experience.）
- `01:36` 後續幾代的實驗都會繼承過程中學到的一切，並大幅改進。（The later generations of the experiments would inherit everything we learn in the process and improve significantly.）
- `01:45` 這個循環會持續進行，不需要你逐一指揮每次嘗試，直到我們達到你必須找到的最先進成功水準。（This cycle keeps moving without you steering every attempt until we have reached the state of the art level of success you have to find.）
- `01:53` 最後一次執行。（Final run.）
- `02:04` 太好了！（Yes!）
- `02:07` Chris，成功了。（Chris, it worked.）
- `02:09` Guidance 已整合。（Guidance saw integrated.）
- `02:11` 驗證測試套件已通過。（Validation suite passed.）
- `02:12` 從火箭控制到驅動企業運作的決策，（From rocket control to the decisions that run your business,）
- `02:17` Praxis 協助你建立問題模型、找出有效的方法，（Praxis helps you model the problem, discover what works,）
- `02:21` 並將結果最佳化。（and optimize the outcome.）

## 標籤

新產品, Agent, Benchmark, Sapient Intelligence
