# Mimic 推出 FLUX-mimic 結合影片與機器人資料

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：mimic (@mimicrobotics) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥🔥🔥 · 日期：2026-07-24

> 原始來源：https://x.com/mimicrobotics/status/2080307032746336367

## 證據與延伸閱讀

- [Mimic 推出 FLUX-mimic 結合影片與機器人](https://www.mimicrobotics.com/blog/introducing-flux-mimic)

## 中文摘要

Mimic 推出 FLUX-mimic 結合影片與機器人資料。

<video src="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/1784871856903-a5a2kk34.mp4" poster="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/07433858efa8ea88.jpg" controls playsinline preload="metadata" style="max-width:100%;height:auto;display:block;margin:1rem 0"></video>
> Black Forest Labs 與 mimic 合作展示新型多模態影片動作模型 FLUX-mimic，並於 Audi 生產實驗室驗證其在複雜非結構化任務中的應用。

**核心架構與數據優勢**
- Mimic 與 Black Forest Labs 合作推出 FLUX-mimic，將影片動作模型（VAM）架構應用於 FLUX 3 視訊生成模型中。
- VAM 架構透過結合影片與世界模型來預測機器人動作，利用大規模影片預訓練所學得的物理互動與人類行為先驗，能以高出傳統視覺語言動作模型（VLA）最高 10 倍的資料效率進行學習。
- 資料金字塔的最底層為無機器人動作的大規模人類影片，中間層為穿戴式系統錄製的示範資料，最頂層則是稀缺的機器人遙控操作與部署資料。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/a52e7251cb08bb17.png)
> 數據收集方式從機器人遠端操控（Teleoperation）、穿戴式裝備數據（Wearables）過渡至人類通用影片（Human video），可擴展性（Scalability）從 10¹ 小時顯著提升至 10⁶ 小時以上，但硬體對齊度（Hardware Alignment）隨之降低。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/cdb3d8f645cf6d66.png)
> 在 mimic-video 實驗中，輸入專家影片作為影片骨幹輸入可使預訓練與微調影片模型的成功率均達到 100%，而輸入預測影片時微調模型成功率為 49%。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/11cdd482fb4f37b3.png)
> The FLUX-mimic model 架構圖展示了從語言指令與影片觀察出發，透過 FLUX Video Stream 與動作解碼器來產生機器人動作的流程。

**效能表現與實際應用**
- 根據 Mimic 的評估，FLUX-mimic 在一項軟質物件組裝測試中，不需針對單一任務進行微調或後訓練，就能開箱即用地達到 95% 的成功率。
- 此模型已實際進駐工廠生產線，並與汽車製造商如 Audi 合作，測試並部署於複雜的多步驟操作與軟質材料（如密封件與電纜）的組裝任務中。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/6bf2d328c0f0fdcd.png)
> FLUX-mimic 在軟體裝組任務（Soft body kitting task）中展現出高達 95% 的自主成功率，顯著超越 Flow matching 與 π0.5 等比較模型。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/94f3267f8a407878.png)
> 白色機械手呈現勝利手勢，下方連接帶有「M1.0」字樣的圓柱形底座。

**邊緣端即時推論與未來展望**
- 雖然視訊模型通常運算需求沈重，但 FLUX-mimic 透過部分視訊去噪、模型量化、即時分塊（RTC）及帶有快取的調適性動作去噪等最佳化設計，能在單張 NVIDIA RTX 5090 GPU 上於機器人邊緣端即時執行。
- 透過 Mimic 的 mimic-ipc 軟體中介軟體，感測器、模型與致動器之間的資料傳輸得以維持極低延遲，確保模型能對世界做出即時反應。
- 為了讓機器人能透過動態草圖、目標圖像或影片示範來接收精確上下文，團隊正積極規劃讓 FLUX-mimic 支援任意形式的提示介面，進一步降低操作門檻。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/8eea138671d78d6a.png)
> FLUX-mimic 在累計套用各項優化後，推論延遲從 Baseline 的 456 ms 顯著降至 101 ms，達到 4.51 倍的加速效果。

## 媒體內容

**Black Forest Labs 與 mimic 合作展示新型多模態影片動作模型 FLUX-mimic，並於 Audi 生產實驗室驗證其在複雜非結構化任務中的應用。**

**影片中的 Prompt 與操作**

操作步驟：

1. （00:00）機器人手臂進行零件夾取與安裝測試
2. （00:58）雙機器人手臂自主執行管線分裝與搬運
3. （01:53）機器人執行預測畫面與真實世界畫面的對比展示
4. （02:31）雙機器人手臂執行車門膠條安裝
5. 系統介面顯示多視角攝影機監控畫面與程式執行記錄

**逐字稿**

- `00:05` 今天我們看到了一個在圖像、影片和音訊方面以多模態方式受到限制的模型（Today we've seen a model that is restrained in a multimodal way on images, video and audio）
- `00:10` 它實際上能夠操縱機器人並與實體世界互動。（that actually manipulate robots and interact with the physical world.）
- `00:17` 要做到這一點的唯一方法，就是賦予機器人一定的智慧層（The only way to do that is to give a robot a certain intelligence layer）
- `00:22` 能夠直覺地理解世界以及如何在世界中採取行動。（that intuitively understands the world and how to act within the world.）
- `00:30` 我們的模型從第一天起就經過訓練，能夠預測世界的動態。（Our model has been trained to predict the dynamics of the world from day one.）
- `00:35` Flux 3 真正打破了這種不對稱性，讓世界的動態成為了一等公民。（Flux 3 really breaks this asymmetry and makes the dynamics of the world a first-class citizen.）
- `00:43` Mimic 基本上是證明這種新型多模態架構方法有多麼強大的理想夥伴。（Mimic is basically an ideal partner to prove how powerful this new multimodal architectural approach actually is.）
- `00:49` 這就是為什麼我們選擇與他們合作。（And this is why we chose to partner with them.）
- `00:53` 去年在 Mimic，我們推出了 Mimic Video。（Last year at Mimic, we introduced Mimic Video.）
- `00:56` VLA 和影片動作模型（Video Action Models）之間有什麼區別？（What is the difference between VLA's and Video Action Models?）
- `01:00` 這兩者都涉及從大規模行為資料中訓練機器人。（Both of them are about training robots from large-scale behavior data.）
- `01:04` 它們的預訓練非常適合機器人技術。（Their pre-training is much better suited for robotics.）
- `01:07` Flux Mimic 是一個新一代的影片動作模型。（Flux Mimic is a next-generation video action model.）
- `01:10` 它建立在 flux 骨幹網路之上，但經過機器人資料的訓練來預測機器人的動作。（It's built on top of the flux backbone, but trained on robot data to predict robot actions.）
- `01:19` 所以我們現在所做的事情與傳統自動化有著根本上的不同。（So what we're doing now is fundamentally different from conventional automation.）
- `01:23` 有了這個，我們突然能夠解決全新層級的非結構化任務，這些任務以前被認為幾乎是自動化無法完成的。（With this, we can suddenly solve a new level of unstructured tasks that were previously considered pretty much impossible for automation to do.）
- `01:31` 這意味著任何位置改變、物件的不同方向、處理不可預測的軟體、電纜等等。（So that means anything with changing positions, different orientations of objects, handling unpredictable soft bodies, cables, and so on.）
- `01:40` 我們有許多令人興奮的合作夥伴關係正在進行中。（We have a lot of exciting partnerships going on.）
- `01:43` 其中一個亮點就是我們與 Audi 的合作。（And one of the highlights is our partnership with Audi.）
- `01:46` 從根本上來說，Audi 想要一個靈活、可靠且具成本效益的自動化解決方案。（Fundamentally, Audi wants an automation solution that is flexible, reliable, and cost-effective.）
- `01:54` 對我們而言，這意味著主要目標確實是減少整合心力。（And for us, this means that the main goal is really to reduce integration efforts.）
- `02:06` 透過與 Mimic 合作，我們一直在測試和部署 Flux Mimic。（Partnering with Mimic, we've been testing and deploying Flux Mimic.）
- `02:10` 我們已經看到這些機器人解決了複雜的軟體操作，這對傳統自動化來說簡直是不可能的。（We've seen these robots solve complex soft body manipulations that would have been simply impossible for conventional automation.）
- `02:19` 這可以對協助我們的員工、提高效率以及在生產和物流的各個流程中擴展靈活自動化產生巨大的影響。（This can have a huge impact in assisting our employees, increasing efficiency, and scaling flexible automation across our processes in production and logistics.）
- `02:31` 對我們來說，與 Mimic 和 Black Forest Lab 這樣的先驅公司合作，對於推動實體 AI 的極限以及在真實世界的生產環境中驗證這些創新至關重要。（For us, partnering with pioneer companies as Mimic and Black Forest Lab is essential for pushing the frontiers of physical AI and validating these innovations in real-world production environments.）
- `02:52` 我們的模型已經進行過大規模的動態理解訓練，我們也整合了實體基礎的資料，這隨後讓我們能夠更快地將模型調整到需要較少資料收集的新應用領域。（Our model was already trained at scale for dynamic understanding, and we also integrate physically grounded data, which then allows us to adapt the model much quicker to new application areas requiring less data collection there.）
- `03:09` 我們已經看到，透過最新一批的模型，根據任務的難度，我們只需要短短 30 分鐘的資料就能讓任務運作得非常好。（We've seen that with the latest batch of models, it can take us as low as 30 minutes of data to get a task to work really well, depending on the difficulty of the task.）
- `03:17` 過去，我們總是發現我們需要至少 30 小時的資料才能讓某些東西運作得很好。（In the past, we always saw that we needed at least 30 hours of data for something to work really well.）
- `03:23` 有了這種內建的動態理解，我們發現模型能夠從失敗中恢復。（With this built-in dynamics understanding, we have seen that the model is able to recover from failure.）
- `03:37` 這真的是最令人興奮的部分，我們的模型現在首次走出了實驗室和虛擬世界，真正進入了真實世界。（That's the really exciting part, where for the first time now, our models are moving out of the labs and out of the virtual world, really into the real world.）

**數據收集方式從機器人遠端操控（Teleoperation）、穿戴式裝備數據（Wearables）過渡至人類通用影片（Human video），可擴展性（Scalability）從 10¹ 小時顯著提升至 10⁶ 小時以上，但硬體對齊度（Hardware Alignment）隨之降低。**

**數據表**

| 項目 | X | Y |
| --- | --- | --- |
| Teleoperation | 10¹ h | 高硬體對齊 |
| Kinematically matched wearables | 10³ h | 中高硬體對齊 |
| Comfort optimized wearables | 約 10⁴ h | 中低硬體對齊 |
| Ego-centric and general video | 10⁶ h | 低硬體對齊 |

**在 mimic-video 實驗中，輸入專家影片作為影片骨幹輸入可使預訓練與微調影片模型的成功率均達到 100%，而輸入預測影片時微調模型成功率為 49%。**

**數據表**

|   | pretrained video model | finetuned video model |
| --- | --- | --- |
| predicted video | 3% | 49% |
| expert video | 100% | 100% |

**FLUX-mimic 在軟體裝組任務（Soft body kitting task）中展現出高達 95% 的自主成功率，顯著超越 Flow matching 與 π0.5 等比較模型。**

**數據表**

|   | 自主成功率 |
| --- | --- |
| FLUX-mimic | 95% |
| Flow matching baseline ST-FT | 70% |
| FLUX-mimic frozen backbone | 65% |
| π0.5 (mimic data mix) | 55% |
| π0.5 (mimic data mix) frozen backbone | 0% |

**FLUX-mimic 在累計套用各項優化後，推論延遲從 Baseline 的 456 ms 顯著降至 101 ms，達到 4.51 倍的加速效果。**

**數據表**

| 項目 | 數值 |
| --- | --- |
| Baseline | 456 ms (1.00x) |
| + Partial video denoising | 364 ms (1.25x) |
| + Compilation | 186 ms (2.45x) |
| + Quantization | 175 ms (2.60x) |
| + Adaptive action denoising with cache | 101 ms (4.51x) |

## 標籤

新產品, 硬體, Robot, Mimic
