# Black Forest Labs 發布 FLUX 3 支援音訊影片生成

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：Black Forest Labs (@bfl_ai) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥🔥🔥🔥 · 日期：2026-07-24

> 原始來源：https://x.com/bfl_ai/status/2080308988961554582

## 證據與延伸閱讀

- [FLUX 3 是大型語言模型與視覺基礎模型](https://bfl.ai/blog/flux-3)
- [訓練中加入動作預測後恢復品質，證明實體AI共用backbone](https://bfl.ai/blog/flux-3-mimic)

## 中文摘要

Black Forest Labs 發布 FLUX 3 支援音訊影片生成。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/707ca0786385b153.jpg)
> 章魚眼睛的高畫質特寫照片，清晰展現其金黃色的瞳孔與獨特的皮膚紋理。

Black Forest Labs 於 2026 年 7 月 23 日正式發表全新多模態基礎模型 FLUX 3，採用單一統一架構同步學習影像、影片與音訊，並將應用範圍延伸至真實世界的物理人工智慧（Physical AI）與機器人動作預測。FLUX 3 Video 目前已開放早期存取（Early Access），使用者可透過 [FLUX 3 影片早期存取申請](https://bfl.ai/models/flux-3?utm_source=x&utm_medium=social&utm_campaign=flux3_launch) 提出申請，官方技術網誌亦同步公開詳細架構與規劃，詳見 [FLUX 3 官方技術說明](https://bfl.ai/blog/flux-3)。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/ad0cfce38f98e016.jpg)
> 簡約插畫風格的夜景圖，描繪海面上的一座燈塔向右側投射出明亮的黃色光束，一輪弦月高掛於深藍色夜空中。

**核心功能與模態整合**
- 影像、影片與音訊聯合學習：透過名為 Self-Flow 的自研方法，打破各模態獨立訓練的限制，讓使用者或模型理解物件如何聚合、物體如何移動及事件如何發出聲響，進而建立真實世界的物理動態表示。
- FLUX 3 Video：支援最長 20 秒、720p 解析度的影片與原生音訊聯合生成，涵蓋文字轉影片、圖片轉影片、影片轉影片、關鍵影格過渡及多語種對話等功能，並在早期評測中超越 Grok Imagine Video、Kling V3 Pro、Runway Gen-4.5 等多款同級模型。
- FLUX 3 Image：具備高度靈活的風格多樣性與精準的跨語言文字渲染能力，後續將於幾週內陸續開放早期存取。
- 開放權重與分階段推出：官方預計在接下來幾個月內，依序推出原生音訊影片生成、合作夥伴動作預測、影像合成編輯 API，以及多模態主幹網路（Backbone）的開放權重版本（FLUX 3 Dev）。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/11db54727bf38eb6.jpg)
> 淺黃色背景襯托的產品照，一張帶有亮藍色管狀金屬結構與紅色絨毛坐墊的現代設計感椅。

**機器人動作預測與 FLUX-mimic 實戰**
- 結合視覺與動作：Black Forest Labs 與機械手臂公司 mimic robotics 合作開發 [FLUX-mimic 專文說明](https://bfl.ai/blog/flux-3-mimic)，將 FLUX 3 影片主幹網路結合機器人深度學習，解碼出支援精細抓取與生產部署的動作模型。
- 奧迪（Audi）產線驗證：該技術目前已實際部署於奧迪的車輛與物流產線中進行測試。
- 跨領域模型架構：透過將動作預測納入訓練課程，模型在經歷短暫的微調適應期後，成功在單一主幹網路上同時支撐內容創作與物理人工智慧兩大應用家族。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/139f6268c8498db0.jpg)
> 一輛深色轎車在空曠的停車場地面上迴轉並摩擦出環狀胎痕。

**推出時程與未來規劃**
- 循序漸進的發布計畫：所有核心能力皆會經過嚴格的安全測試與早期存取階段，以確保系統穩定推出。
- 下一代模型研發：團隊目前已著手研發下一代模型，目標是將感知、動作與語言預測更進一步整合至同一個統一模型中。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/924de0bbb99576d1.png)
> 「FLUX 3」字樣醒目置中，周圍環繞著多個代表圖像、影片、動作與音訊等多媒體生成主題的預覽畫面與標籤。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/90e9008eaec939a5.png)
> 一隻帶有斑點的灰色馬匹在草原上奔馳，背景是陰暗的天空，且空中隱約有魚類漂浮。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/c94857f25165eccc.png)
> 一棵大型樹木的茂密樹冠與枝幹剪影，整體以深綠色調與黑色的高對比風格呈現。

<video src="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/1784826211988-gvj8wbn0.mp4" poster="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/ce011a1638213fb3.jpg" controls playsinline preload="metadata" style="max-width:100%;height:auto;display:block;margin:1rem 0"></video>
> 各種高品質 AIGC 影片生成與視覺效果的宣傳短片，最後預告 FLUX 3 即將推出。

<video src="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/1784826282978-yan081a1.mp4" poster="https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/8d8062148bd19867.jpg" controls playsinline preload="metadata" style="max-width:100%;height:auto;display:block;margin:1rem 0"></video>
> 影片介紹由 Black Forest Labs 與 mimic 合作開發的全新多模態視覺動作模型 FLUX-mimic，展示其在機器人自動化、軟體操作與工業製造（如 Audi 產線應用）上的突破性進展。

## 媒體內容

**影片介紹由 Black Forest Labs 與 mimic 合作開發的全新多模態視覺動作模型 FLUX-mimic，展示其在機器人自動化、軟體操作與工業製造（如 Audi 產線應用）上的突破性進展。**

**影片中的 Prompt 與操作**

操作步驟：

1. （00:00）機器人手臂進行管線與零件組裝
2. （00:24）機器人自主分類收納黑色管狀零件
3. （00:58）雙機器人手臂進行物件抓取與互動演示
4. （02:04）FLUX-mimic 機器人進行車門零組件與膠條組裝
5. （02:31）雙機器人手臂進行 Audi 車門飾條/膠條自動化安裝

**逐字稿**

- `00:05` Today we've seen a model that is restrained in a multimodal way on images, video and audio
- `00:10` that actually manipulate robots and interact with the physical world.
- `00:17` The only way to do that is to give a robot a certain intelligence layer
- `00:22` that intuitively understands the world and how to act within the world.
- `00:30` Our model has been trained to predict the dynamics of the world from day one.
- `00:35` Flux 3 really breaks this asymmetry and makes the dynamics of the world a first-class citizen.
- `00:43` Mimic is basically an ideal partner to prove how powerful this new multimodal architectural approach actually is.
- `00:49` And this is why we chose to partner with them.
- `00:53` Last year at Mimic, we introduced Mimic Video.
- `00:56` What is the difference between VLAs and Video Action Models?
- `01:00` Both of them are about training robots from large-scale behavior data.
- `01:04` Their pre-training is much better suited for robotics.
- `01:07` Flux Mimic is a next-generation video action model.
- `01:10` It's built on top of the Flux backbone, but trained on robot data to predict robot actions.
- `01:19` So what we're doing now is fundamentally different from conventional automation.
- `01:23` With this, we can suddenly solve a new level of unstructured tasks that were previously considered pretty much impossible for automation to do.
- `01:31` So that means anything with changing positions, different orientations of objects, handling unpredictable soft bodies, cables, and so on.
- `01:40` We have a lot of exciting partnerships going on.
- `01:43` And one of the highlights is our partnership with Audi.
- `01:46` Fundamentally, Audi wants an automation solution that is flexible, reliable, and cost-effective.
- `01:54` And for us, this means that the main goal is really to reduce integration efforts.
- `02:06` Flux Mimic
- `02:07` We have been testing and deploying Flux Mimic.
- `02:10` We've seen these robots solve complex soft body manipulations that would have been simply impossible for conventional automation.
- `02:18` This can have a huge impact in assisting our employees, increasing efficiency, and scaling flexible automation across our processes in production and logistics.
- `02:32` For us, partnering with pioneer companies as Mimic and Black Forest Lab is essential for pushing the frontiers of physical AI
- `02:42` and validating these innovations in real-world production environments.
- `02:52` Our model was already trained at scale for dynamic understanding.
- `02:57` And we also integrate physically grounded data, which then allows us to adapt the model much quicker to new application areas requiring less data collection there.
- `03:09` We've seen that with the latest batch of models, it can take us as low as 30 minutes of data to get a task to work really well, depending on the difficulty of the task.
- `03:17` In the past, we always saw that we needed at least 30 hours of data for something to work really well.
- `03:23` With this built-in dynamics understanding, we have seen that the model is able to recover from failure.
- `03:37` That's the really exciting part where, for the first time now, our models are moving out of the labs and out of the virtual world, really into the real world.

## 標籤

新產品, VLM, 硬體, Black Forest Labs
