Hourly ·
Black Forest Labs Unveils FLUX 3: Multimodal Model Spanning Video, Image, Audio, and Robotics
FLUX 3 outperforms Grok, Kling, Runway, and Luma in early video benchmarks. The same backbone powers FLUX-mimic, a robotics model already deployed at Audi factories.
Black Forest Labs has released FLUX 3, a multimodal foundation model that jointly learns from images, video, and audio within a single unified architecture. The model is now available in Early Access.
FLUX 3 can generate videos up to 20 seconds long with native audio, supporting text-to-video, image-to-video, video-to-video, keyframe-to-video, and multilingual dialogue. Agentic chaining allows assembling individual clips into multi-shot sequences lasting several minutes.
In preliminary evaluations, FLUX 3 was preferred over Grok Imagine Video in 69% of comparisons, Kling v3 Pro in 60%, and Runway Gen-4.5 in 77%. It beat Luma Ray 3.2 in 93% of matchups. The model showed particular strength in capturing human facial expressions, synchronizing audio with physical events, and producing multilingual output.
The same model backbone also powers FLUX-mimic, a video-action model built with mimic robotics and deployed at Audi. FLUX-mimic handles soft-body manipulation tasks -- seals, cables, and flexible materials -- that conventional automation has never been able to touch, running at 101ms reaction time on a single NVIDIA RTX 5090.
FLUX 3 is built on BFL's Self-Flow approach for aligning multimodal generation and understanding. The company plans to roll out FLUX 3 Video, FLUX 3 Image, and an open-weight FLUX 3 Dev over the coming weeks.
Sources: Black Forest Labs -- Hacker News Discussion
黑森林实验室推出FLUX 3:跨视频、图像、音频和机器人技术的多模态模型
FLUX在早期视频基准测试中优于Grok、Kling、Runway和Luma。相同的底层架构支撑着[K 已经在奥迪工厂部署的FLUX-mimic机器人模型。
← 小时版 小时 · 2026-07-24 12:00 UTC 黑森林实验室发布FLUX 3:跨视频、图像、[K 音频和机器人多模态基础模型 FLUX 3在早期视频基准测试中超过Grok、Kling、Runwa[5D[K Runway和Luma。相同的主干结构驱动了FLUX-mimic,一个已在奥迪工厂部署的机器人模[K 型。黑森林实验室发布了FLUX 3,这是一个跨图像
More Hourlies Stories
Content on Anagnorisis is summarized, paraphrased, and editorialized from publicly available sources for length and clarity. Original sources are linked where available. All trademarks belong to their respective owners.
