Black Forest Labs released Flux 3 on July 2026, a multimodal foundation model generating 20-second video with native audio. BFL's internal tests claim 52% preference over ByteDance's Seedance 2.0, though independent benchmarks are pending.
Key facts
- Flux 3 generates video with native audio up to 20 seconds
- BFL claims 93% preference over Luma Ray 3.2
- 52% preference over Seedance 2.0 in internal tests
- Flux-mimic tested at Audi via Mimic Robotics partnership
- Flux 3 Image early access planned within weeks
Black Forest Labs (BFL) has released Flux 3, a multimodal foundation model that learns from images, video, and audio simultaneously, according to The Decoder. For the first time in BFL's product line, Flux 3 generates videos with native audio up to 20 seconds long, supporting text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven clip chaining for longer sequences. BFL describes the model as a step toward "real-world visual intelligence," aiming to build world models that perceive, predict, and act across physical and digital environments.
Key Takeaways
- Black Forest Labs released Flux 3, a multimodal model generating 20-second video with native audio.
- Internal tests claim 52% preference over Seedance 2.0, but independent results are pending.
Internal benchmarks show narrow lead over Seedance 2.0
In early evaluations using 10-second clips at 720p, BFL reports that Flux 3 was preferred over Luma Ray 3.2 in 93 percent of comparisons, over Runway Gen-4.5 in 77 percent, and over Grok Imagine Video in 69 percent. The margins narrow against stronger competitors: Flux 3 was preferred over Kling v3 Pro 60 percent of the time, over Happy Horse v1 at 59 percent, over Happy Horse 1.1 at 57 percent, and over both Seedance 2.0 — ByteDance's video generation model first observed on Jimeng AI in February 2026 — and Gemini Omni Flash at 52 percent each. BFL explicitly states these results are preliminary, and no independent tests are available yet. Matching Seedance 2.0, which has already reached Hollywood production pipelines, would place Flux 3 among the top video models.
Robotics and the path to world models
Flux 3 is based on Self-Flow, BFL's proprietary approach, and the company expects it to improve image generation for complex prompts and accurate multilingual text rendering. BFL plans to release Flux 3 Image in early access within weeks. Separately, the company worked with Mimic Robotics to develop Flux-mimic, a video-action model now being tested on production tasks at Audi. This robotics application underscores BFL's stated ambition to build a foundation model that can act in physical environments, not just generate media.

What to watch
Watch for independent third-party benchmarks comparing Flux 3 to Seedance 2.0 and Gemini Omni Flash, expected within weeks. Also track the Flux 3 Image early access release and any disclosed Audi production results for Flux-mimic.

Source: the-decoder.com









