Black Forest Labs Ships Flux 3: One Model for Images, Video, and Audio — Beats Runway in 77% of Comparisons
Black Forest Labs has launched Flux 3 in early access. It is a single foundation model trained simultaneously across images, video, and audio — not a suite of specialized models stapled together.
The underlying architecture is Self-Flow, BFL’s approach to aligning multimodal generation and understanding in a shared parameter space. The premise: images, video, and audio are projections of the same physical reality, and training a model across all three forces it to learn constraints that no single modality can teach — sound must match the physics of an impact, motion must obey mass, the future must follow from the past.
Capabilities
Flux 3 Video generates clips up to 20 seconds with native audio in a single pass. Supported modes:
- Text-to-video with audio
- Image-to-video (animation from a starting frame, or image as visual reference)
- Video-to-video transfer (transplant a character or scene element into a new context)
- Generative continuation from input video and audio
- Keyframe-to-video for controlled transitions
- Multilingual dialogue generation
- Agentic chaining of individual clips into multi-shot sequences
The model handles a wide style range — from camcorder-style footage to animation and cinematics — and BFL highlights strong typography generation in motion.
Early Benchmark Numbers
BFL published preliminary human-preference comparisons. All results involve 10-second text-to-video clips at 720p with audio. Results are vendor-reported and described as preliminary, with further improvements expected during early access:
| Competitor | Flux 3 Win Rate |
|---|---|
| Luma Ray 3.2 | 93% |
| Runway Gen-4.5 | 77% |
| Grok Imagine Video | 69% |
| Kling v3 Pro | 60% |
| Happy Horse 1.1 | 57% |
| Seedance 2.0 | 52% |
| Gemini Omni Flash | 52% |
The Seedance 2.0 comparison is the one to watch. ByteDance’s Seedance 2.0 has held both Arena Video top spots since its launch, with a 79-point ELO lead over Google Veo 3.1. Flux 3 shows only a narrow 52% win rate in early testing — better, but not a decisive lead. That comparison will sharpen when Arena adds Flux 3.
Architecture Note
Self-Flow versus standard flow matching: BFL reports lower generation error (Fréchet distance) per modality when jointly trained, and higher success rates on manipulation tasks through finetuning. The key claim is that modality constraints act as mutual supervision — learning video teaches the model things about images that image-only training cannot provide.
This is architecturally distinct from building separate text-to-image and text-to-video models and combining outputs at inference time.
Access
Flux 3 is now available in early access at bfl.ai. Pricing and API details have not been announced.