GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Black Forest Labs Ships Flux 3: One Model for Images, Video, and Audio — Beats Runway in 77% of Comparisons

Black Forest Labs has launched Flux 3 in early access. It is a single foundation model trained simultaneously across images, video, and audio — not a suite of specialized models stapled together.

The underlying architecture is Self-Flow, BFL’s approach to aligning multimodal generation and understanding in a shared parameter space. The premise: images, video, and audio are projections of the same physical reality, and training a model across all three forces it to learn constraints that no single modality can teach — sound must match the physics of an impact, motion must obey mass, the future must follow from the past.

Capabilities

Flux 3 Video generates clips up to 20 seconds with native audio in a single pass. Supported modes:

  • Text-to-video with audio
  • Image-to-video (animation from a starting frame, or image as visual reference)
  • Video-to-video transfer (transplant a character or scene element into a new context)
  • Generative continuation from input video and audio
  • Keyframe-to-video for controlled transitions
  • Multilingual dialogue generation
  • Agentic chaining of individual clips into multi-shot sequences

The model handles a wide style range — from camcorder-style footage to animation and cinematics — and BFL highlights strong typography generation in motion.

Early Benchmark Numbers

BFL published preliminary human-preference comparisons. All results involve 10-second text-to-video clips at 720p with audio. Results are vendor-reported and described as preliminary, with further improvements expected during early access:

CompetitorFlux 3 Win Rate
Luma Ray 3.293%
Runway Gen-4.577%
Grok Imagine Video69%
Kling v3 Pro60%
Happy Horse 1.157%
Seedance 2.052%
Gemini Omni Flash52%

The Seedance 2.0 comparison is the one to watch. ByteDance’s Seedance 2.0 has held both Arena Video top spots since its launch, with a 79-point ELO lead over Google Veo 3.1. Flux 3 shows only a narrow 52% win rate in early testing — better, but not a decisive lead. That comparison will sharpen when Arena adds Flux 3.

Architecture Note

Self-Flow versus standard flow matching: BFL reports lower generation error (Fréchet distance) per modality when jointly trained, and higher success rates on manipulation tasks through finetuning. The key claim is that modality constraints act as mutual supervision — learning video teaches the model things about images that image-only training cannot provide.

This is architecturally distinct from building separate text-to-image and text-to-video models and combining outputs at inference time.

Access

Flux 3 is now available in early access at bfl.ai. Pricing and API details have not been announced.