Black Forest Labs launches FLUX 3, a multimodal model generating video with synced audio
Black Forest Labs
Black Forest Labs released FLUX 3, a frontier model jointly trained across images, video, and audio in a unified architecture. It can generate clips up to 20 seconds long with synchronized audio (dialogue, sound effects, ambience) from a single prompt, and the company also unveiled an offshoot action model, Flux-mimic, for robotics with partner Mimic Robotics AG. Video and Action are in early access via API and select partners; image generation and an open-weight Dev tier are planned for later.
Why it matters
Marks a shift from separate image/video/audio pipelines to a single jointly-trained multimodal generative model, and extends the same architecture toward robotic control — a notable step toward unified "world models" for visual and physical intelligence.
Importance: 4/5
Frontier flagship release from Black Forest Labs (new multimodal architecture) confirmed by official post + 2 independent media outlets.