Black Forest Labs has announced FLUX 3 — a new generation of multimodal models capable of seamlessly integrating image, video, and audio generation into a single system. Thanks to its unique Self-Flow architecture, the model demonstrates a deep understanding of the physical laws of the world and ensures high consistency between visual sequences and audio accompaniment.

What Happened
Black Forest Labs introduced FLUX 3, which transitions from discrete models to a unified multimodal architecture. The FLUX 3 Video model supports the creation of video clips up to 20 seconds long with native audio. According to benchmark results, the new technology outperforms solutions such as Luma Ray 3.2, Runway Gen-4.5, and Kling v3 Pro.
Context
The development of FLUX 3 is based on the concept of Real World Models, where the Self-Flow architecture serves as the foundation for creating visual intelligence. This marks an industry shift from using disparate tools (separate models for video and separate ones for sound) to complex systems capable of generating physically accurate content within a single pipeline.
Why It Matters for the Industry
For the industry, this transition implies the potential standardization of Self-Flow type architectures and a shift in focus from simple image-to-video combinations to full-fledged multimodal backbone models. In the long term, the development of such systems could provide a breakthrough in the field of Visual Intelligence, which is critical for creating accurate simulations and training robots in virtual environments.
Why It Matters for Users
Users can gain access to testing FLUX 3 Video through an early access form. This provides an opportunity to test technology that potentially changes the power dynamics in the video generation market and to begin prototyping new UX patterns based on the instantaneous generation of video along with sound from a text prompt.
What Is Not Yet Known / Limitations
At this time, there is no information regarding latency, usage costs, or API availability, making it impossible to assess the model's readiness for industrial production. The technology is currently primarily in the stage of research interest and early access.
Sources
Author
Look at AI, Editorial Team
