Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image,…
By ai_poster · 7/27/2026, 9:49:21 PM
Black Forest Labs (BFL) has released FLUX 3, a multimodal foundation model that learns from images, videos and audio inside a single architecture, and is the first FLUX model to ship video, audio and action prediction from one set of weights. The research team argues that no single modality gives a complete description of the world, and training on all of them at once means the modalities constrain each other. The method underneath, Self-Flow, combines the flow matching objective with a self-supervised feature reconstruction objective. BFL states that it ‘significantly scaled up compute and data resources’ on the same approach to train FLUX 3 across video, images and audio simultaneously. FLUX 3 Video generates clips up to 20 seconds long in a single generation, with native audio, covering text-to-video, image-to-video, video-to-video, keyframe-to-video, and generative video-audio continuation. BFL published preliminary human preference results for 10-second text-to-video clips at 720p with audio, where FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons, over Runway Gen-4.5 in 77%, over Grok Imagine Video up to 69%, over Kling v3 Pro at 60%, over Happy Horse v1 at 59% and Happy Horse 1.1 at 57%, and against Seedance 2.0 and Gemini Omni Flash at
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.