parallelquant
July 26, 2026 · MarkTechPost

Black Forest Labs releases FLUX 3, a multimodal flow model

Black Forest Labs released FLUX 3, a foundation model trained jointly on images, video, and audio in a single architecture. It's the first FLUX model to generate video, audio, and robot-action predictions from one shared set of weights.

Why it matters: This fits a broader push toward models that learn a unified representation of the world rather than separate ones per modality, echoing recent world-model work like Induction Labs' Photon-1 and the Open Dreamer reproduction of Dreamer 4. Folding robot-action prediction into the same weights as image/video/audio generation suggests foundation-model labs increasingly see robotics control as just another modality to learn jointly, not a separate specialty.

Related updates