parallelquant
July 23, 2026 · The Decoder

Black Forest Labs' Flux 3 adds native audio to video generation

Black Forest Labs released Flux 3, a multimodal model trained on images, video, and audio that generates videos up to 20 seconds long with native synchronized sound, a first for the company. BFL's own benchmarks put it just ahead of Bytedance's Seedance 2.0, though independent evaluations aren't out yet. The company says it's already testing Flux 3 on robotics tasks as a step toward a broader world model.

Why it matters: Native audio closes a real feature gap with commercial leaders like Seedance and Veo, and BFL's stated ambition of building toward a world model echoes the same generation-as-simulation framing showing up elsewhere in robotics AI. Because the benchmark claims are self-reported, treat the 'ahead of Seedance' ranking as provisional until outside evals confirm it.

Related updates