AI news story

Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs

Black Forest Labs has released Flux 3, a multimodal foundation model that learns from images, video, and audio and can…

  • Generative
  • Source: The Decoder
  • Published: 2026-07-23

Editor's take

Black Forest Labs' Flux 3 now generates video with integrated audio up to 20 seconds, a notable advancement in multimodal AI. This development directly challenges established players like Seedance, which has previously led in synthesizing video with sound. The ability to create synchronized audio and visual content within a single generative model is crucial for more realistic and immersive media production.

The implications extend beyond entertainment, impacting areas like synthetic data generation for robotics and autonomous systems, where accurate audio-visual correlation is vital. Flux 3's reported performance, exceeding Seedance 2.0 in BFL's internal benchmarks, suggests a tightening competitive landscape and a potential shift in the foundation model hierarchy.

Future developments to monitor include independent verification of Flux 3's performance against Seedance and other emerging multimodal models like Google's Lumiere. The duration limit of 20 seconds is also a key factor to watch; extending this capability will be critical for practical application in longer-form content creation.