Black Forest Labs opened early access to FLUX 3 on July 23, and the company's own explanation of why it built the model is more interesting than any spec sheet. Images, video, and audio, in BFL's framing, are each an incomplete "projection of the same underlying reality, captured by different sensors, each of which loses some information in the process." Train a model on only one of those projections and it learns that projection. Train it on all three at once, and the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past, and the model ends up learning something closer to how the world actually holds together.
That's a real architectural choice, not marketing language, and it shows up directly in what FLUX 3 can do. It generates video up to 20 seconds long with native audio baked in from the same pass, not layered on afterward, along with image synthesis and editing. Less expected: the same backbone extends into action prediction. Black Forest Labs partnered with mimic robotics to build FLUX-mimic, a video-action model now running on production lines at Audi. Creative generation and physical robot control, coming out of the same underlying model, is the detail that actually matters here, more than any single benchmark.
What BFL is actually claiming is easy to overstate, so the precise version matters. The company calls FLUX 3 "our first model" built entirely on this joint-training principle, not the first multimodal model industry-wide, and Google's Veo 3 already generates audio and video jointly in a single pass. BFL's own early comparisons, preferring FLUX 3 over other video models in the majority of side-by-side tests, are preliminary and self-reported, an early signal rather than a settled result. The actual signal is that a serious lab spent its resources dissolving the boundary between "image model," "video model," and "audio model" into one shared representation, and that a growing list of others are converging on the same bet from different angles, not any single superlative claim.
Once that boundary stops being architecturally necessary, it stops being a useful way to think about the tools too. A creator working with something like this doesn't start by deciding "I'm making a video" or "I'm making an image." They start with a character, a world, a physical premise, and the medium that premise ends up taking, a still frame, a twenty-second clip, a sound design choice, becomes a downstream decision rather than the first one. That's a genuinely different working process than choosing a video generator and prompting it, and it's closer to how a director thinks about a scene than how someone thinks about operating a tool.
None of this means today's category labels vanish overnight. Plenty of tools will keep shipping as "the video one" or "the image one" for a while yet, because rollout is slower than research. But the underlying premise, that video, image, and audio are separate things requiring separate models, is the part that's actually getting old, and the creators who benefit most won't be the ones who mastered a specific video model. They'll be the ones who stopped organizing their process around the question of which medium to open first.