Zeitlich begrenztes Angebot! Sichere dir ein Jahr grenzenlose Kreativität mit Jahresplänen mit BIS ZU 27 % RABATT.

Plan ansehen ›
Kulturelle Gespräche

Dein nächstes Lieblings-Videomodell steuert vielleicht auch eine Fabrikstraße

O
Emily Watterson
Jul 19, 2026 · 7 Minuten Lesezeit
Your Next Favorite Video Model Might Also Run a Factory Line

Black Forest Labs hat den frühen Zugang zu FLUX 3 geöffnet on July 23, and the company's own explanation of why it built the model is more interesting than any spec sheet. Images, video, and audio, in BFL's framing, are each an incomplete "projection of the same underlying reality, captured by different sensors, each of which loses some information in the process." Train a model on only one of those projections and it learns that projection. Train it on all three at once, and the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past, and the model ends up learning something closer to how the world actually holds together.

Das ist eine echte architektonische Entscheidung, kein Marketing-Sprech, und sie zeigt sich direkt darin, was FLUX 3 leisten kann. Es generiert Videos von bis zu 20 Sekunden Länge mit nativem Audio, das im selben Durchgang entsteht und nicht nachträglich draufgelegt wird – dazu Bildsynthese und -bearbeitung. Weniger zu erwarten: dasselbe Grundgerüst erstreckt sich bis in die Aktionsvorhersage. Black Forest Labs kooperiert mit mimic robotics um FLUX-mimic zu bauen, ein Video-Action-Modell, das jetzt an den Produktionslinien von Audi läuft. Dass kreative Generierung und die Steuerung physischer Roboter aus demselben zugrunde liegenden Modell stammen – das ist hier das Detail, das wirklich zählt, mehr als jeder einzelne Benchmark.

Was BFL tatsächlich behauptet, lässt sich leicht überzeichnen – deshalb kommt es auf die genaue Version an. Das Unternehmen nennt FLUX 3 "unser erstes Modell", das vollständig auf diesem Joint-Training-Prinzip aufbaut, not the first multimodal model industry-wide, and Google's Veo 3 already generates audio and video jointly in a single pass. BFL's own early comparisons, preferring FLUX 3 over other video models in the majority of side-by-side tests, are preliminary and self-reported, an early signal rather than a settled result. The actual signal is that a serious lab spent its resources dissolving the boundary between "image model," "video model," and "audio model" into one shared representation, and that a growing list of others are converging on the same bet from different angles, not any single superlative claim.

Once that boundary stops being architecturally necessary, it stops being a useful way to think about the tools too. A creator working with something like this doesn't start by deciding "I'm making a video" or "I'm making an image." They start with a character, a world, a physical premise, and the medium that premise ends up taking, a still frame, a twenty-second clip, a sound design choice, becomes a downstream decision rather than the first one. That's a genuinely different working process than choosing a video generator and prompting it, and it's closer to how a director thinks about a scene than how someone thinks about operating a tool.

Nichts davon bedeutet, dass die heutigen Kategorielabels über Nacht verschwinden. Viele Tools werden noch eine Weile als „das für Videos“ oder „das für Bilder“ auf den Markt kommen, weil das Ausrollen langsamer läuft als die Forschung. Aber die zugrunde liegende Prämisse – dass Video, Bild und Audio getrennte Dinge sind, die getrennte Modelle brauchen – ist der Teil, der wirklich veraltet, und am meisten profitieren werden nicht die Creator, die ein bestimmtes Video-Modell gemeistert haben. Sie werden diejenigen sein, die aufgehört haben, ihren Prozess um die Frage herum zu organisieren, welches Medium sie zuerst öffnen.

Grenzenlos kreativ sein

Schließ dich Millionen von Creators an, die mit OpenArt Bilder, Videos, Charaktere und Storys erstellen – alles auf einer Plattform.

Kostenlos loslegen →