¡Oferta por tiempo limitado! Desbloquea un año de creatividad ilimitada con planes anuales con HASTA UN 27 % DE DESCUENTO.

Ver plan ›
Conversaciones culturales

Puede que tu próximo modelo de vídeo favorito también gestione una línea de producción

O
Emily Watterson
Jul 19, 2026 · 7 minutos de lectura
Your Next Favorite Video Model Might Also Run a Factory Line

Black Forest Labs abrió el acceso anticipado a FLUX 3 on July 23, and the company's own explanation of why it built the model is more interesting than any spec sheet. Images, video, and audio, in BFL's framing, are each an incomplete "projection of the same underlying reality, captured by different sensors, each of which loses some information in the process." Train a model on only one of those projections and it learns that projection. Train it on all three at once, and the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past, and the model ends up learning something closer to how the world actually holds together.

Es una decisión arquitectónica real, no lenguaje de marketing, y se nota directamente en lo que FLUX 3 puede hacer. Genera vídeo de hasta 20 segundos con audio nativo integrado en la misma pasada, no añadido después, junto con síntesis y edición de imágenes. Menos esperado: la misma base se extiende a la predicción de acciones. Black Forest Labs se ha asociado con mimic robotics para construir FLUX-mimic, un modelo de vídeo-acción que ya funciona en las líneas de producción de Audi. Que la generación creativa y el control de robots físicos surjan del mismo modelo subyacente es el detalle que realmente importa aquí, más que cualquier benchmark concreto.

Lo que BFL afirma realmente es fácil de exagerar, así que la versión precisa importa. La empresa llama a FLUX 3 «nuestro primer modelo» construido íntegramente sobre este principio de entrenamiento conjunto, not the first multimodal model industry-wide, and Google's Veo 3 already generates audio and video jointly in a single pass. BFL's own early comparisons, preferring FLUX 3 over other video models in the majority of side-by-side tests, are preliminary and self-reported, an early signal rather than a settled result. The actual signal is that a serious lab spent its resources dissolving the boundary between "image model," "video model," and "audio model" into one shared representation, and that a growing list of others are converging on the same bet from different angles, not any single superlative claim.

Once that boundary stops being architecturally necessary, it stops being a useful way to think about the tools too. A creator working with something like this doesn't start by deciding "I'm making a video" or "I'm making an image." They start with a character, a world, a physical premise, and the medium that premise ends up taking, a still frame, a twenty-second clip, a sound design choice, becomes a downstream decision rather than the first one. That's a genuinely different working process than choosing a video generator and prompting it, and it's closer to how a director thinks about a scene than how someone thinks about operating a tool.

Nada de esto significa que las etiquetas de categoría de hoy vayan a desaparecer de la noche a la mañana. Muchas herramientas seguirán presentándose como «la de vídeo» o «la de imagen» durante un tiempo, porque el despliegue va más lento que la investigación. Pero la premisa de fondo, que vídeo, imagen y audio son cosas separadas que requieren modelos separados, es la parte que realmente está quedándose obsoleta, y los creadores que más se beneficiarán no serán los que dominaron un modelo de vídeo concreto. Serán quienes dejaron de organizar su proceso en torno a la pregunta de qué medio abrir primero.

Crea sin límites

Únete a millones de creadores que usan OpenArt para generar imágenes, vídeos, personajes e historias, todo en una sola plataforma.

Empieza gratis →