Seus modelos favoritos — todos em um só lugar, com gerações ilimitadas.

Aproveite até 27% OFF por tempo limitado! ›
Novidades do OpenArt

Seu próximo modelo de vídeo favorito também pode estar rodando uma linha de fábrica

O
Emily Watterson
19 de jul. de 2026 · leitura de 7 minutos
Your Next Favorite Video Model Might Also Run a Factory Line

A Black Forest Labs abriu o acesso antecipado ao FLUX 3 on July 23, and the company's own explanation of why it built the model is more interesting than any spec sheet. Images, video, and audio, in BFL's framing, are each an incomplete "projection of the same underlying reality, captured by different sensors, each of which loses some information in the process." Train a model on only one of those projections and it learns that projection. Train it on all three at once, and the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past, and the model ends up learning something closer to how the world actually holds together.

Essa é uma escolha arquitetural de verdade, não linguagem de marketing, e aparece diretamente no que o FLUX 3 consegue fazer. Ele gera vídeo de até 20 segundos com áudio nativo integrado na mesma passagem, não adicionado depois, além de síntese e edição de imagem. O menos esperado: o mesmo backbone se estende à previsão de ações. A Black Forest Labs firmou parceria com a mimic robotics para criar o FLUX-mimic, um modelo de vídeo-ação que agora roda em linhas de produção da Audi. Geração criativa e controle físico de robôs, saindo do mesmo modelo de base, é o detalhe que realmente importa aqui, mais do que qualquer benchmark isolado.

O que a BFL está de fato afirmando é fácil de exagerar, então a versão precisa importa. A empresa chama o FLUX 3 de "nosso primeiro modelo" construído inteiramente sobre esse princípio de treinamento conjunto, not the first multimodal model industry-wide, and Google's Veo 3 already generates audio and video jointly in a single pass. BFL's own early comparisons, preferring FLUX 3 over other video models in the majority of side-by-side tests, are preliminary and self-reported, an early signal rather than a settled result. The actual signal is that a serious lab spent its resources dissolving the boundary between "image model," "video model," and "audio model" into one shared representation, and that a growing list of others are converging on the same bet from different angles, not any single superlative claim.

Once that boundary stops being architecturally necessary, it stops being a useful way to think about the tools too. A creator working with something like this doesn't start by deciding "I'm making a video" or "I'm making an image." They start with a character, a world, a physical premise, and the medium that premise ends up taking, a still frame, a twenty-second clip, a sound design choice, becomes a downstream decision rather than the first one. That's a genuinely different working process than choosing a video generator and prompting it, and it's closer to how a director thinks about a scene than how someone thinks about operating a tool.

Nada disso significa que os rótulos de categoria atuais vão sumir da noite para o dia. Muitas ferramentas vão continuar sendo lançadas como "a de vídeo" ou "a de imagem" por um bom tempo, porque a adoção é mais lenta que a pesquisa. Mas a premissa por trás disso, de que vídeo, imagem e áudio são coisas separadas que exigem modelos separados, é a parte que de fato está ficando ultrapassada, e os criadores que mais se beneficiam não serão os que dominaram um modelo de vídeo específico. Serão os que pararam de organizar seu processo em torno da pergunta de qual mídia abrir primeiro.

Crie sem limites

Junte-se a milhões de criadores que usam a OpenArt para gerar imagens, vídeos, personagens e histórias — tudo em uma só plataforma.

Comece grátis →