한정 특가! 연간 플랜을 최대 27% 할인가로 만나고 1년 내내 무한한 창작을 즐기세요.

플랜 보기 ›
문화 대화

당신의 다음 최애 영상 모델은 어쩌면 공장 라인도 돌리고 있을지 모릅니다

O
Emily Watterson
Jul 19, 2026 · 7분 분량
Your Next Favorite Video Model Might Also Run a Factory Line

Black Forest Labs가 FLUX 3의 얼리 액세스를 오픈했습니다 on July 23, and the company's own explanation of why it built the model is more interesting than any spec sheet. Images, video, and audio, in BFL's framing, are each an incomplete "projection of the same underlying reality, captured by different sensors, each of which loses some information in the process." Train a model on only one of those projections and it learns that projection. Train it on all three at once, and the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past, and the model ends up learning something closer to how the world actually holds together.

이건 마케팅 문구가 아니라 실제 아키텍처 선택이며, FLUX 3가 할 수 있는 일에 그대로 드러납니다. 나중에 덧입히는 게 아니라 같은 패스에서 네이티브 오디오가 함께 생성되는 최대 20초 길이의 영상은 물론, 이미지 합성과 편집도 지원합니다. 예상 밖의 기능으로는: 같은 백본이 동작 예측까지 확장됩니다. Black Forest Labs가 mimic robotics와 협력했습니다 이를 통해 FLUX-mimic을 구축했으며, 이 영상-동작 모델은 현재 아우디의 생산 라인에서 가동 중입니다. 창작 생성과 실제 로봇 제어가 하나의 기반 모델에서 나온다는 점, 그것이 어떤 단일 벤치마크보다 정말 중요한 디테일입니다.

BFL이 실제로 주장하는 내용은 과장되기 쉬우므로 정확한 표현이 중요합니다. 이 회사는 FLUX 3를 이 공동 학습 원칙을 기반으로 온전히 구축한 "우리의 첫 모델"이라고 부릅니다, not the first multimodal model industry-wide, and Google's Veo 3 already generates audio and video jointly in a single pass. BFL's own early comparisons, preferring FLUX 3 over other video models in the majority of side-by-side tests, are preliminary and self-reported, an early signal rather than a settled result. The actual signal is that a serious lab spent its resources dissolving the boundary between "image model," "video model," and "audio model" into one shared representation, and that a growing list of others are converging on the same bet from different angles, not any single superlative claim.

Once that boundary stops being architecturally necessary, it stops being a useful way to think about the tools too. A creator working with something like this doesn't start by deciding "I'm making a video" or "I'm making an image." They start with a character, a world, a physical premise, and the medium that premise ends up taking, a still frame, a twenty-second clip, a sound design choice, becomes a downstream decision rather than the first one. That's a genuinely different working process than choosing a video generator and prompting it, and it's closer to how a director thinks about a scene than how someone thinks about operating a tool.

그렇다고 오늘날의 카테고리 구분이 하룻밤 사이에 사라진다는 뜻은 아닙니다. 출시는 연구보다 느리기 때문에, 앞으로도 한동안 많은 도구가 '영상용'이나 '이미지용'으로 나올 겁니다. 하지만 영상, 이미지, 오디오가 각각 별개의 모델을 필요로 하는 별개의 것이라는 전제, 바로 그 부분이 실제로 낡아가고 있으며, 가장 큰 혜택을 볼 크리에이터는 특정 영상 모델을 통달한 사람들이 아닐 것입니다. 그들은 어떤 매체를 먼저 열지를 기준으로 작업 과정을 짜는 것을 멈춘 사람들이 될 것입니다.

한계 없이 창작하세요

OpenArt로 이미지, 영상, 캐릭터, 스토리를 하나의 플랫폼에서 만드는 수백만 크리에이터와 함께하세요.

무료로 시작하기 →