최근 AI 에이전트 파이프라인으로 촬영한 2분짜리 브랜드 필름은 기획안부터 완성본까지 3일이 걸렸습니다. 전통적인 방식이었다면 대략 두 달이 걸렸을 작업이죠. invideo가 직접 문서화한 제작 사례에 따르면. A ninety-second horror short pulled around four hundred separate video generations before it was done. Numbers like that used to be the whole story: AI got fast enough to make an entire scene, so why would you need a person in the room. The more interesting number is the other one buried in that same case study: more than 40 percent of the finished shots in one documented project weren't a single generation at all. They were stitched together from the strongest few seconds of several different attempts, a practice invideo's own team describes bluntly: "Prompt, eight tries, Frankenstein the keepers."
그 디테일이 속도 관련 주장들보다 실제로 무슨 일이 벌어지는지를 더 잘 말해 줍니다. 명령 한 번으로 전체 장면을 렌더링할 수 있는 모델이라도 결정을 내리는 사람의 필요를 없애지는 못합니다. 다만 그 모든 결정을 더 앞당기고, 더 빠르게 실행할 수 있게 해 줄 뿐입니다. The workflow behind these AI-made shorts isn't one person typing a single magic sentence and getting a finished film back. It's a full crew of separate roles, a producer role holding the script and characters, a storyboard role visualizing shots before anything gets generated, a cinematography role taking direction like "hold the shot longer" or "track the actor through the doorway," a costume role, a production design role, each one scoped to its own job the way a real set is. The AI executes inside each of those roles. The structure of the roles themselves is still the same structure a film crew has always had, because someone still has to decide what the shot should look like before anything generates it.
생성이 아니라 판단이 병목이라는 가장 명확한 증거는 영상이 완성된 후에 벌어집니다. 잘 문서화된 제작 과정에는 조립된 러프 컷을 다시 보내 페이싱 문제, 사운드 이슈, 그리고 감정의 톤이 어긋난 순간을 집중적으로 찾아내는 비평 단계가 포함됩니다. 바로 다음과 같은 것들입니다 invideo의 자체 워크플로 노트는 이 단계를 가장 자주 건너뛰면서도 사람 편집자가 놓치는 오류를 잡아내는 단계로 꼽습니다. 곱씹어 볼수록 참으로 이상한 문장입니다. 자동화된 검수가 숙련된 편집자가 놓친 부분을 짚어낸다니 말이죠. AI가 더 나은 안목을 가졌다는 뜻은 아닙니다. 이제 안목이 과정의 다른 지점에서, 즉 컷 하나하나가 아니라 완성된 전체를 대상으로 적용되고 있으며, 그 지적이 옳은지 판단하고 실행에 옮기는 사람은 여전히 필요하다는 뜻입니다.
So what is the creator actually making, if not the pixels. They're making the same thing a director has always made: the specific sequence of choices that turns raw coverage into a scene that lands the way it's supposed to. Which take gets used, which two generations get spliced together because neither one alone had the whole shot, which note from the critique pass actually matters and which one gets ignored, all of it decided by someone rather than generated by anything. None of that shows up if you only look at "who generated the footage," and all of it is the actual difference between a finished piece of work and four hundred generations sitting in a folder. 모델이 장면을 만들 수 있습니다. 어떤 버전의 장면이 남길 만한 가치가 있는지는 여전히 스스로 판단하지 못합니다.