Limited-time offer! Unlock a year of limitless creativity with annual plans at UP TO 27% OFF.

Voir le forfait ›
Cultural Conversations

Le plan Frankenstein : que reste-t-il quand l'IA réalise toute la prise

O
Emily Watterson
Jul 23, 2026 · 7 minutes read
The Frankenstein Shot: What's Left When AI Makes the Whole Take

A two-minute brand film shot with an AI-agent pipeline recently ran three days from script to finished cut, against roughly two months for the traditional equivalent shoot, according to invideo's own documented production case. A ninety-second horror short pulled around four hundred separate video generations before it was done. Numbers like that used to be the whole story: AI got fast enough to make an entire scene, so why would you need a person in the room. The more interesting number is the other one buried in that same case study: more than 40 percent of the finished shots in one documented project weren't a single generation at all. They were stitched together from the strongest few seconds of several different attempts, a practice invideo's own team describes bluntly: "Prompt, eight tries, Frankenstein the keepers."

Ce détail en dit plus sur ce qui se passe réellement que les promesses de vitesse. A model that can render a whole scene on command doesn't remove the need for a person making decisions, it just moves all those decisions earlier and makes them faster to act on. The workflow behind these AI-made shorts isn't one person typing a single magic sentence and getting a finished film back. It's a full crew of separate roles, a producer role holding the script and characters, a storyboard role visualizing shots before anything gets generated, a cinematography role taking direction like "hold the shot longer" or "track the actor through the doorway," a costume role, a production design role, each one scoped to its own job the way a real set is. The AI executes inside each of those roles. The structure of the roles themselves is still the same structure a film crew has always had, because someone still has to decide what the shot should look like before anything generates it.

La preuve la plus claire que le goulot d'étranglement est le jugement, et non la génération, c'est ce qui se passe une fois les images créées. Les productions documentées intègrent une étape où le montage brut assemblé est renvoyé pour une passe de critique cherchant spécifiquement les problèmes de rythme, de son, et les moments où le registre émotionnel sonne faux, ce genre de chose Les notes de workflow d'invideo décrivent cette étape comme la plus souvent négligée, mais celle qui repère les erreurs que les monteurs humains manquent. C'est une phrase vraiment étrange à assimiler : un passage automatisé qui repère ce qu'un monteur chevronné a laissé passer. Ça ne veut pas dire que l'IA a meilleur goût. Ça veut dire que le goût s'applique désormais à un autre moment du processus, sur l'ensemble assemblé plutôt que plan par plan, et que quelqu'un doit toujours décider si la remarque est juste et agir en conséquence.

So what is the creator actually making, if not the pixels. They're making the same thing a director has always made: the specific sequence of choices that turns raw coverage into a scene that lands the way it's supposed to. Which take gets used, which two generations get spliced together because neither one alone had the whole shot, which note from the critique pass actually matters and which one gets ignored, all of it decided by someone rather than generated by anything. None of that shows up if you only look at "who generated the footage," and all of it is the actual difference between a finished piece of work and four hundred generations sitting in a folder. Le modèle peut créer la scène. Il ne peut toujours pas décider, tout seul, quelle version de la scène valait la peine d'être conservée.

Crée sans limites

Rejoins des millions de créateurs qui utilisent OpenArt pour générer des images, des vidéos, des personnages et des histoires — le tout sur une seule plateforme.

Commence gratuitement →