In February 2026, ByteDance dropped Seedance 2.0. Within 24 hours it was everywhere, because a filmmaker generated a Hollywood-quality fight scene with a two-line prompt. The Deadpool screenwriter saw it and publicly posted that it was likely over for screenwriters. Elon Musk replied. The Motion Picture Association issued a formal statement the same day. ByteDance pulled the ability to generate real celebrity likenesses, tightened the guardrails, and the internet held its breath wondering if the model had been gutted in the process.
It wasn't. The underlying capability, the action, the physics, the motion, the emotional performance, all survived intact. And honestly, those conversations need to happen. Generative AI is moving faster than any legal framework can keep up with, and how the industry handles that matters. But that is a different video. What we can tell you is that Seedance 2.0 is the most capable video model on the market right now, and it is live on OpenArt.
Ermin Monzon, our Head of Content Media & Education, broke down why the model matters technically. Our in-house creator Bob stress-tested every reference type he could think of. And our collaborator Mia Meow built a 10-tip workflow guide from weeks of testing. This Handbook pulls the strongest material from all three to give you one complete picture.
Ce que Seedance 2.0 peut faire → Ermin's overview
Before we get into workflow tips, it is worth understanding why this model is generating the reaction it is. Ermin walked through a series of demos that show where Seedance 2.0 pulls away from everything else on the market.
The fight scene that broke the internet is the obvious headline, but the demo that surprised Ermin more was a dancer with neon light trails sweeping around them in orbital arcs. It looks like a Nike commercial or a music video with a serious production budget. The reason this is technically significant is that the light trails are not composited over the dancer. The model generated both simultaneously. The motion and the effect interact with each other: the trails bend around the dancer's body, and the lighting from the trails actually reflects on the floor beneath them. That is called coherent VFX generation. The model understands that light behaves physically. It does not just paint glowing lines on top of video. For anyone doing music video content, UGC ads, or anything with a visual effects component, this is where Seedance 2.0 stands apart.
Another demo worth watching: a skier launching off a mountain with an alpine range filling the background. The physics of the snow spray, the body rotation, the landing. All generated, all physically plausible. The model handles action and motion at a level that no other publicly available tool is matching right now.
Text-to-video: the single-take prompt structure → 0:50 in Mia's video
Everything you hear in Mia's text-to-video demos, the audio, the ambient sound, was generated at the same time as the video. One prompt, one take. Other models struggle to hold this many details in a single generation. Seedance 2.0 actually follows what you give it.
Mia's prompt structure for single-take shots is straightforward. Start with "continuous single take" and your camera actions. Then describe what the camera sees as it moves. End with control and style keywords: no cuts, seamless transition, cinematic, high definition. That formula alone will get you strong results.
Pour des scènes plus complexes, comme un fight sequence at 2:13, Mia recommends starting with a clear beginning and end state. Describe the fight scene, describe how it ends (everyone on the ground, for example), and let the model fill the middle. She also found that describing specific camera angles creates a more cinematic feel, and adding a film genre reference to the end of the prompt helps set the visual tone. One thing she discovered through iteration: if you describe your characters physically in the prompt rather than leaving it vague, the model does a better job designing the scene and keeping characters visually distinct.
Le conseil clé de Mia sur le prompting : le modèle est bien entraîné au langage du cinéma. Des termes comme « tracking shot », « crane shot », « whip pan » et des indices de genre comme « film de guerre brut » ou « thriller néo-noir » font toute la différence.

Le tutoriel complet de Mia couvre ses 10 astuces de workflow, y compris celles qu'on a résumées tout au long de ce guide. Regarde-le en parallèle ou reviens-y plus tard.
Le système de références : taguer images, vidéos et audio dans ton prompt → 4:27 dans la vidéo de Mia · → Le guide approfondi des références de Bob
C'est le mécanisme central qui alimente tout le reste dans Seedance 2.0 sur OpenArt, et une fois que tu l'auras compris, chaque section qui suit prendra tout son sens.
L'idée est simple : tu importes des images, des vidéos et des fichiers audio comme références, puis tu les tagues directement dans ton prompt texte grâce au symbole @. Tape @ et choisis le fichier importé que tu veux référencer. Ensuite, tu dis exactement au modèle quoi faire de chacun. Utilise le visage de l'image un. Reprends le mouvement de caméra de la vidéo un. Synchronise les lèvres sur l'audio un. Le modèle lit ton prompt, examine les fichiers tagués et combine le tout en une seule génération.
Mia's tutoriel multimodal à 4:27 offre l'explication la plus claire de son fonctionnement. Elle a importé une vidéo de danse, deux images de personnages et une chanson, puis a demandé au modèle d'utiliser la vidéo comme référence de mouvement de caméra, l'image une comme danseur de gauche, l'image deux comme danseur de droite et l'audio comme musique de fond. Le tout tagué avec @, le tout dans un seul prompt.
Bob pushed this further. In one of his demos, he referenced four different files in a single prompt: the face from image one, the overalls from image three, the shirt from image four, and a location from a reference video. The prompt was: the man in image one is walking down the street from the video wearing the overalls from image three and carrying the shirt from image four, talking to the camera. The model held consistency across every element. That is a remarkable amount of compositing happening from a single text prompt.

A few things both creators learned about how references behave:
Si ta vidéo de référence contient de l'audio et que tu tagues aussi un fichier audio séparé, the model tends to favor the video's existing audio over your uploaded track. Mia's workaround: upload the video without audio if you want your own music or voiceover to take priority.
Les références ne se limitent pas à des éléments concrets. You can also point to a video and say "I just want it to look like this" without copying anything else. Bob used this to transfer a visual style (fisheye lens, flickering light) from one video to an entirely new scene.
You do not have to use all reference types at once. A single character image in a text prompt is a perfectly valid use of the system. The power scales with how many elements you layer in, but it works fine at every level of complexity.

Lip sync, voice, and performance → 7:57 in Mia's video · → L'analyse audio approfondie de Bob
La vidéo complète de Bob couvre en profondeur les références dans les images, l'audio et la vidéo. Les sections sur le lip sync et le clonage vocal sont là où elle brille vraiment.
This is the section with the most tips from the most creators, and for good reason. Seedance 2.0's audio capabilities are the feature that separates it from everything else on the market right now.
Bob's transcript trick
This is the single most useful tip across all three videos. When you are doing lip sync with an audio reference, include the actual transcript of the words in your prompt alongside the audio file. Here is why.
Bob tested this with a clip of himself talking to camera. In his first attempt, he tagged the audio file and described the scene but did not include the exact words being spoken. The model listened to the audio and tried to replicate it, but got details wrong ("CGI 2.0" instead of "Seedance 2.0"). When he modified the prompt to explicitly spell out the words being said, the audio file combined with the written transcript produced dramatically better results. The lip sync was tighter, the words were accurate, and the overall performance was more convincing.

À retenir : chaque fois que tu fais du lip sync, fournis la transcription réelle avec le fichier audio. Tu es quasi assuré d'obtenir un meilleur résultat.
Fais correspondre la durée de ton audio à celle de ta vidéo
Mia l'a appris à ses dépens. Si ta référence audio dure 11 secondes et que tu règles la vidéo sur 15 secondes, le modèle est obligé d'étirer ou de compresser l'audio pour qu'il rentre, et tu perds presque à coup sûr le timing. Fais toujours correspondre le réglage de durée à la longueur réelle de ton fichier audio.
Voice cloning
Bob demonstrated voice cloning with a 15-second clip of Ermin's voice (Uncle Monz around the OpenArt office). The setup: record about 15 seconds of a voice in a way that captures the qualities you want duplicated. Tag that audio as a voice reference rather than a lip sync source. Then write the actual dialogue in the prompt, and the model generates new speech in that voice.
Le premier essai était trop rapide pour le débit naturel d'Ermin. Bob a donné au modèle 14 secondes complètes au lieu de 9, et le résultat ressemblait beaucoup plus à la façon dont Ermin parle vraiment. La leçon : le rythme compte autant que l'échantillon vocal. Donne au modèle assez de temps pour délivrer les mots à la bonne vitesse pour cette voix.
Tu peux réutiliser la même référence vocale sur plusieurs vidéos pour garder une voix cohérente pour un personnage.
Mia's performance directing tips → 12:20 in Mia's video
Mia found that you can direct the emotional performance and energy of a character through the prompt itself. The words you write influence how the character delivers the dialogue, not just what they say. If you write "he says it excitedly," the character's body language and vocal energy shift. If you write "she whispers nervously," the whole performance changes.
Cela s'applique aussi à plusieurs personnages dans la même scène. Chez 11:10 in Mia's video, she demonstrates giving two characters different vocal personalities and emotional states in one prompt, and the model keeps them distinct.
Mia's negative guardrail tip
Quand tu veux des transitions nettes d'une scène à l'autre sans morphing entre elles (par exemple, une narration sur quatre images différentes), Mia a appris à ajouter des garde-fous négatifs au prompt : « pas de morphing, pas de ghosting, pas de coupes de caméra ». Sans ça, le modèle mélange parfois les scènes de façons qui ressemblent à des glitchs visuels. Avec ces garde-fous, tu obtiens des transitions nettes d'une scène à la suivante.

Advanced references: motion, storyboards, and style → 13:22 dans la vidéo de Mia · → Références vidéo de Bob
Une fois que tu comprends comment fonctionne le système de références pour les visages et les voix, voici ce que tu peux lui indiquer d'autre.
Motion and camera transfer
If you have a shot where you like the way the camera moves, you can transfer that camera motion to an entirely new scene. Bob demonstrated this with a video of a dog jumping up and down as the camera dollied around it. He brought in that video as a reference and said "use that video as a reference for the camera and character movement to guide a scarecrow jumping on a trampoline." The result transferred both the camera dolly and the jumping action to the new character and scene.
Il a aussi montré le transfert de mouvement avec de vraies séquences. En utilisant une clip of Emily running at 19:15, il a dit « remplace la femme de la vidéo un par l'homme de l'image un et il neige ». Le modèle a échangé le personnage, ajouté un temps hivernal et conservé le mouvement de course intact. Certains détails d'arrière-plan ont changé (l'ambiance générale est devenue grise et hivernale), mais le mouvement principal a tenu.
Character Swap Split Screen LRB.mp4
Du storyboard à la vidéo → 13:22 dans la vidéo de Mia
Tu peux importer une grille de type planche de BD ou un storyboard, et le modèle tentera de l'animer case par case. Mia a testé ça avec une grille dessinée à la main et a constaté que le modèle suivait les cases dans l'ordre, mais que la logique directionnelle entre les cases n'était pas toujours évidente pour lui à partir des seuls visuels.
Two tips that made this work better. First, describe the narrative logic between panels in your prompt. The model needs to understand the story progression, not just the visual sequence. Second, if your storyboard has text or captions on it, add "text overlay no captions on screen" to the prompt to prevent text from bleeding into the generated video.
Mia précise aussi que tu peux utiliser de vraies pages de bandes dessinées pour ça, mais attention à tout ce qui touche à une propriété intellectuelle existante.
Références de style et réplication de templates → 14:45 dans la vidéo de Mia
Mia's template replication tip is a practical one for anyone producing content at volume. If you have a video with a visual style you like, you can reference it and tell the model to replicate just the style, not the content. She tested this by pointing to a video and saying she only wanted the fisheye lens effect and the flickering double-exposure look, then combined it with character images and outfit references to create a fashion sequence. The model picked up the abstract visual qualities without copying the specific content of the reference video.
C'est utile pour garder une identité visuelle cohérente sur une série de vidéos. Crée une vidéo qui te plaît, puis reprends son style pour toutes les générations suivantes.
Post-generation: extend and edit → 16:06 in Mia's video
Une vidéo n'est pas terminée une fois générée. Seedance 2.0 te permet de la prolonger, l'éditer et la réécrire après coup.

Étendre. Importe une vidéo que tu veux prolonger, choisis d'étendre vers l'avant (ou l'arrière), indique le nombre de secondes, ajuste ton réglage de durée en conséquence et décris ce que tu veux voir dans la nouvelle portion. Mia a prolongé une vidéo de 10 secondes vers l'avant, puis de 10 secondes vers l'arrière, et les a combinées en une scène fluide plus longue que la limite de 15 secondes par génération. C'est un vrai déblocage. La limite de 15 secondes était autrefois un plafond strict. Tu peux désormais construire des scènes fluides plus longues, sans coupures.
Modifier. Si tu veux modifier quelque chose qui s'est déjà produit dans une vidéo générée, tu peux dialoguer avec le modèle à propos de la vidéo existante et lui dire quoi changer. Ça fonctionne comme l'édition vidéo vers vidéo de Wan 2.7, mais avec la meilleure compréhension de Seedance 2.0 de ce qui se trouve dans la scène.
En résumé
Seedance 2.0 is the most capable video model available on OpenArt right now, and the reference system is the reason to pay attention. The ability to tag images, video, and audio into a single prompt and have the model hold consistency across all of them is not something any other model is doing at this level. The lip sync is strong. The voice cloning works. The motion transfer is clean. The coherent VFX generation, where lighting and effects actually interact with the scene physically, is new territory entirely.
Trois créateurs, trois vidéos, une conclusion partagée : ce modèle écoute. Il suit les prompts complexes et à éléments multiples mieux que tout ce qui existe. Le mieux à faire, c'est de te lancer et de commencer à lui demander ce que tu veux, parce qu'il a une imagination incroyable. Ce sont les mots de Bob, et il a raison.