Seedance 2.5 es uno de los modelos de IA más capaces de OpenArt Generador de vídeo con IA, and a lot of what makes it powerful comes down to how you prompt it.
Crea 30 segundos completos de vídeo de una vez. Admite hasta 50 archivos de referencia. Te deja editar una parte de un clip sin rehacer el resto. Cada una de esas funciones tiene su forma de comunicarte con ella, y casi nadie llega a aprender la sintaxis.
Esta guía cubre lo que realmente funciona, incluida la redacción exacta que usan los propios ejemplos de ByteDance.
Conclusiones clave
- Describe toda la escena y lo que cambia durante los 30 segundos, no un solo momento congelado.
- Etiqueta cada archivo de referencia en tu prompt e indica la parte exacta que quieres usar.
- Split your 30 seconds into timestamps like "0-8s" so each part gets its own direction.
- Put music, sound effects, and voice lines in their own brackets so they do not get mixed up.
- Fix one wrong element with a targeted edit instead of generating the whole video again.
Describe la escena, no solo el sujeto
"A woman walking down a city street" is a fine starting point, but it leaves a lot up to the model.
Seedance 2.5 does much better when you describe the full scene. Say what the subject is doing, what the place looks like, and how the camera moves. "A woman in a red coat walks down a rainy city street at night, slow tracking shot from the side, streetlights reflected in the puddles."
Same subject, completely different video. The more you tell it, the less it has to guess.
Use references instead of describing looks
This is the biggest change Seedance 2.5 makes possible. You get 50 reference slots, so you do not need to spend half your prompt describing someone's face.
Upload a picture of the character and let the model copy it. Do the same for products, places, and styles. If you have a video clip with the mood or colors you want, add that too.
Una buena regla: si puedes mostrarlo, no lo describas. Reserva tus palabras para lo que una imagen no puede transmitirle al modelo, como la acción, el movimiento de cámara y el ritmo.
Label every reference in your prompt
This is the step most people skip, and it is the difference between a reference that helps and one that quietly ruins your video.
Uploading a file is not enough. Seedance 2.5 will not reliably guess that image 3 is the location and image 4 is the jacket. You have to say so in the prompt text. ByteDance's own example prompts do this on every single reference.
The pattern looks like this:
@Image 1 define la cara, el pelo y la chaqueta verde de la mujer.
@Image 2 defines the coffee shop, the window, and the morning light.
@Video 1 define la velocidad al caminar y el movimiento de cámara.
@Audio 1 defines her voice.
Luego escribe la acción después de las etiquetas. La numeración funciona igual para todos los tipos de archivo, así que puedes señalar @Imagen 7 o @Vídeo 2 de la misma forma.
Name the part of the picture you want
Every reference brings along things you did not ask for. A background, a stranger in the corner, a color you do not want. Those can show up in your video if the label is vague.
La propia guía de ByteDance recomienda indicar qué parte de una referencia usar, en lugar de enumerar lo que hay que dejar fuera. Así que escribe la etiqueta de forma concreta:
@Image 1 define únicamente la cara y la chaqueta verde de la mujer.
@Image 2 defines the coffee shop counter and the window light only.
One extra word does a lot of work here. Saying "face only" is more reliable than uploading a headshot and hoping the model works out that you did not want the wall behind her.
Straight-up "do not" instructions are documented for a smaller set of things: captions, logos, watermarks, background music, and overall look. Those are safe to write plainly, like "no captions" or "no background music, keep the room sound."
Agrupa varias imágenes de lo mismo
Si subes cuatro ángulos de un mismo producto, el modelo puede tratarlos como cuatro productos distintos. Di que son el mismo objeto y dilo claramente:
@Image 1 define el frente de la lámpara de escritorio. @Image 2 define el lado izquierdo de la misma lámpara. @Image 3 define la parte trasera de la misma lámpara. Las tres imágenes muestran una única lámpara. El vídeo debe contener solo una lámpara.
That last sentence is the one doing the work.
Mix your reference types
Seedance 2.5 accepts up to 30 images, 10 video clips, and 10 audio clips in one generation. That is 50 files total, and each type controls a different part of the result. Keep images under 4K, and keep your video clips adding up to 30 seconds or less. The same 30 second limit applies to your audio clips added together.
Las fotos contienen caras, productos y lugares. Los clips de vídeo aportan movimiento, ritmo y el aspecto general. Los clips de audio dan forma a las voces y al sonido de fondo.
Los mejores resultados suelen combinar dos o tres tipos. Para un anuncio de producto, eso podría ser una foto del producto, un clip corto para el estilo y un archivo de audio para el sonido.
You do not need to fill all 50 slots. Eight or fewer people or products in one video stays reliable, and consistency starts slipping past that. Short reference clips of 5 to 10 seconds work better than long ones.
If you use the same character often, save it once in the Biblioteca de personajes de OpenArt instead of hunting for the picture every time.
Divide tus 30 segundos en marcas de tiempo
Seedance 2.5 mantiene una historia completa a lo largo de 30 segundos. Por eso tu prompt debe decir qué ocurre durante ese tiempo, no describir un momento congelado.
La forma más clara de hacerlo es escribir los rangos de tiempo directamente en el prompt. Los propios ejemplos de ByteDance usan este formato:
0-5s: The chef sets an empty plate on the counter and picks up a spoon.
5-15s: Slow push in as sauce is drizzled across the plate in one line.
15-25s: A hand places three scallops down, one at a time.
25-30s: Pull back to a wide shot of the finished plate under a warm overhead light.
Las marcas de tiempo funcionan mejor que «empieza con, luego, y por último» porque el modelo capta tanto el orden como el ritmo. Si un momento se siente apresurado en tu primer resultado, amplía solo ese rango de tiempo y deja el resto igual.
Three rules keep timestamps working. Use whole seconds, never half seconds. Leave no gaps in the timeline, so the next range picks up where the last one ended. Give each range one clear action, because an overstuffed range gets beats dropped and an empty one gets filled with something you did not ask for.
Las marcas de tiempo tampoco funcionan como contador. «Asiente tres veces en un segundo» no servirá, así que describe la acción en lugar del número de repeticiones.
Si quieres más de 30 segundos, no necesitas unir nada. Genera el primer clip y luego pide una extensión que continúe a partir de él y mantenga los mismos personajes, lugar y estilo.
Name the camera move
Seedance 2.5 reads camera direction straight from your prompt. If you want a specific shot, name it.
Words that work: tracking shot, slow push in, wide establishing shot, handheld, aerial, dolly forward, orbit, rack focus, low angle, top down. "Cinematic" is not a camera move and tells the model nothing.
Pick one move per beat instead of stacking three. A single clear move that runs the full ten seconds looks far better than three fighting each other.
También puedes copiar el movimiento de un clip que ya tengas. Súbelo como referencia de vídeo y etiquétalo, por ejemplo: «@Video 1 define solo el movimiento de cámara y el ritmo». El modelo sigue ese movimiento en lugar de adivinar a partir de tus palabras.
Write the sound into your prompt
Seedance 2.5 genera el audio y el vídeo a la vez, así que puedes moldear el sonido con palabras. Describe el ambiente: «espacio interior tranquilo», «calle concurrida con tráfico», «música que va creciendo hasta los últimos segundos».
Cuando un prompt tiene música, efectos de sonido, líneas de voz y subtítulos en pantalla a la vez, las frases sueltas se lían. Cuatro tipos de corchetes las mantienen separadas:
| What you want | Corchetes a usar | Example |
|---|---|---|
| Music | ( ) | (soft piano plays under the scene) |
| Sound effects | < > | <suena una campana a lo lejos> |
| Voice lines | { } | {Hello, welcome back.} |
| Subtítulos en pantalla | 【 】 | 【Chapter One】 |
Provienen de la guía de prompts de la propia ByteDance, escrita originalmente para Seedance 2.0 y que sigue vigente para la 2.5. Ten en cuenta que los corchetes de la música son los anchos, no los paréntesis normales.
No tienes que usar corchetes en cada línea. Recurre a ellos cuando una frase normal pueda interpretarse de dos maneras, como una línea hablada que sigue apareciendo como subtítulo. Los ejemplos más recientes de ByteDance a menudo solo usan comillas con el nombre del hablante, y eso también funciona.
You can also turn things off in plain words. "No captions" and "no background music, keep the room sound" are both documented and both work.
Consigue diálogos de voz en el idioma adecuado
Two rules here come straight from ByteDance. Keep all your voice lines in one language, because mixing two in the same prompt causes problems. And any time the line is not in English or Chinese, name the language right before it.
Dice en japonés: {もう大丈夫です}
The same order helps when you want a specific accent or delivery. Put the language and the accent first, then how it is said, then the line itself.
Spoken language: American English. She says it fast and a little annoyed: {I told you this would happen.}
Naming it up front is more reliable than adding "in an American accent" at the end of the prompt.
Fix one thing instead of redoing the whole video
La edición por regiones es para cuando el 90 % de tu clip está bien y solo una cosa no encaja. Una etiqueta incorrecta en una botella, un fondo que no cuadra, un ángulo de cara que no dio en el clavo.
Before you regenerate, check whether the problem sits in one spot. If it does, target just that spot and leave everything else alone. You keep every part of the clip that already worked.
La redacción sigue la misma estructura que tu primer prompt. Nombra el clip, di qué se mantiene y luego di qué cambia:
Edita @Video 1. Mantén igual la persona, la acción y el movimiento de cámara. Cambia solo la botella que tiene en la mano por la de @Image 2 e iguala la iluminación original.
Using @Video 1, replace only the background with the place in @Image 3. Keep the person, her clothes, and her movement exactly the same.
Edit @Video 1. 0-5s: no change. 5-12s: replace the text on the sign with the text in @Image 4, matching the original font and lighting. 12-20s: no change.
Fíjate en que el último combina marcas de tiempo con una edición, así que solo se modifica la parte central del clip.
Two things to expect when you edit. The result keeps the same shape and length as the clip you fed in, so you cannot crop or trim in the same step. And clips of 20 seconds or less edit more reliably than longer ones.
Weak prompts vs strong prompts
The gap between a usable result and a wasted generation usually comes down to three things.
| What you are setting | Débil | Strong |
|---|---|---|
| Cámara | A city skyline at dusk | Low aerial gliding over a skyline at dusk, one slow push toward a single lit tower |
| Change over time | A climber on a ridge | A climber pulls over the ridge, stands, and turns to the valley as the camera pulls back |
| References | Use these pictures | @Image 1 defines the woman's face and coat only. @Image 2 defines the market. The woman walks through the market. |
Five mistakes that waste generations
- Describing a photo instead of a video. Say what changes across the 30 seconds, not what one frame looks like.
- Leaving references unlabeled. If you upload five files, say what each one is for.
- Writing labels too wide. "@Image 1 defines the woman" pulls in her background too. Write "her face and jacket only."
- Stacking camera moves. Three moves in one beat fight each other. Pick one.
- Cramming a five-scene story into one paragraph. Use timestamps, or build it out shot by shot in OpenArt Director.
30 ejemplos de prompts de Seedance 2.5
Copy any of these, swap in your own subject, and add your reference labels on top.
Product and ecommerce
- A skincare bottle sits on a marble counter in soft window light. Slow 360 orbit at eye level, ending straight on the label. Quiet room tone with a soft chime on the final frame.
- A leather bag sits open on a linen backdrop. Close push in on the stitching, then pull back to a three-quarter angle. Soft studio sound and a light fabric rustle.
- An earbud case opens in slow motion, the lid lifting to show a small blue light inside. Camera tilts down from high to a tight close-up. A clean click on the lid, then a short brand chime.
- A running shoe hits a puddle in slow motion with droplets hanging in the air. Side tracking shot at ground level, then an orbit around the frozen splash. Sharp splash sound stretched long.
- A coffee bag tears open and beans pour into a glass jar. Overhead shot that shifts to a side view of the falling beans. Rich pouring sound, ending on warm cafe tone.
UGC and social ads
- A woman in a bright kitchen holds up a skincare bottle and talks straight to the camera about her morning routine. Handheld selfie angle, slight natural shake, warm window light. Casual room tone and clear voice lines.
- A man sits in a cafe explaining one feature of his app, gesturing as he talks. Static medium shot at eye level with the cafe soft behind him. Cafe chatter under his voice.
- A creator takes the first bite of a snack and reacts before describing the flavor. Handheld close-up at a kitchen counter. Crunch synced to the bite, then a natural spoken reaction.
- A woman unpacks a travel organizer on a hotel bed and shows how it fits in a suitcase. Overhead angle that shifts to a side view. Zipper and fabric sounds with casual narration.
- A trainer demonstrates one resistance band move in a home gym, explaining it between reps. Handheld shot following the movement. Breathing synced to the effort.
Characters and story scenes
- Two friends sit across a small table in late afternoon light. 0-10s: wide shot of the room. 10-20s: medium two-shot as the one on the left speaks. 20-30s: the other leans in to answer. Handheld with slight natural movement. Warm room noise under the voices.
- A man walks into an empty office at night and stops when the lights flicker on by themselves. Slow dolly forward behind him, then a cut to his face. Low hum, one sharp electrical snap, then silence.
- A girl in a cloak reaches toward a glowing stone in a forest clearing at dusk. Slow push in with a gentle rise as light spills out. Wind chimes and a soft swell timed to the burst of light.
- A grandmother teaches her grandson to fold dough at a kitchen table. Static medium shot, warm overhead light, no cuts. Kitchen sounds and quiet conversation.
- A courier runs up six flights of stairs with a package. Handheld camera following one step behind, going up with him. Footsteps, breathing, and a stairwell echo.
Fashion and beauty
- A model walks toward the camera down an empty street at dusk in an oversized coat. Steady tracking shot at hip height moving backward at her pace. Street noise with footsteps synced to each stride.
- A hand blends foundation across a cheek in extreme close-up. Static shot, soft even lighting. Quiet room tone with a light blending sound.
- A model turns in place in a long gown under a single overhead light. Camera orbits at the same speed as the turn. Fabric movement audible with a low string swell.
Comida
- A chef drizzles sauce across a plate in slow motion. Overhead close-up, warm kitchen light. Drizzle sound stretched long, ending on kitchen room tone.
- Steam rises off a bowl of noodles as chopsticks lift them into frame. Close side angle at table height. Restaurant chatter with a soft slurp on the lift.
- A pizza comes out of a wood oven with the cheese still bubbling and lands on a board. One continuous camera move from oven to table. Fire crackle and sizzling cheese.
Nature and travel
- A whale glides past a diver in clear blue water with sunlight coming down in shafts. Wide static underwater shot, the diver small against the whale. Muffled water sound and low whale song.
- A hot air balloon drifts over a valley of vineyards at sunrise with mist in the rows below. Aerial shot following the drift. Distant burner flame and soft wind.
- A lighthouse beam sweeps across a rocky coast in thick fog. Slow static wide shot as the beam turns through frame. Foghorn, waves, and gulls.
Sports and action
- A skateboarder lands a trick down a set of stairs in a sunlit park. Low tracking shot on the approach, then a static hold on the landing. Wheels on concrete and a short cheer.
- A swimmer dives off the block in slow motion with water exploding around the entry. Side camera at water level that follows underwater. Splash stretched long, then muffled water tone.
- A cyclist takes a mountain switchback at speed, leaning into each turn. Aerial shot tracking from above and descending with the road. Wind rush building with speed.
Corporate and explainer
- A founder walks through an office explaining what the company does, straight to camera, with people working softly out of focus behind her. Steady tracking shot alongside her. Office noise under a clear voice.
- A laptop screen shows a dashboard as a cursor moves through three menus, each click lighting up a panel. Static overhead shot of the desk. Soft keyboard and click sounds.
- A team stands around a whiteboard as one person draws a simple three-step diagram. Static medium shot from the side, natural office light. Marker squeak and quiet room tone.
Three full sample prompts
Here is how the pieces come together.
Sample 1: Product brand film
Mode: Text to Video. References: product photo plus a brand style clip.
@Image 1 defines the serum bottle, the label, and the cap only. @Video 1 defines the color and the pace only, not the product.
A skincare serum bottle sits on white marble with soft light from the left. 0-10s: wide shot, slow push in as the light warms. 10-20s: a hand enters from the right and turns the bottle slowly. 20-30s: tight close-up on the label with the background soft. Quiet indoor room tone.
What this demonstrates: narrow reference labels, a three-part timeline, one camera move per beat, and sound named in a few words.
Sample 2: Multi-character scene
Mode: Reference Generation. References: two character photos, a cafe photo, and a style clip.
@Image 1 defines the first person's face, hair, and blue shirt only. @Image 2 defines the second person's face and glasses only. @Image 3 defines the cafe, the window seat, and the afternoon light. @Video 1 defines the handheld camera feel only.
Two people sit across a small cafe table in late afternoon light. 0-8s: wide shot of the room with warm light through the window. 8-20s: medium two-shot as the person on the left picks up a cup and speaks. 20-30s: the person on the right leans forward to answer. Handheld with slight natural movement. Warm cafe noise with soft chatter behind the voices.
What this demonstrates: character photos doing the work on faces so the text can handle action and timing, with "only" on each label so two stray backgrounds stay out of the shot.
Sample 3: Cinematic landscape
Mode: Image to Video. References: a landscape photo plus a style clip.
@Image 1 defines the cliff, the ocean, and the golden hour light and sets the first frame. @Video 1 defines the grade and pacing only.
A lone figure stands at the edge of a cliff with the ocean ahead. 0-15s: wide aerial from high above and behind, descending slowly toward eye level. 15-25s: hold at a medium distance behind the figure. 25-30s: pull back wide with the full horizon in view. (orchestral score building toward the end) <wind and distant waves>
What this demonstrates: one long camera move described as a timeline, an image reference that sets the opening frame, and brackets separating the score from the sound effects.
The bottom line
Write the scene and what changes across it. Label every reference and name the part of it you want. Split the 30 seconds into timestamps and name one camera move per beat.
Seedance 2.5 rewards being specific, and most of that specificity is syntax you can copy. Start with a prompt above that is close to your idea, swap in your own subject, and try it on OpenArt.
Frequently asked questions
How many reference files can I use in a Seedance 2.5 prompt?
Up to 30 images, 10 video clips, and 10 audio clips in one generation, which is 50 files total. Label each one in the prompt text so the model knows what it is for.
How do I write timestamps in a Seedance 2.5 prompt?
Write plain time ranges on their own lines, like "0-8s:" followed by what happens. ByteDance's official examples use this format, and it controls both the order of events and how long each one lasts.
Why does my product show up twice in the video?
You uploaded more than one angle of it and the model read them as separate objects. Label them as the same thing and add a closing line like "all three images show one lamp, the video must contain only one lamp."
Can I make a video longer than 30 seconds?
Yes. Generate your first 30 seconds, then ask for an extension that continues from that clip and keeps the same characters, place, and style. You can repeat that to reach several minutes without stitching clips together.
Why is my English dialogue spoken with the wrong accent?
Name the language and accent before the line rather than after it. Write "Spoken language: American English," then how it is said, then the line in curly brackets.
Do Seedance 2.5 prompts work on Seedance 2.0?
Most single-scene prompts work across all AI video models. Timestamps do not give the best results however, because Seedance 2.0 reads shot numbers instead. Rewrite "0-5s" and "5-15s" as "Shot 1" and "Shot 2" and the same prompt works. Uploading several angles of one subject is also 2.5 only.