Seedance 2.5 is one of the most capable AI models on OpenArt's AI video generator, and a lot of what makes it powerful comes down to how you prompt it.
It makes a full 30 seconds of video at once. It takes up to 50 reference files. It lets you edit one part of a clip without redoing the rest. Each of those features has a way you talk to it, and most people never learn the syntax.
This guide covers what actually works, including the exact wording ByteDance's own examples use.
Key takeaways
- Describe the whole scene and what changes over the 30 seconds, not one frozen moment.
- Label every reference file in your prompt, and name the exact part of it you want.
- Split your 30 seconds into timestamps like "0-8s" so each part gets its own direction.
- Put music, sound effects, and voice lines in their own brackets so they do not get mixed up.
- Fix one wrong element with a targeted edit instead of generating the whole video again.
Describe the scene, not just the subject
"A woman walking down a city street" is a fine starting point, but it leaves a lot up to the model.
Seedance 2.5 does much better when you describe the full scene. Say what the subject is doing, what the place looks like, and how the camera moves. "A woman in a red coat walks down a rainy city street at night, slow tracking shot from the side, streetlights reflected in the puddles."
Same subject, completely different video. The more you tell it, the less it has to guess.
Use references instead of describing looks
This is the biggest change Seedance 2.5 makes possible. You get 50 reference slots, so you do not need to spend half your prompt describing someone's face.
Upload a picture of the character and let the model copy it. Do the same for products, places, and styles. If you have a video clip with the mood or colors you want, add that too.
A good rule: if you can show it, do not describe it. Save your words for the things a picture cannot tell the model, like the action, the camera move, and the timing.
Label every reference in your prompt
This is the step most people skip, and it is the difference between a reference that helps and one that quietly ruins your video.
Uploading a file is not enough. Seedance 2.5 will not reliably guess that image 3 is the location and image 4 is the jacket. You have to say so in the prompt text. ByteDance's own example prompts do this on every single reference.
The pattern looks like this:
@Image 1 defines the woman's face, hair, and green jacket.
@Image 2 defines the coffee shop, the window, and the morning light.
@Video 1 defines the walking speed and the camera move.
@Audio 1 defines her voice.
Then write the action after the labels. Numbering works the same way for every file type, so you can point at @Image 7 or @Video 2 the same way.
Name the part of the picture you want
Every reference brings along things you did not ask for. A background, a stranger in the corner, a color you do not want. Those can show up in your video if the label is vague.
ByteDance's own guidance is to say which part of a reference to use, rather than listing what to leave out. So write the label narrow:
@Image 1 defines the woman's face and green jacket only.
@Image 2 defines the coffee shop counter and the window light only.
One extra word does a lot of work here. Saying "face only" is more reliable than uploading a headshot and hoping the model works out that you did not want the wall behind her.
Straight-up "do not" instructions are documented for a smaller set of things: captions, logos, watermarks, background music, and overall look. Those are safe to write plainly, like "no captions" or "no background music, keep the room sound."
Group multiple pictures of the same thing
If you upload four angles of one product, the model may treat them as four different products. Say they are the same object and say it plainly:
@Image 1 defines the front of the desk lamp. @Image 2 defines the left side of the same desk lamp. @Image 3 defines the back of the same desk lamp. All three images show one lamp. The video must contain only one lamp.
That last sentence is the one doing the work.
Mix your reference types
Seedance 2.5 accepts up to 30 images, 10 video clips, and 10 audio clips in one generation. That is 50 files total, and each type controls a different part of the result. Keep images under 4K, and keep your video clips adding up to 30 seconds or less. The same 30 second limit applies to your audio clips added together.
Pictures hold faces, products, and places. Video clips carry movement, timing, and the overall look. Audio clips shape voices and background sound.
The best results usually combine two or three types. For a product ad, that might be a product photo, a short clip for the style, and an audio file for the sound.
You do not need to fill all 50 slots. Eight or fewer people or products in one video stays reliable, and consistency starts slipping past that. Short reference clips of 5 to 10 seconds work better than long ones.
If you use the same character often, save it once in the OpenArt Characters library instead of hunting for the picture every time.
Split your 30 seconds into timestamps
Seedance 2.5 holds a full story across 30 seconds. So your prompt should say what happens across that time, not describe one frozen moment.
The clearest way to do that is to write time ranges directly into the prompt. ByteDance's own examples use this format:
0-5s: The chef sets an empty plate on the counter and picks up a spoon.
5-15s: Slow push in as sauce is drizzled across the plate in one line.
15-25s: A hand places three scallops down, one at a time.
25-30s: Pull back to a wide shot of the finished plate under a warm overhead light.
Timestamps beat "opens with, then, and finally" because the model gets both the order and the pacing. If a beat feels rushed in your first result, widen that one time range and keep the rest.
Three rules keep timestamps working. Use whole seconds, never half seconds. Leave no gaps in the timeline, so the next range picks up where the last one ended. Give each range one clear action, because an overstuffed range gets beats dropped and an empty one gets filled with something you did not ask for.
Timestamps are also not a counter. "Nods three times in one second" will not work, so describe the action instead of the number of repeats.
If you want more than 30 seconds, you do not need to stitch anything. Generate the first clip, then ask for an extension that continues from it and keeps the same characters, place, and style.
Name the camera move
Seedance 2.5 reads camera direction straight from your prompt. If you want a specific shot, name it.
Words that work: tracking shot, slow push in, wide establishing shot, handheld, aerial, dolly forward, orbit, rack focus, low angle, top down. "Cinematic" is not a camera move and tells the model nothing.
Pick one move per beat instead of stacking three. A single clear move that runs the full ten seconds looks far better than three fighting each other.
You can also copy the movement from a clip you already have. Upload it as a video reference and label it, like "@Video 1 defines the camera move and the pacing only." The model follows that motion instead of guessing at your words.
Write the sound into your prompt
Seedance 2.5 makes the audio and the video at the same time, so you can shape the sound with words. Describe the room: "quiet indoor space," "busy street with traffic," "music building to the last few seconds."
When a prompt has music, sound effects, voice lines, and on-screen captions all at once, plain sentences get muddled. Four bracket types keep them apart:
| What you want | Brackets to use | Example |
|---|---|---|
| Music | ( ) | (soft piano plays under the scene) |
| Sound effects | < > | <a bell rings in the distance> |
| Voice lines | { } | {Hello, welcome back.} |
| On-screen captions | 【 】 | 【Chapter One】 |
These come from ByteDance's own prompt guide, first written for Seedance 2.0 and still read by 2.5. Note that the music brackets are the wide ones, not regular parentheses.
You do not have to use brackets for every line. Reach for them when a plain sentence could be read two ways, like a spoken line that keeps showing up as a caption instead. ByteDance's newer examples often just use quotation marks with a speaker name, and that works too.
You can also turn things off in plain words. "No captions" and "no background music, keep the room sound" are both documented and both work.
Get voice lines in the right language
Two rules here come straight from ByteDance. Keep all your voice lines in one language, because mixing two in the same prompt causes problems. And any time the line is not in English or Chinese, name the language right before it.
She says in Japanese: {もう大丈夫です}
The same order helps when you want a specific accent or delivery. Put the language and the accent first, then how it is said, then the line itself.
Spoken language: American English. She says it fast and a little annoyed: {I told you this would happen.}
Naming it up front is more reliable than adding "in an American accent" at the end of the prompt.
Fix one thing instead of redoing the whole video
Region editing is for when 90% of your clip is right and one thing is not. A wrong label on a bottle, a background that does not fit, a face angle that missed.
Before you regenerate, check whether the problem sits in one spot. If it does, target just that spot and leave everything else alone. You keep every part of the clip that already worked.
The wording follows the same shape as your first prompt. Name the clip, say what stays, then say what changes:
Edit @Video 1. Keep the person, the action, and the camera move the same. Change only the bottle in her hand to the one in @Image 2, and match the original lighting.
Using @Video 1, replace only the background with the place in @Image 3. Keep the person, her clothes, and her movement exactly the same.
Edit @Video 1. 0-5s: no change. 5-12s: replace the text on the sign with the text in @Image 4, matching the original font and lighting. 12-20s: no change.
Notice that the last one combines timestamps with an edit, so only the middle of the clip is touched.
Two things to expect when you edit. The result keeps the same shape and length as the clip you fed in, so you cannot crop or trim in the same step. And clips of 20 seconds or less edit more reliably than longer ones.
Weak prompts vs strong prompts
The gap between a usable result and a wasted generation usually comes down to three things.
| What you are setting | Weak | Strong |
|---|---|---|
| Camera | A city skyline at dusk | Low aerial gliding over a skyline at dusk, one slow push toward a single lit tower |
| Change over time | A climber on a ridge | A climber pulls over the ridge, stands, and turns to the valley as the camera pulls back |
| References | Use these pictures | @Image 1 defines the woman's face and coat only. @Image 2 defines the market. The woman walks through the market. |
Five mistakes that waste generations
- Describing a photo instead of a video. Say what changes across the 30 seconds, not what one frame looks like.
- Leaving references unlabeled. If you upload five files, say what each one is for.
- Writing labels too wide. "@Image 1 defines the woman" pulls in her background too. Write "her face and jacket only."
- Stacking camera moves. Three moves in one beat fight each other. Pick one.
- Cramming a five-scene story into one paragraph. Use timestamps, or build it out shot by shot in OpenArt Director.
30 Seedance 2.5 prompt examples
Copy any of these, swap in your own subject, and add your reference labels on top.
Product and ecommerce
- A skincare bottle sits on a marble counter in soft window light. Slow 360 orbit at eye level, ending straight on the label. Quiet room tone with a soft chime on the final frame.
- A leather bag sits open on a linen backdrop. Close push in on the stitching, then pull back to a three-quarter angle. Soft studio sound and a light fabric rustle.
- An earbud case opens in slow motion, the lid lifting to show a small blue light inside. Camera tilts down from high to a tight close-up. A clean click on the lid, then a short brand chime.
- A running shoe hits a puddle in slow motion with droplets hanging in the air. Side tracking shot at ground level, then an orbit around the frozen splash. Sharp splash sound stretched long.
- A coffee bag tears open and beans pour into a glass jar. Overhead shot that shifts to a side view of the falling beans. Rich pouring sound, ending on warm cafe tone.
UGC and social ads
- A woman in a bright kitchen holds up a skincare bottle and talks straight to the camera about her morning routine. Handheld selfie angle, slight natural shake, warm window light. Casual room tone and clear voice lines.
- A man sits in a cafe explaining one feature of his app, gesturing as he talks. Static medium shot at eye level with the cafe soft behind him. Cafe chatter under his voice.
- A creator takes the first bite of a snack and reacts before describing the flavor. Handheld close-up at a kitchen counter. Crunch synced to the bite, then a natural spoken reaction.
- A woman unpacks a travel organizer on a hotel bed and shows how it fits in a suitcase. Overhead angle that shifts to a side view. Zipper and fabric sounds with casual narration.
- A trainer demonstrates one resistance band move in a home gym, explaining it between reps. Handheld shot following the movement. Breathing synced to the effort.
Characters and story scenes
- Two friends sit across a small table in late afternoon light. 0-10s: wide shot of the room. 10-20s: medium two-shot as the one on the left speaks. 20-30s: the other leans in to answer. Handheld with slight natural movement. Warm room noise under the voices.
- A man walks into an empty office at night and stops when the lights flicker on by themselves. Slow dolly forward behind him, then a cut to his face. Low hum, one sharp electrical snap, then silence.
- A girl in a cloak reaches toward a glowing stone in a forest clearing at dusk. Slow push in with a gentle rise as light spills out. Wind chimes and a soft swell timed to the burst of light.
- A grandmother teaches her grandson to fold dough at a kitchen table. Static medium shot, warm overhead light, no cuts. Kitchen sounds and quiet conversation.
- A courier runs up six flights of stairs with a package. Handheld camera following one step behind, going up with him. Footsteps, breathing, and a stairwell echo.
Fashion and beauty
- A model walks toward the camera down an empty street at dusk in an oversized coat. Steady tracking shot at hip height moving backward at her pace. Street noise with footsteps synced to each stride.
- A hand blends foundation across a cheek in extreme close-up. Static shot, soft even lighting. Quiet room tone with a light blending sound.
- A model turns in place in a long gown under a single overhead light. Camera orbits at the same speed as the turn. Fabric movement audible with a low string swell.
Food
- A chef drizzles sauce across a plate in slow motion. Overhead close-up, warm kitchen light. Drizzle sound stretched long, ending on kitchen room tone.
- Steam rises off a bowl of noodles as chopsticks lift them into frame. Close side angle at table height. Restaurant chatter with a soft slurp on the lift.
- A pizza comes out of a wood oven with the cheese still bubbling and lands on a board. One continuous camera move from oven to table. Fire crackle and sizzling cheese.
Nature and travel
- A whale glides past a diver in clear blue water with sunlight coming down in shafts. Wide static underwater shot, the diver small against the whale. Muffled water sound and low whale song.
- A hot air balloon drifts over a valley of vineyards at sunrise with mist in the rows below. Aerial shot following the drift. Distant burner flame and soft wind.
- A lighthouse beam sweeps across a rocky coast in thick fog. Slow static wide shot as the beam turns through frame. Foghorn, waves, and gulls.
Sports and action
- A skateboarder lands a trick down a set of stairs in a sunlit park. Low tracking shot on the approach, then a static hold on the landing. Wheels on concrete and a short cheer.
- A swimmer dives off the block in slow motion with water exploding around the entry. Side camera at water level that follows underwater. Splash stretched long, then muffled water tone.
- A cyclist takes a mountain switchback at speed, leaning into each turn. Aerial shot tracking from above and descending with the road. Wind rush building with speed.
Corporate and explainer
- A founder walks through an office explaining what the company does, straight to camera, with people working softly out of focus behind her. Steady tracking shot alongside her. Office noise under a clear voice.
- A laptop screen shows a dashboard as a cursor moves through three menus, each click lighting up a panel. Static overhead shot of the desk. Soft keyboard and click sounds.
- A team stands around a whiteboard as one person draws a simple three-step diagram. Static medium shot from the side, natural office light. Marker squeak and quiet room tone.
Three full sample prompts
Here is how the pieces come together.
Sample 1: Product brand film
Mode: Text to Video. References: product photo plus a brand style clip.
@Image 1 defines the serum bottle, the label, and the cap only. @Video 1 defines the color and the pace only, not the product.
A skincare serum bottle sits on white marble with soft light from the left. 0-10s: wide shot, slow push in as the light warms. 10-20s: a hand enters from the right and turns the bottle slowly. 20-30s: tight close-up on the label with the background soft. Quiet indoor room tone.
What this demonstrates: narrow reference labels, a three-part timeline, one camera move per beat, and sound named in a few words.
Sample 2: Multi-character scene
Mode: Reference Generation. References: two character photos, a cafe photo, and a style clip.
@Image 1 defines the first person's face, hair, and blue shirt only. @Image 2 defines the second person's face and glasses only. @Image 3 defines the cafe, the window seat, and the afternoon light. @Video 1 defines the handheld camera feel only.
Two people sit across a small cafe table in late afternoon light. 0-8s: wide shot of the room with warm light through the window. 8-20s: medium two-shot as the person on the left picks up a cup and speaks. 20-30s: the person on the right leans forward to answer. Handheld with slight natural movement. Warm cafe noise with soft chatter behind the voices.
What this demonstrates: character photos doing the work on faces so the text can handle action and timing, with "only" on each label so two stray backgrounds stay out of the shot.
Sample 3: Cinematic landscape
Mode: Image to Video. References: a landscape photo plus a style clip.
@Image 1 defines the cliff, the ocean, and the golden hour light and sets the first frame. @Video 1 defines the grade and pacing only.
A lone figure stands at the edge of a cliff with the ocean ahead. 0-15s: wide aerial from high above and behind, descending slowly toward eye level. 15-25s: hold at a medium distance behind the figure. 25-30s: pull back wide with the full horizon in view. (orchestral score building toward the end) <wind and distant waves>
What this demonstrates: one long camera move described as a timeline, an image reference that sets the opening frame, and brackets separating the score from the sound effects.
The bottom line
Write the scene and what changes across it. Label every reference and name the part of it you want. Split the 30 seconds into timestamps and name one camera move per beat.
Seedance 2.5 rewards being specific, and most of that specificity is syntax you can copy. Start with a prompt above that is close to your idea, swap in your own subject, and try it on OpenArt.
Frequently asked questions
How many reference files can I use in a Seedance 2.5 prompt?
Up to 30 images, 10 video clips, and 10 audio clips in one generation, which is 50 files total. Label each one in the prompt text so the model knows what it is for.
How do I write timestamps in a Seedance 2.5 prompt?
Write plain time ranges on their own lines, like "0-8s:" followed by what happens. ByteDance's official examples use this format, and it controls both the order of events and how long each one lasts.
Why does my product show up twice in the video?
You uploaded more than one angle of it and the model read them as separate objects. Label them as the same thing and add a closing line like "all three images show one lamp, the video must contain only one lamp."
Can I make a video longer than 30 seconds?
Yes. Generate your first 30 seconds, then ask for an extension that continues from that clip and keeps the same characters, place, and style. You can repeat that to reach several minutes without stitching clips together.
Why is my English dialogue spoken with the wrong accent?
Name the language and accent before the line rather than after it. Write "Spoken language: American English," then how it is said, then the line in curly brackets.
Do Seedance 2.5 prompts work on Seedance 2.0?
Most single-scene prompts work across all AI video models. Timestamps do not give the best results however, because Seedance 2.0 reads shot numbers instead. Rewrite "0-5s" and "5-15s" as "Shot 1" and "Shot 2" and the same prompt works. Uploading several angles of one subject is also 2.5 only.