Get the Most from Seedance 2.0 Prompts
Last Updated
Aug 5, 2026Fresh
Models Tested
Seedance 2.0, Seedance 2.0 HD
Introduction to Seedance 2.0 Video Generation
Seedance 2.0 generates video and sound together in a single pass: spoken dialogue with synced lips, ambient sound, and effects arrive already embedded in your 15-second clip. Developed by ByteDance, the model powers video creation on Ambience AI, so there is no separate soundtrack step to manage.
Sound changes how you prompt. A silent-video prompt describes what the camera sees. A Seedance prompt also scripts what the viewer hears, and the difference between crisp dialogue and garbled audio comes down to a few concrete rules.
This guide collects those rules from our own test generations: how to budget spoken words, structure all 15 seconds, keep stock music out of your mix, and reuse characters across clips with references. Every technique here was validated on real generations before it was written down.
Understanding Seedance 2.0's Capabilities
Seedance 2.0 is a native-audio video model: picture and sound come from one generation, so lips match words and footsteps land on frames. Clips run 15 seconds, which is three times the length of a standard 5-second AI clip and enough room for a scene with a beginning, middle, and end.
Core Capabilities
Native Audio in One Pass
- Lip-synced dialogue - Quoted lines in your prompt become spoken words on screen
- Ambient sound and effects - Rain, espresso machines, and footsteps generate with the picture
- 15-second clips - Room for a full scene, a product beat, or a short story
- 720p standard, 1080p HD - Choose Seedance 2.0 HD when resolution matters more than credits
Generation Modes
- Text-to-Video - Describe the scene and the sound, and Seedance builds both
- Image-to-Video - Supply a first frame and animate it with sound
- References - Anchor characters, products, and styles across clips (Team and Business plans)
Credits and Generation Times
- Seedance 2.0: 900 credits text-to-video, 825 image-to-video, 900 with references
- Seedance 2.0 HD: 1950 credits text-to-video, 1875 image-to-video, 1950 with references
- Generations typically finish in 2 to 5 minutes depending on mode and resolution
- Because audio is native, no separate music or voiceover step runs afterward
Writing Dialogue That Sounds Right
Dialogue is Seedance's signature feature and its easiest one to get wrong. The model speaks exactly what you put in quotation marks, so the way you write those quotes decides whether the delivery sounds natural or garbled.
The Word Budget
- About 20 spoken words per 15-second clip. That is the comfortable total across all lines.
- About 10 words per line. Short lines stay crisp. Long lines are where speech starts to smear.
- One sentence per quote. Never pack two sentences into a single quoted line. Split them into separate quotes with an action between.
In our test generations, garbled speech tracked line length rather than clip duration. A single 21-word quote transcribed as clean speech for five seconds and then dissolved into phoneme soup for the rest of the clip. The same brief, rewritten as two short quotes, transcribed word for word across the full 15 seconds.
The Anatomy of a Dialogue Prompt
1. Scene and Subject
Who is on camera, where they are, and what the setting sounds like.
2. Quoted Lines
The exact words, in quotation marks, one sentence per quote, inside the word budget.
3. Voice and Lip Direction
How the line is delivered, plus a restraint clause: "warm excited delivery, realistic lip articulation, no exaggerated mouth opening".
4. Audio Mix and Suppressors
One sentence ranking dialogue above ambience, then the stacked list of sounds you do not want.
Worked Example
Prompt (this exact one transcribed word for word in testing):
"Shot 1: A warm painterly illustrated barista, already sliding a cup across the sunlit bar, medium close-up, locked camera, facing camera. Soft cheerful delivery, realistic lip articulation, no exaggerated mouth opening, no head turns while speaking. She says 'One flat white, extra warm.' She wipes the steam wand, espresso machine hissing. Shot 2: She smiles and says 'Mornings start slow here.' Closing: she rests both hands on the counter in a held medium close-up for the final two seconds, silence holds final 0.5s. Dialogue clean and prominent, ambient cafe murmur subtle. - No music, no library audio, no voiceover narration, no on-screen text, no subtitles, no logo"
Spoken words: 9 of a comfortable 20
Sentences per quote: 1, with an action beat between the two quotes
Voice and lips: "Soft cheerful delivery, realistic lip articulation, no exaggerated mouth opening, no head turns while speaking"
Structure: Shot labels, never second ranges
Closing beat: A named hold for the final two seconds, with silence on the last half second
Audio mix: One mix sentence plus the stacked suppressor list
Voice Descriptors and Lip Restraint
Name the delivery, then rein in the mouth. Without a restraint clause the model over-acts: exaggerated jaw movement, grimaces on emphasis, and head turns that pull the lips out of sync. This is the clause we run on every dialogue prompt:
"Warm excited delivery, realistic lip articulation, no exaggerated mouth opening, no head turns while speaking."
- Favor a medium close-up with a locked camera. Lip sync degrades as the camera moves.
- Separate lines with physical action so the model has a natural pause to animate
- Add an accent or age when it matters: "American accent, warm excited delivery"
- Use contractions and plain words. Formal phrasing reads as stiff when spoken aloud.
Structuring the Full 15 Seconds
Fifteen seconds is long for AI video. Without structure the model front-loads the action and lets the last third drift, so the clip feels like it ends twice. The fix is to script beats that span the whole runtime.
Use Shot Labels, Skip Timestamps
Number your beats and let the model handle timing:
"Shot 1: [camera and action] ... Shot 2: [camera and action] ... Closing: [held frame]"
Second ranges like "0-5s:" and "5-10s:" read as text the model tries to honor literally. ByteDance documents precise-timing control as unstable, and timed segments produce abnormal output. Shot labels give the same structure with none of the damage.
- Two or three shots fit comfortably in one generation
- One camera instruction and at most one short quoted line per shot
- Order shots by what happens, never by clock time
A Three-Beat Template
Beat 1: Establish
Wide or medium shot that sets the place, the subject, and the ambient sound bed.
Beat 2: Deliver
The core action or the dialogue line, usually in a closer shot so lips and hands read clearly.
Beat 3: Close
A short, silent action beat that resolves the scene: a smile, a wipe of the counter, a look out the window.
End on Silence
Left undirected, the model invents its own transition and the last seconds degrade in both picture and audio. Direct the ending instead, naming the frame and the silence:
"Closing: she holds a thumbs up in a static medium close-up for the final two seconds, silence holds final 0.5s."
- Put dialogue in the first two thirds of the clip
- Reserve the last beat for a held frame, a pull-back, or a specific gesture
- Name the silence explicitly so the audio has room to land before the cut
Controlling Music and Ambient Sound
Left to its own judgment, Seedance often adds a generic stock music bed under your scene. Sometimes that is what you want. For dialogue, product demos, and anything that should feel real, it usually is not, and you have to opt out explicitly.
The Audio Mix Line
End every dialogue prompt with one sentence that ranks the mix, naming the ambience you do want:
"Dialogue clean and prominent, ambient cafe murmur subtle."
Ranking beats listing. Telling the model which element sits on top keeps speech intelligible even when the room is busy, and naming the ambience gives it something to fill the bed with instead of reaching for music.
Stack the Music Suppressors
Seedance has no separate negative prompt field. Instead, close the prompt with a stacked suppressor list on its own dash:
"- No music, no library audio, no voiceover narration, no on-screen text, no subtitles, no logo"
Each phrase catches a different default the model falls back to. A single "no music" still let library beds and narrator voiceovers through in our tests. The full stack, paired with the mix sentence above, kept every arm clean and free of burned-in captions.
Sound Words That Work
Ambience
room tone, rain on glass, distant traffic, wind through leaves, cafe murmur, birdsong
Effects
footsteps on gravel, espresso machine hiss, keyboard clicks, a zipper, a striking match
Delivery and Tone
warm low voice, bright and quick, hushed, deadpan, with a light laugh, half-whispered
When You Do Want Music
Describe it like a brief to a composer, rank it below the dialogue, and drop "no music" from the suppressor list:
"Soft lo-fi piano throughout. Dialogue clean and prominent, music low under the voice. - No library audio, no voiceover narration, no on-screen text, no subtitles, no logo"
First Frames and Image-to-Video
Seedance can start a clip two ways, and choosing deliberately saves credits and retries.
Let Seedance Compose the Frame
Text-to-video builds the opening image from your description. Best when the scene is new and you care more about motion and sound than an exact look.
Describe the opening frame with the same care you would give an image prompt: subject, setting, light, and mood.
Supply Your Own First Frame
Image-to-video animates an image you provide, and the output keeps its aspect ratio. Best for continuing a look you already have: a generated still, a product render, or a brand scene.
Your prompt then describes motion and sound rather than appearance: what moves, what is said, and what we hear.
Image-to-Video Prompting Tips
- Do not re-describe the image. The model can see it. Describe the change.
- Anchor motion to things in the frame: "steam rises from the cup", "she turns toward the window"
- Dialogue rules apply unchanged: word budget, one sentence per quote, audio mix line
- One important limit: Seedance declines photoreal photos of real people as input. See the references section below for the details and the alternative.
References for Consistent Characters
References solve the hardest problem in AI video: keeping the same character, product, or style across clips. You attach media alongside your prompt, then cite each item with an @ tag so the model knows what role it plays. References are available on Team and Business plans.
| Reference Type | How to Cite | Limit | Best For |
|---|---|---|---|
| Images | @Image1 through @Image9 | Up to 9 | Characters, products, wardrobe, and visual style |
| Videos | @Video1 through @Video3 | Up to 3 | Motion style, pacing, and camera behavior |
| Audio | @Audio1 through @Audio3 | Up to 3 | Voice character and music direction |
Reference Prompt Example
"Shot 1: The hiker from @Image1 crests a ridge at golden hour, wearing the jacket from @Image2, medium close-up, locked camera. Tired warm delivery, realistic lip articulation, no exaggerated mouth opening. She plants her trekking pole and says 'Worth every step.' Closing: she looks back at the valley in a held frame for the final two seconds, silence holds final 0.5s. Dialogue clean and prominent, wind and boots on rock subtle. - No music, no library audio, no voiceover narration, no on-screen text, no subtitles, no logo"
- Cite every attached reference at least once, and give each a clear role
- Keep the @ tag attached to a noun: "the hiker from @Image1", never a bare tag on its own
- All dialogue and structure rules from earlier sections still apply
The Photoreal Input Restriction
Seedance declines input images that show photoreal people. Illustrated, stylized, and 3D-rendered characters work, and characters described purely in the prompt work. This applies to first frames and to references alike.
To animate a real photo of a real person, use MiniMax H3 instead. It is built for exactly that job, and our MiniMax H3 prompting guide covers how its reference citations differ from Seedance's.
Troubleshooting & Solutions
These are the issues we hit most in testing, with the fixes that actually resolved them.
Garbled or Smeared Speech
Problem: Words blur together or the mouth stops matching the audio
Solution: Shorten the line to 10 words or fewer, keep one sentence per quote, and split long speeches into two quotes with action between them
Stock Music Appears
Problem: A generic music bed plays under the scene uninvited
Solution: Rank the mix ("Dialogue clean and prominent, ambient subtle") and append the full suppressor stack, including "no library audio"
Speech Cut Off at the End
Problem: The clip ends mid-word
Solution: Move all dialogue into the first two thirds of the clip and script a silent closing beat, named explicitly: "she smiles in silence"
Character Drift
Problem: The subject's face or outfit changes between clips
Solution: Attach the character as an image reference and cite it with @Image1, or reuse the same first frame with image-to-video
Dead Air in the Last Seconds
Problem: The subject sits idle after the action finishes
Solution: Script three distinct beats with shot labels so every stretch of the 15 seconds has a job
Input Image Rejected
Problem: A first frame or reference with a photoreal person is declined
Solution: Use a stylized or illustrated version of the character, describe them in the prompt, or switch to MiniMax H3 for real photos
Start Creating Videos with Real Sound
Seedance 2.0 folds picture, dialogue, and sound design into one prompt. The craft is in the constraints: a 20-word speech budget, one sentence per quote, shot labels for structure, an audio mix line to control the soundtrack, and a silent beat to close. Follow those and the model rewards you with clips that feel produced rather than generated.
From here, explore the rest of the video lineup. The MiniMax H3 guide covers the other native-audio model, including how to animate real photos of real people. The Kling guide and WAN guide cover our silent models, which remain the right choice for quick, low-credit clips you plan to score with generated music.
Ready to hear your first clip? Head to our AI video generator, pick Seedance 2.0, and start with the worked example above. Or browse the full suite of AI creative tools on the Ambience AI homepage.
Sources & Citations
This guide combines our own test generations with official documentation and authoritative technical sources.
- ByteDance Seed: Seedance Model Page - Official model overview
- BytePlus ModelArk - Platform documentation and prompting guidance
Ready to Create Videos with Real Sound?
Put the word budget, shot labels, and audio mix line to work. Your first Seedance 2.0 clip is one well-crafted prompt away.