Seedance 2.0 — The Complete Human-Friendly Prompting Guide
Original technical spec sources: ByteDance Seed Official Blog · ZenCreator Guide · Forbes · fal.ai · Atlabs AI · Freepik Blog · Imagine.Art · TechCrunch
Welcome: Why Seedance 2.0 Is Worth Your Attention
If you've been using AI video generators and hitting the same walls — silent outputs that need audio layered in post, 8-second clips that feel abruptly cut off, characters who lose their face between shots — Seedance 2.0 was built specifically to solve those problems.
Launched by ByteDance's Seed AI team on February 10, 2026, Seedance 2.0 is a fundamentally different kind of video model. Here's what makes it genuinely special:
It generates video and audio simultaneously. Not video first, then audio pasted on top. A single pass through a Dual-Branch Diffusion Transformer produces picture and sound as one unified artifact. That means dialogue is lip-synced because it was designed that way from frame one — not because of a post-processing alignment trick. Sound effects land exactly when things happen on screen because both were computed together. This is the architecture difference that matters most.
It runs for 15 seconds with multiple shots. Most competitors top out at 8 seconds and a single continuous camera move. Seedance 2.0 lets you cut between angles, switch from slow motion to normal speed, and tell a story with a beginning, middle, and end — all in a single generation. That's not a small upgrade. That's a workflow change.
It can hold up to 12 reference files at once. Nine images, three video clips, and three audio clips — all tagged with an @ system and all influencing the output in parallel. Want to hand the model a character face, a costume reference, an environment photo, a storyboard sketch, a camera movement clip, and a music track all at once? That's legal. That's the vision.
Its physics engine is something people actually comment on. Fabric drapes and folds correctly. Liquids pour with realistic viscosity. Lighting bounces off surfaces the way light actually behaves. Forbes specifically highlighted the "hyper-real outputs" and physical plausibility as distinguishing characteristics.
The usable output rate is over 90%. In practice, that means you're not running a generation six times hoping one comes out okay. Most generations are publication-ready.
This guide will take you from zero to fluent — covering the reference system, camera vocabulary, audio direction, multi-shot scripting, and the difference between prompts that work and prompts that waste credits.
Table of Contents
- Model Identity & Access
- Technical Specifications
- Core Architecture Insight
- Quick Start — 3 Copy-Paste Prompts
- The Multimodal Reference System
- Multimodal Reference Masterclass
- Prompt Formula: Subject → Action → Camera → Style
- Camera Movement Vocabulary
- Audio Direction
- Key Prompting Best Practices
- Video Extension
- Video Editing
- Shot Size Guide
- Example Prompts (Production-Ready)
- Common Mistakes → Better Versions
- Seedance 2.0 vs. Veo 3.1 — Practical Comparison
- Prompt Building Workshop
- Cheat Sheet
1. Model Identity & Access
| Field | Value |
|---|---|
| Product Name | Seedance 2.0 |
| Developer | ByteDance Seed AI |
| Launch Date | February 10, 2026 |
| Category | Video generation with native audio |
| Architecture | Dual-Branch Diffusion Transformer (audio + video generated simultaneously) |
| Access | Dreamina (Jimeng AI), CapCut (rolling out March 2026), fal.ai, various third-party platforms |
| API Availability | Announced for Q3 2026 |
| Predecessor | Seedance 1.0 (silent), Seedance 1.5 Pro (basic audio) |
Where to use it right now:
- Dreamina / Jimeng AI — ByteDance's primary creative platform, full feature access
- fal.ai — API-friendly, good for developers and power users
- CapCut — Consumer app integration rolling out through 2026 per TechCrunch
2. Technical Specifications
| Parameter | Value |
|---|---|
| Max Resolution | 2K (up to 2048px) |
| Aspect Ratios | 16:9, 4:3, 1:1, 3:4, 9:16 |
| Max Duration | 15 seconds per generation |
| Multi-Shot | Supported within single generation (multiple cuts/angles) |
| Audio | Native dual-channel stereo; dialogue, SFX, ambient, music — all generated in single pass |
| Audio Channels | 2-channel stereo |
| Audio Tracks | Multi-track parallel: background music + ambient SFX + character voiceover |
| Max Image Inputs | 9 images |
| Max Video Inputs | 3 clips (15 seconds total) |
| Max Audio Inputs | 3 clips (15 seconds total) |
| Total Max Reference Files | Up to 12 (9 images + 3 videos) + 3 audio |
| Character Consistency | Maintained across multi-shot sequences |
| Video Extension | Supported (continuation from existing clip) |
| Video Editing | Supported (targeted modifications to clips, characters, actions) |
| Generation Time | ~60s (standard) to ~10 min (15s with multiple references) |
| Watermarking | Invisible watermark for off-platform identification |
| Usable Output Rate | 90%+ (reported) |
A note on generation time: Patience pays off here. The longer generation times (especially with many references) aren't a bug — they're the model doing serious work. A 10-minute generation that produces a publication-ready 15-second clip is still faster than a traditional shoot day.
3. Core Architecture Insight
Seedance 2.0 uses a Dual-Branch Diffusion Transformer that generates video and audio simultaneously in a single pass. This is fundamentally different from competitors that handle audio as a separate post-processing step.
Why this matters in practice:
- Lip-synced dialogue is baked into the output, not layered. When a character speaks, the model computed the lip movements and audio waveform together. They match because they're the same computation.
- Sound effects align precisely with visual events. A sword hitting steel rings at exactly the frame the contact happens. Footsteps land on beat with each step. No manual sync needed.
- The audio-visual coherence is noticeably superior. If you've ever been in a loud restaurant where the person's lips are slightly out of sync with their voice, you know how distracting even small misalignment is. Seedance 2.0 eliminates that class of problem by design.
- No post-production audio syncing needed. Your output is already a finished, synchronized artifact.
This architecture is the core reason Seedance 2.0 is worth learning separately from other video models — the workflow implications are real.
4. Quick Start — 3 Copy-Paste Prompts
You can use these immediately, with no reference files needed. Each is text-to-video only.
Quick Start Prompt 1: Dramatic Character Moment
A woman in her early 40s with silver-streaked hair, wearing a long charcoal coat, stands at the edge of a cliff overlooking the ocean at dusk. The wind moves her coat. Camera slowly pushes in from medium shot to close-up on her face. She closes her eyes, breathes deeply, then opens them — a slight smile crossing her lips.
Audio: Wind, crashing waves below, no music. The sound builds slightly as the camera closes in, then softens as she smiles.
16:9, 12 seconds. Golden-hour light fading to deep blue at horizon. Cinematic film grain.Quick Start Prompt 2: Product Reveal
A premium mechanical watch, matte black dial with rose gold indices, rests on a dark slate surface. Camera starts wide, then cranes down to an overhead close-up on the face. The second hand sweeps smoothly. A beam of afternoon light crosses the dial during the final three seconds.
Audio: Soft ambient room tone, the faint tick of the movement. No music.
16:9, 10 seconds. Dramatic single-source side lighting from camera left. Shallow depth of field.Quick Start Prompt 3: Energy Sports Sequence
A surfer drops into a massive overhead wave on an overcast morning. Wide tracking shot from the beach as he carves along the face. Scene cuts to a tight tracking shot from the water — spray explodes from the rail. Final shot: wide aerial pull-back as the wave closes out and the surfer kicks off the back.
Audio: Ocean roar, the hiss of the board on water, wind. Distant seagulls. No music.
16:9, 15 seconds. Moody grey skies, deep blue-green water.5. The Multimodal Reference System
This is Seedance 2.0's most powerful differentiator. Reference files are tagged with @ syntax and can influence characters, environments, camera style, motion choreography, and sound design simultaneously.
5.1 Reference Types and Roles
| Input Type | Max Count | What It Controls | Tag Syntax |
|---|---|---|---|
| Images | 9 | Characters, environments, style, props, storyboards | @Image1, @Image2, ... |
| Videos | 3 (15s total) | Camera movement, action choreography, motion rhythm | @Video1, @Video2, ... |
| Audio | 3 (15s total) | Music, dialogue, sound effects, pacing/rhythm | @Audio1, @Audio2, ... |
| Text | Unlimited | Scene direction, story, actions | Direct in prompt |
5.2 Reference Use Cases
| Scenario | Reference Setup |
|---|---|
| Character-driven narrative | @Image1 = character face, @Image2 = outfit, @Image3 = setting |
| Storyboard-to-video | @Image1 = storyboard sketch (text-based shot descriptions) |
| Motion transfer | @Video1 = reference choreography or camera movement |
| Audio-driven | @Audio1 = music track or voiceover, video syncs to rhythm |
| Style transfer | @Image1 = visual style reference, @Video1 = camera style reference |
| Product commercial | @Image1 = product photo, @Image2 = brand style guide, @Audio1 = background music |
5.3 Storyboard Reference (R2V — Reference to Video)
Seedance 2.0 can directly interpret text-based storyboards from reference images:
Refer to the shooting script in @Image1 (storyboard with shot scale, camera movement, visuals, copy). Character from @Image2. Scene from @Image3. Props from @Image4. Create a 15-second healing short film.How to prepare a storyboard image for best results:
- Clearly label each panel with numbers (Panel 1, Panel 2, etc.)
- Include camera direction notes directly on the sketch (e.g., "WIDE → MEDIUM, dolly in")
- Add simple action notes ("character looks left," "hand enters frame")
- Even rough stick-figure sketches work — the model interprets spatial composition and camera framing, not artistic quality
6. Multimodal Reference Masterclass
The @ reference system is what separates a power user from someone just typing text prompts. This section walks through five complete workflow examples — each one demonstrating a different way to build a production-ready video using references.
Workflow 1: Character-Driven Narrative
The goal: Generate a 15-second narrative scene with a specific, consistent character across multiple shots.
What you're uploading:
@Image1— A clear, front-facing, well-lit photo of your character (real person, illustrated character sheet, or AI-generated portrait)@Image2— A full-body costume reference (can be a separate photo, outfit flatlay, or design sketch)@Image3— The environment (a photo of a real location, an illustration of the setting, or an AI-generated scene)
Why this works: The model reads @Image1 for facial structure, @Image2 for clothing details, and @Image3 for spatial context and lighting. It then keeps all three consistent as the character moves through space.
The prompt:
@Image1 is the main character — a woman in her late 20s, use her exact face and expression style. @Image2 is her outfit — replicate the jacket and bag details precisely. @Image3 is the setting — a narrow Kyoto alley at twilight.
The character walks slowly down the alley toward camera (wide establishing shot). Scene cuts to medium shot as she pauses and looks at a lit paper lantern above her. Close-up on her face as warm lantern light catches her eyes.
Audio: Soft footsteps on stone, distant temple bells, gentle ambient street noise. No music.
16:9, 15 seconds. Warm orange lantern light against deep blue twilight. Shallow depth of field on the close-up.Pro tips:
- Use a well-lit, uncluttered character photo — the cleaner the reference, the more faithful the rendering
- If you want multiple characters, assign a separate @Image reference to each one and mention them explicitly in the prompt
- The model preserves identity across all cuts, so feel free to write multi-shot sequences
Workflow 2: Storyboard-to-Video
The goal: Turn a pre-planned shot list or storyboard into a generated video without writing a long text prompt.
What you're uploading:
@Image1— A storyboard image (can be hand-drawn, digital, or even a grid of photo references with handwritten notes)@Image2— Character reference@Image3— Environment reference
Why this works: Seedance 2.0 reads the spatial composition of each storyboard panel and interprets it as shot framing. Handwritten camera notes on the image (arrows for movement, labels for shot size) are processed as direction.
The prompt:
@Image1 contains a 4-panel storyboard:
Panel 1: Wide shot — empty diner exterior, rainy night, neon sign reflected in wet pavement
Panel 2: Medium shot — man enters through glass door, shaking rain off coat
Panel 3: Close-up — his hand wraps around a coffee mug on the counter
Panel 4: Wide shot — he sits alone by the window, looking out at the rain
Character from @Image2. Diner environment aesthetic from @Image3.
Produce this as a continuous 15-second sequence. Let the transitions feel cinematic — slow and weighted.
Audio: Rain on glass, door chime as he enters, coffee cup set on counter, ambient diner sounds. A gentle, melancholy jazz piano begins softly at the final wide shot.
16:9. Cool desaturated tones with warm neon color from the sign.Pro tips:
- Label your panels clearly — "Panel 1:", "Panel 2:", etc. in the prompt
- The model handles the cuts between your panels automatically; you don't need to specify exact cut timings
- If you want specific timing (e.g., "spend 5 seconds on Panel 4"), say so explicitly
Workflow 3: Motion Transfer
The goal: Take a specific camera movement or physical performance from a reference video clip and apply it to your own scene and characters.
What you're uploading:
@Video1— A short clip (up to 15 seconds) demonstrating the camera movement or motion you want to transfer@Image1— Your character or subject reference@Image2— Your setting
Why this works: The model extracts the camera path and motion signature from @Video1 — the timing, speed, arc, and rhythm of movement — and applies it to your subject. You get the motion pattern without the original footage's content.
The prompt:
@Video1 is a reference clip showing a slow 360-degree orbit around a subject at medium height — use this exact camera movement and speed.
@Image1 is the subject: a hand-thrown ceramic vase, ash glaze, about 30cm tall, sitting on a wooden plinth.
@Image2 is the environment: a calm Japanese pottery studio with diffused natural light from shoji screens.
Apply the orbit movement from @Video1 to this scene. The camera orbits smoothly around the vase, completing about 270 degrees in 10 seconds.
Audio: Near silence — the faint creak of the wooden floor, distant birdsong outside. No music.
16:9, 10 seconds. Soft, even white light from shoji screens. No harsh shadows.Pro tips:
- Your reference clip doesn't need to be high quality — the model is reading motion data, not visual style
- If you only want to borrow the camera movement (not the subject's physical performance), say so: "Use the camera path from @Video1 but apply it to the scene described below"
- To borrow a physical performance (a dance move, a martial arts strike, a gymnastics routine), describe it similarly: "Transfer the movement of the performer in @Video1 to the character described below"
Workflow 4: Audio-Driven Video
The goal: Generate a video that synchronizes visually to a music track or voiceover — where the pacing, cuts, and energy follow the audio.
What you're uploading:
@Audio1— Your music track, voiceover, or ambient soundscape@Image1— Visual style or setting reference (optional but helpful)
Why this works: The Dual-Branch architecture means audio and video are computed in relation to each other. When you provide an audio reference, the video pacing, cut rhythm, and energy level respond to it. Fast BPM → more dynamic cuts. Slow, meditative audio → longer, quieter shots.
The prompt:
@Audio1 is a 12-second electronic music track — driving, rhythmic, BPM approximately 128. Use this as the audio track for the video. Let the visual energy and edit rhythm match the track's intensity.
@Image1 is the visual style reference: kinetic urban photography, motion blur, neon-lit streets at night.
Generate a montage sequence: a skateboarder in a lit parking garage, fast grinding shots, trick landings, sparks from grinds. The cuts should land on the beat of @Audio1.
16:9, 12 seconds. High contrast, pushed greens and magentas. Motion blur on fast sections.Pro tips:
- Mention the approximate BPM or emotional energy of your audio track in the text — this helps the model align the visual pacing even more precisely
- For voiceover-driven video, include the full script text in the prompt alongside the audio reference: "The narrator in @Audio1 says: '[your script]'. Generate video that visually illustrates these words."
- Works especially well for music videos, brand montages, and any content where the sound design is the primary creative element
Workflow 5: Product Commercial
The goal: Generate a polished product commercial with brand-consistent visuals and audio.
What you're uploading:
@Image1— Clean product photo (well-lit, isolated or in context)@Image2— Brand style guide, mood board, or visual aesthetic reference@Audio1— Background music or brand audio signature
Why this works: The model treats @Image1 as the hero subject, reads @Image2 for color palette and visual language, and uses @Audio1 to set the emotional register of the commercial.
The prompt:
@Image1 is the product: a premium olive oil in a dark glass bottle with a hand-lettered label. This is the hero of the video — it must be clearly visible and beautifully rendered throughout.
@Image2 is the brand style: artisan Italian, warm terracotta and cream tones, rustic stone surfaces, Mediterranean light.
@Audio1 is the background track: a gentle acoustic guitar piece with a warm, unhurried feel. Use it as the audio foundation.
Shot sequence:
— Wide shot: The bottle sits on a worn stone kitchen counter beside fresh herbs and tomatoes. Soft morning light from a nearby window.
— Slow dolly-in to medium close-up on the label.
— Overhead shot: A hand tips the bottle, pouring a thin golden stream of oil into a white ceramic bowl. The oil catches the light.
— Final wide shot: The bottle, bowl, and ingredients arranged as a still life. The camera locks off.
Audio: @Audio1 throughout. Add natural sounds — the faint pour of oil, a gentle clink as the bottle is set down. No voiceover.
16:9, 15 seconds. Warm morning light, soft shadows, slightly matte color treatment.Pro tips:
- Always specify that the product is "the hero" — this signals the model to keep it sharp, well-lit, and compositionally central
- For text on packaging (labels, logos), include a high-resolution product photo as @Image1 and mention the text details explicitly
- The reveal → detail → hero shot structure works reliably for products: start wide, move in close to show texture/quality, end with a beauty shot
7. Prompt Formula: Subject → Action → Camera → Style
Every strong Seedance 2.0 prompt has four components. Think of it as a recipe — you can riff on the proportions, but you shouldn't skip ingredients.
Subject → Action → Camera → Style| Component | Description | Example |
|---|---|---|
| Subject | Who/what appears (include age, material, type) | "A woman in her 30s with short black hair, wearing a denim jacket" |
| Action | Single clear verb in present tense | "walking slowly through a rainy street" |
| Camera | Shot size + movement + angle + lens type | "Medium shot, slow dolly-in, eye level, 50mm lens" |
| Style | Visual anchor + lighting + color treatment | "Cinematic film look, tungsten lighting, desaturated teal tones" |
Why the order matters: The model reads top-to-bottom and builds context progressively. "Subject" establishes what exists in the world. "Action" animates it. "Camera" tells the model how to observe it. "Style" sets the aesthetic register. Putting camera instructions before subject description can cause the model to assign camera behavior before it knows what it's pointing at.
Building a prompt step by step:
- Start with subject: "A red fox kit, around 6 weeks old, fluffy winter coat"
- Add action: "cautiously stepping across a frozen pond"
- Add camera: "Low angle tracking shot, just above ice level, following alongside"
- Add style: "Documentary nature film aesthetic, overcast natural light, muted cool tones, slight vignette"
Complete: "A red fox kit, around 6 weeks old with a fluffy winter coat, cautiously steps across a frozen pond. Low angle tracking shot, just above ice level, following alongside. Documentary nature film aesthetic, overcast natural light, muted cool tones, slight vignette."
8. Camera Movement Vocabulary
Seedance 2.0 has genuinely director-level camera control. The model responds to specific cinematography terminology — and being specific here is the single easiest way to upgrade your output quality.
| Movement | What This Looks Like | When to Use | Prompt Language |
|---|---|---|---|
| Dolly-in | The camera physically moves forward toward the subject, like walking toward someone. The perspective compresses slightly. It feels intimate and focused. | Revealing emotion, building tension, drawing attention to a specific detail. Works brilliantly for face reveals and product close-ups. | "Slow dolly-in on character's face" |
| Pan | The camera stays in place but rotates left or right, like turning your head. Space unfolds to the side. | Revealing what's next to the subject — a landscape, a crowd, a second character entering frame. | "Slow pan right revealing cityscape" |
| Tracking | The camera moves alongside a moving subject, staying at roughly the same distance. Like a car driving next to a runner. | Following action. Running, walking, driving, skating — anything in motion that you want to stay with. | "Tracking shot alongside running athlete" |
| Crane | The camera rises or falls through vertical space. Dramatic and theatrical — it reveals scale. | Big reveals, establishing grandeur, the "pull back to show the whole world" feeling. | "Crane down from sky to ground level" |
| Handheld | Slight, organic shake — it feels like a real person is holding the camera. Intimate and raw. | Documentary feel, personal moments, chaos and urgency. Anything where polish would feel emotionally wrong. | "Shaky handheld follow, natural movement" |
| Gimbal/Steadicam | Smooth, floating movement — the camera glides without shake. Polished and cinematic. | Professional narrative work, commercial content, any scene where you want movement without the roughness of handheld. | "Stabilized gimbal tracking, polished" |
| Push-in | Similar to dolly-in but often slower and more subtle — the world slowly closes in around the subject. | Building dread, dawning realization, slow-burn emotional moments. | "Slow push-in as character realizes..." |
| Pull-back | The camera moves away from the subject, gradually revealing more of the world around them. The subject gets smaller as context grows. | Scale reveals, loneliness, the "small person in a big world" feeling. Also great for ending shots. | "Camera pulls back to reveal vast landscape" |
| Dutch angle | The camera is tilted on its axis — the horizon is diagonal. Everything feels slightly wrong. | Psychological unease, villains, something-is-off moments, fever dreams. Don't overuse. | "Tilted Dutch angle, 15 degrees" |
| Orbit | The camera circles around the subject in an arc, from any height. Like walking around something to examine it. | Product reveals, showcasing a 3D object, showing a character from all sides, impressive architectural shots. | "Camera orbits around the product" |
| POV | The camera becomes the character's eyes — first-person perspective. | Immersive experiences, gaming aesthetic, following a character's visual attention. | "POV shot from the driver's seat" |
Multi-Shot Cuts
Use the keyword "lens switch" or "scene cuts" to signal a cut within a single generation. This is one of Seedance 2.0's most powerful and unique features.
A samurai faces his opponent. Camera slowly pushes in. Scene cuts to a fast-panning profile shot. They charge and clash in ultra-slow motion — swords meet, sparks fly, bamboo leaves scatter. Normal speed resumes as they land back-to-back.Tips for multi-shot sequences:
- Describe each shot segment in clear temporal order ("first... then... finally...")
- Use time/speed cues:
"ultra-slow motion","normal speed resumes","time-lapse of...","freeze frame for a moment" - Transitions can be explicit:
"hard cut to","cross-dissolve into","smash cut to" - Aim for 2–4 shots in 15 seconds — more than that feels rushed
9. Audio Direction
This is where Seedance 2.0 pulls away from the field. Because audio is generated in the same pass as video, you can direct sound with the same specificity as a foley artist on a real production. Here's how to get the most from each audio category.
9.1 Dialogue
The model generates lip-synced speech. You can have multiple speakers, emotional delivery direction, whispers, shouts, and even singing.
- Supports multiple speakers
- Include emotional direction with dialogue (tone, pace, volume)
- Supports Chinese dialects, traditional opera, and singing
- For multilingual scenes, specify the language
Examples:
The woman turns to camera and says with quiet confidence, "Everything changes today."Two old men argue loudly across a chess table. First man: "You always play the same way, you predictable fool." Second man, leaning back and laughing: "And you always fall for it."A child whispers to the camera, eyes wide: "I think there's something in the basement." Beat of silence. Then a creak from below.The chef narrates in a warm, unhurried voice: "The secret is patience. Always patience." As she speaks, her hands work the dough.A street performer sings the opening bars of a folk song — a capella, clear and true, the melody floating above the ambient market noise.9.2 Sound Effects and Foley
Seedance 2.0 captures subtle foley nuances that most generators miss — the kind of sounds a Foley artist would record separately in post.
Things it handles well:
- Fabric rustling, glass scratching, bubble wrap popping
- Impact sounds, footsteps on various surfaces (wood, gravel, concrete, wet stone)
- Environmental sounds (rain, wind, traffic, forest)
- Material interactions (metal on metal, liquid on glass, paper tearing)
Examples:
SFX: The scratch of frosted glass, soft rustling of plush fabric, gentle tapping on acrylic.Audio: Heavy boots on wet cobblestones — each step a crisp splash. Raincoat fabric against itself as the figure walks. Distant thunder, fading.SFX: A match struck, the hiss of ignition, the soft pop of a gas burner catching. Then the steady blue hiss of the flame.Audio: Leather gloves pulled on — the tight snap of the cuff. Metal zippers. The click of a briefcase clasp. These sounds in sequence, intimate and deliberate.SFX: A vinyl record dropping onto the turntable — the faint crackle of static, then the warm hiss before the music begins.9.3 ASMR
Seedance 2.0 excels at ASMR-style close-up audio — the soft, textural, hyper-present sounds that make viewers lean in. This is a genuinely strong use case.
The model captures:
- Tactile material interactions (peeling, unboxing, folding)
- Near-silence environments with a single sound source
- Layered soft sounds (whisper + fabric + ambient)
Examples:
Close-up of hands slowly unwrapping a luxury chocolate bar. Gold foil crinkles and peels back to reveal dark chocolate with a perfect snap. The chocolate is placed on a ceramic plate with a soft tap.
Audio: ASMR style — crisp foil crinkling, satisfying chocolate snap, gentle ceramic tap. No music, no voice. Room ambient only.
9:16, 10 seconds. Warm, soft overhead lighting. Shallow depth of field.Close-up of hands peeling the protective film off a new phone screen. The film comes away slowly, smoothly, catching the light.
Audio: ASMR — the slow, clean peel of the protective film, a gentle crinkle, the satisfying reveal of the pristine screen beneath. Room ambient: nearly silent.
9:16, 12 seconds. Bright, clean white-box lighting.A person slowly stirs a mug of hot tea. The spoon circles in the liquid, tapping softly against the ceramic sides.
Audio: ASMR — the gentle clink of spoon on ceramic, the soft sound of liquid moving, a quiet breath. Steam ambient. No music.
1:1, 8 seconds. Soft morning window light.Hands fold a crisp white linen napkin with deliberate, precise movements — each fold clean and sharp.
Audio: ASMR — the crisp rustle of starched linen, each fold landing with a soft pressing sound. Near silence otherwise.
9:16, 10 seconds. Flat, even studio light. White surface.9.4 Action SFX
For combat, sports, and high-energy sequences, you want to specify the specific impact layers — the model can generate complex sound design if directed.
Examples:
Audio: The ring of steel on steel as swords clash, the grunt of physical exertion, bamboo splintering, the whoosh of a missed strike, then sudden silence as one fighter wins. No music.Audio: Racing engine at high rev, tire squeal on tarmac, wind buffeting the car at speed, the crack of gravel against the wheel well on a corner. No music — pure mechanical SFX.Audio: The thwack of a basketball on hardwood, the swish of net, crowd reaction building from murmur to roar. Sneakers squeaking. An echoing sports hall.9.5 Ambient / Atmosphere
Ambient audio is often underused but critically important. It tells the viewer where they are without a single word.
Examples:
Audio: Cinematic orchestral swell building through the sequence. Deep bass, warm strings.Audio: A Tokyo street at rush hour — layered train announcements, bicycle bells, heeled shoes on tile, the hiss of bus doors. Dense and alive but not chaotic.Audio: Deep forest ambience — wind in pine canopy, a distant stream, occasional bird call. No human sounds. A sense of remoteness.Audio: Hospital corridor — fluorescent hum, soft footsteps, a distant PA announcement, the rubber squeak of shoe soles on linoleum. Institutional quiet.Audio: Open ocean from a small boat — rhythmic water slap against the hull, rope creaking, wind, and nothing else. Vast.9.6 Music-Driven Sequences
When music is the emotional driver, describe it with the same specificity you'd give a film composer.
Examples:
Audio: A melancholy solo piano — sparse, minor key, unhurried. Each note allowed to decay fully before the next. The music begins on the first shot and fades at the end.Audio: Driving electronic music, 130BPM, with a hard drop at the 8-second mark. Visual cut rhythm should align with the beat. High energy throughout.Audio: Traditional Korean gayageum — a single instrument, contemplative and formal. The music carries the emotional weight of the scene without dialogue.10. Key Prompting Best Practices
10.1 Core Rules
- One dominant action per shot. Complex scenes with multiple simultaneous actions reduce quality. If you need complexity, use multi-shot structure.
- Use positive constraints. Seedance 2.0 does NOT support negative prompts. Describe what you want, not what to avoid. Instead of "no harsh shadows," write "soft, diffused light from an overcast sky."
- Leverage references over text. When you have visual/audio assets, use
@tags — they're more precise than text descriptions. A character photo is worth a thousand words of description. - Camera language matters. The model excels at executing specific camera directions. Vague directions ("nice camera movement") produce mediocre results. Specific directions ("slow dolly-in, eye level, 50mm lens") produce director-quality output.
- Keep prompts structured. Subject → Action → Camera → Style order produces best results.
- Always specify aspect ratio. The model needs to know whether to compose for wide screen (16:9), portrait (9:16), or square (1:1). Don't leave this out.
- Always direct the audio. Sound is generated in the same pass. If you don't describe it, the model makes its own choices — which may not serve your vision.
10.2 For Complex Multi-Shot Narratives
- Write detailed action descriptions with temporal flow — what happens first, then, finally
- Include camera transitions:
"scene cuts abruptly","switches to ultra-slow motion","normal speed resumes" - Describe emotional beats:
"tension builds","relief washes over","moment of realization" - Use 2–4 shots for a 15-second clip — give each shot enough time to breathe
10.3 For Product/Commercial Work
- Upload product image as @Image1 and brand style guide as @Image2
- Include specific camera choreography: reveal → detail → hero shot
- Specify lighting progression across the clip
- Name the product or describe it very specifically — the model should have no doubt what the "hero" is
10.4 For Character Consistency
- Upload character reference as @Image1 (clear, front-facing, well-lit)
- The model maintains identity across multi-shot sequences
- For multiple characters, assign separate reference images
- Reference the character by name in your prompt text: "The woman from @Image1..." — this helps the model track who is who
11. Video Extension
Seedance 2.0 supports continuation from existing clips — you can extend a generated video forward in time.
Extend the video. The man on horseback gallops toward a blossoming tree, reaches up to break a branch of cherry blossoms. Other riders appear from behind the hill, dismount in a circle, and present flowers.Tips:
- Reference the existing content for continuity — briefly describe what just happened so the model has context
- Describe new actions and camera changes clearly
- The model preserves subject appearance and scene style across the extension
- Think of it as writing the next scene in a screenplay — you know what happened, now tell the model what happens next
- Audio continues naturally through extensions; you can specify if the audio character should change
12. Video Editing
Targeted modifications to existing generated clips:
- Change specific characters, actions, or storylines while preserving the rest of the composition
- Modify visual elements (color, lighting, specific objects) while keeping the camera work intact
- Alter pacing or add action beats to an otherwise good clip
- The model maintains logical flow and visual consistency across edits
Example editing prompts:
Edit the clip: Replace the character's red jacket with a dark navy coat. Keep everything else — the setting, camera movement, and timing — exactly the same.Edit the clip: Change the final shot from day lighting to a dramatic sunset. Warm the color grade across the whole clip slightly to match.13. Shot Size Guide
Understanding shot size helps you pair the right camera movement with the right framing.
| Shot Size | What It Includes | Pairing Guidance |
|---|---|---|
| Wide / Establishing | Subject + most or all of environment | Slow dolly or locked-off. Avoid fast pans. Establishes space and context. Lets the viewer orient themselves. |
| Medium | Subject from roughly waist up | Handheld feels personal and intimate; gimbal feels polished and professional. Includes subject + enough context to understand their situation. |
| Close / Close-Up | Face, hands, or a specific detail | Tiny push-ins work best; pans can feel jarring at this scale. Focuses on detail and emotion. Shallow depth of field shines here. |
| Extreme Close-Up (ECU) | A single eye, a fingertip, a texture | Nearly always locked-off or micro push-in. For ASMR, product texture, and intense emotional moments. |
14. Example Prompts (Production-Ready)
Cinematic Action — Bamboo Forest Duel
Two martial artists face off in a bamboo forest. Camera slowly pushes in on their eyes. Lens switch to a wide tracking shot as they charge at each other. The clash happens in ultra-slow motion — swords meet, sparks fly, bamboo leaves scatter. Normal speed resumes as they land back-to-back. The straw-caped fighter's hat splits open.
Audio: Tense silence, then the ring of steel on steel. Bamboo creaking. A single shakuhachi flute note holds over the slow-motion moment.
16:9, 15 seconds.Cinematic Action — Urban Rooftop Chase
A figure in dark clothes sprints across a flat rooftop at night, city lights sprawling below. Wide tracking shot from a parallel rooftop. Scene cuts to a low tracking shot at leg level — boots hitting concrete, coat flying. She reaches the edge and leaps without hesitation. Slow-motion as she drops, arms out. Cut to: she lands in a roll on the next rooftop, comes up running.
Audio: Heavy footfalls on gravel roofing, wind at height, the massive whoosh of the leap, the crack of the landing roll, footsteps again. No music.
16:9, 15 seconds. Blue-grey city-lit night. High contrast.Cinematic Action — Fighter Training Montage
A boxer trains alone in a dim gym at 5am. Wide shot of the empty gym as he skips rope (establishing). Scene cuts to medium tracking shot as he moves to the heavy bag — combinations, fast, rhythmic. Tight close-ups intercut: gloves connecting, sweat, determination in the eyes. Final shot: he stops, leans on the bag breathing hard, then looks up with quiet resolve.
Audio: Rope cutting air, feet on the mat, the heavy bag receiving each combination, exertion breathing. The gym's HVAC hum underneath. No music — the silence makes it harder.
16:9, 15 seconds. Tungsten bulb practicals hanging from ceiling. Deep shadows, warm highlights.Product Commercial — Wireless Speaker
@Image1 is the product (a premium wireless speaker in matte white). @Image2 is the brand style guide (minimalist, warm tones, Scandinavian aesthetic). @Audio1 is a 10-second deep bass music track.
Camera starts with a wide shot of a sunlit Scandinavian living room. Dolly-in toward the speaker on a wooden shelf. The music from @Audio1 begins playing as the camera approaches. Medium close-up reveals the speaker's texture and form. Final shot: overhead angle, the speaker centered on the shelf with plants on either side.
16:9, 12 seconds. Warm afternoon light.Product Commercial — Luxury Fragrance
A dark glass perfume bottle with a gold stopper stands on a black marble surface. A single droplet of water falls from above and bursts in slow motion against the marble beside the bottle.
Camera: Overhead wide establishing shot → slow dolly down to eye-level medium close-up as the droplet falls → extreme slow-motion of the burst → final locked-off medium shot of the bottle, perfectly still.
Audio: The musical tone of glass, the soft impact of the droplet, a quiet resonant hum that fades to silence. No music.
16:9, 12 seconds. Dramatic side lighting from camera right — deep shadows, the bottle partially lit. Black and gold throughout.Product Commercial — Artisan Coffee Brand
@Image1 is the product: a kraft paper bag of single-origin coffee with a hand-stamped label. @Image2 is brand aesthetic: warm, artisanal, independent coffee shop morning light.
Shot sequence:
— Wide shot of a wooden counter, morning light streaming through a window. The bag sits beside a ceramic grinder and a pour-over setup.
— Slow dolly-in to the bag, label filling frame.
— Cut to medium close-up: hands open the bag, and a breath of steam and aroma seems to rise (suggest this in the visual warmth).
— Wide shot: the full pour-over ritual begins, water falling in a perfect thin stream.
Audio: Morning ambient — birds outside, the quiet of early hours, the gentle gurgle of water through the grounds. No music. Almost meditative.
16:9, 15 seconds. Warm golden morning light. Muted, natural color palette.ASMR — Chocolate Unwrap
Extreme close-up of hands slowly unwrapping a luxury chocolate bar. Gold foil crinkles and peels back to reveal dark chocolate with a perfect snap. The chocolate is placed on a ceramic plate with a soft tap.
Audio: ASMR style — crisp foil crinkling, satisfying chocolate snap, gentle ceramic tap. No music, no voice. Room ambient only.
9:16, 10 seconds. Warm, soft overhead lighting. Shallow depth of field.ASMR — Bookshop Sounds
Close-up sequence in an old bookshop. Hands slowly pull a large hardcover book from a packed shelf — the resistance, then release as it comes free. Pages turn slowly, one by one. The spine is opened and pressed flat with a soft creak.
Audio: ASMR — the soft shush of the book sliding out, the papery rustle of pages turning, the leather creak of the spine, ambient old-book-smell silence (suggested through the visual warmth). No music.
9:16, 12 seconds. Warm lamplight, deep shadows between shelves.ASMR — Rain Window
Extreme close-up of raindrops running down a glass window, backlit by city lights at night. A hand reaches up and traces one raindrop's path with a fingertip.
Audio: ASMR — rain pattering on glass, the soft squeak of fingertip on wet glass, ambient rain on the roof above. No music.
9:16, 15 seconds. City lights blurred through the rain — bokeh reds and yellows. Intimate and still.Narrative — The Letter
A woman in her 60s sits at a kitchen table in afternoon light. She holds an envelope, unopened, in both hands. Camera holds on her in a medium shot — we can see she's been sitting here a while. She finally opens it, slowly, with a letter opener. Pulls the paper out and begins to read. Her expression shifts almost imperceptibly — relief? Grief? Both.
Audio: Clock ticking, paper unfolding, her quiet breathing. Traffic faint outside. No music.
16:9, 15 seconds. Dusty afternoon light through a curtained window. Warm, slightly overexposed. Film grain.Narrative — First Day
A teenage boy in a new school uniform sits alone in a cafeteria, tray untouched, watching other students with that particular combination of longing and armor. Wide shot initially — the cafeteria large around him. Slow push-in to medium shot. He takes a breath. Picks up his fork. Starts eating, not looking up.
Audio: Cafeteria din — voices, trays, the PA system. But in a slight way the sound is softened around him, the loneliness creating its own quiet. No music.
16:9, 12 seconds. Fluorescent cafeteria lighting, which is harsh and unkind and correct.Narrative — The Return
A man steps off a train onto a small-town platform. He's in his 40s, carrying too much — bags, a coat over one arm. The platform is empty except for one person waiting at the far end. He spots them. They spot him. A long beat of stillness — distance between them. Then he starts walking, slowly at first, then not slowly at all.
Audio: Train doors closing behind him, hiss of brakes, the crunch of gravel underfoot. Wind. No music until the last two seconds — a single held string note.
16:9, 15 seconds. Overcast autumn light. Leaves on the platform.Sports — Basketball Clutch Moment
A figure skating pair performing a dramatic lift sequence on an empty rink. The woman is thrown into a triple spin while the man holds a steady base. Camera tracks their movement from ice level. They land in perfect synchronization. The audience erupts.
Audio: Blades scraping ice, wind from the spin, crowd gasping then cheering. Orchestral crescendo on the landing.
16:9, 15 seconds. Cool blue-white rink lighting with warm spotlights.Sports — Marathon Final Mile
A marathon runner enters the final straightaway. She's exhausted but accelerating — a closing sprint on depleted legs. Tracking shot from alongside at road level. The finish line banner visible ahead. She breaks the tape. Immediately crumbles to her knees.
Camera: Side tracking shot → wide shot as she crosses → tight close-up on her face immediately after.
Audio: Her labored breathing, crowd cheering, the crack of the timing tape, then the noise of the crowd washing over everything.
16:9, 15 seconds. Bright midday sun. Sweat-soaked kit. Raw and real.Sports — Free Solo Climbing
A climber on a sheer granite face with no rope, hundreds of meters above the valley floor. She reaches for a hold, tests it, moves. The camera is on the wall beside her, tracking her movement — a technically impossible shot that AI makes possible.
Scene cuts to a wide shot from a drone perspective: she is a small figure on an immense wall of stone.
Audio: Wind at altitude, the dry scrape of chalk on rock, her controlled exhale. Silence otherwise. The silence is the point.
16:9, 15 seconds. Harsh alpine sun, deep shadows in the cracks. No music.Nature — Arctic Fox Hunt
A red fox kit, around 6 weeks old with a fluffy winter coat, cautiously steps across a frozen pond. Low angle tracking shot, just above ice level, following alongside. Each step is tentative — the ice creaks.
Scene cuts to close-up on its face: ears forward, eyes fixed on something out of frame. Then it pounces — straight down through the snow. Comes up with a vole.
Audio: Ice creaking under each step, the silent pause before the pounce, then the muffled crunch of impact and the trill of the successful hunt. Wind in sparse trees nearby.
16:9, 15 seconds. Overcast grey-white winter light. Muted blues and the orange of the fox fur.Nature — Storm Formation
Time-lapse of a thunderstorm forming over open plains. Clouds build from nothing — white cumulus climbing, darkening, anvilling out at altitude. Beneath them, the light turns green-grey and wrong. A single bolt drops from the base of the storm.
Audio: Wind building from still to rushing over the time-lapse, the distant bass roll of thunder, a close lightning crack that makes the viewer flinch.
16:9, 15 seconds. The color of the sky is the main visual element — let it be strange and beautiful.Nature — Tidal Pool
Extreme close-up exploration of a tidal pool. A small crab moves between anemones. A sea slug, improbably beautiful, moves across a rock. The camera drifts slowly, discovering each creature — a nature documentary aesthetic at macro scale.
Audio: Gentle water movement in the pool, the sound of the larger ocean beyond, the occasional plop of a small wave refilling the pool. No music. Natural Documentary.
16:9, 15 seconds. Bright tide-pool light — clear water catching sun, dappled on the bottom.Comedy — The Dog and the Stick
A golden retriever in a park is presented with a stick that is approximately 12 feet long. The owner holds one end. The dog looks at the stick. Looks at the owner. Back at the stick. Attempts to pick it up, miscalculates badly, bonks a nearby pigeon. Sits and looks off-camera with the expression of someone who has made mistakes and is comfortable with that.
Audio: Park ambience, the thwack of the dog attempting to grab the stick, the startled pigeon, a brief dog-shaking-head sound. The owner, heard but not seen, suppressing laughter. No music.
16:9, 12 seconds. Bright park day. Slightly comedic handheld energy.Comedy — The Espresso Machine Wins
A man in pajamas confronts his espresso machine at 6am. He attempts every button combination. The machine makes increasingly aggressive sounds. Steam goes somewhere unexpected. He holds a completely empty cup. Stares at it. Stares at the machine. The machine emits one final triumphant hiss.
Audio: The buttons clicking, the machine's mechanical protests, unexpected steam venting, the anticlimactic silence of the empty cup. No music — just the sounds of defeat.
16:9, 10 seconds. Harsh morning kitchen light — fluorescent, merciless.Comedy — Cat Meeting a Mirror
A tabby cat discovers a full-length mirror. Its initial encounter involves: staring, slow approach, full puff, sideways intimidation walk, sniffing the glass, and ultimately concluding the other cat is boring. The cat walks away mid-confrontation.
Audio: Near silence with soft paw pads on the floor, the subtle sounds of a deeply confused animal, and eventually footsteps retreating with dignity intact.
1:1, 12 seconds. Natural apartment light. Slightly elevated camera looking down on the unfolding situation.Character Consistency — Forest Discovery
@Image1 = Main character (a 10-year-old girl with braided pigtails and a yellow raincoat).
@Image2 = Forest setting reference.
The girl discovers a hidden path through the foggy forest. She pushes aside ferns (medium tracking shot). Lens switch to close-up of her face — eyes widen with wonder. Lens switch to her POV — a glowing treehouse is visible through the mist. She runs toward it (wide shot from behind, camera following).
Audio: Soft footsteps on wet leaves, bird calls, a magical shimmer sound when the treehouse appears. Whimsical orchestral undertone.
16:9, 15 seconds. Overcast diffused light, misty atmosphere.Character Consistency — The Detective
@Image1 = Character reference: a man in his 50s, grey stubble, worn brown coat, tired eyes that see too much.
He stands in the rain outside a closed bar, collar up, studying the door. (Wide shot, locked-off.) Scene cuts to his POV: a light on in the upstairs window, a shadow moving behind the glass. Scene cuts back to his face in close-up — no readable expression, but he's already decided something. He pushes the door open and goes in.
Audio: Rain on awning, wet footsteps, the squeak of the door. Inside: a distant jukebox, low voices. No score.
16:9, 15 seconds. Night. Cold blue rain light. Warm bar light spilling under the door.Character Consistency — Grandparent and Child
@Image1 = Elderly grandmother character reference — white hair, warm dark eyes, floral apron.
@Image2 = Child character reference — 8-year-old boy, freckles, striped shirt.
The grandmother teaches the boy to roll dumplings at the kitchen table. He tries to fold one and it falls apart. She gently redoes it, his hands over hers. He tries again. This time it holds. He holds it up triumphantly. She claps.
Audio: Kitchen sounds — the gentle slap of dough, voices in warm murmur (no specific dialogue), her soft laugh when his first one fails, his genuine delight when the second one works.
16:9, 15 seconds. Warm kitchen light, afternoon. Flour on every surface.Storyboard-to-Video (R2V) — Café Scene
Refer to the shooting script in @Image1 (hand-drawn storyboard with 4 panels showing: 1. Wide establishing shot of café exterior, 2. Medium shot of woman entering, 3. Close-up of her ordering at counter, 4. Wide shot of her sitting by the window with coffee).
Character face and appearance from @Image2. Café environment style from @Image3.
Create a 15-second sequence following the storyboard. Warm, golden afternoon light throughout. Gentle background café ambience with soft acoustic guitar.
16:9.15. Common Mistakes → Better Versions
These are the eight most common ways prompts underperform — and exactly how to fix them.
Mistake 1: Using a negative prompt
Why it fails: Seedance 2.0 doesn't support negative prompts. The model may interpret "no people" as a cue to add people, or ignore the instruction entirely.
Why it works: Describing the presence of emptiness ("pristine," "deserted," "empty sand") is positive language. You're telling the model what IS there, not what isn't.
Mistake 2: Overloading one shot with too many actions
Why it fails: Multiple competing simultaneous actions confuse the model's spatial and temporal reasoning. Quality degrades when competing motions need to be generated in the same frame.
Why it works: Each shot has one dominant visual action. The sense of multitasking chaos is created through editing and sound design, not by literally depicting everything at once.
Mistake 3: Vague camera direction
Why it fails: "Nice camera movement" tells the model nothing useful. It will make a generic choice that may not serve the scene.
Why it works: Specific shot size, movement type, speed, and duration. The reveal structure creates a deliberate narrative beat.
Mistake 4: Not specifying aspect ratio
Why it fails: The model defaults to a ratio that may not suit your platform. If you're making a TikTok, 16:9 horizontal video is useless.
Why it works: You've told the model exactly what canvas to compose for.
Mistake 5: Ignoring audio direction
Why it fails: The model will generate some audio, but without direction it may produce something generic or tonally wrong for your vision.
Audio: Deep ocean swells, wind building, the long rumble of thunder, then a crack as lightning strikes closer. The rain begins — not a dramatic moment, just steady rain joining the other sounds. No music."
Why it works: You've specified the emotional character of the audio (building, not suddenly dramatic), the specific sound elements, and the absence of music — which is itself a creative choice.
Mistake 6: Not tagging uploaded references
Why it fails: Without the @Image1 tag, the model may not properly connect your uploaded reference to the scene described.
Why it works: The explicit tag creates an unambiguous link between the uploaded file and its role in the output.
Mistake 7: Expecting more than 15 seconds in a single generation
Why it fails: Seedance 2.0's maximum single generation is 15 seconds. Longer prompts will be truncated or the model will rush through them.
Why it works: You're working with the model's architecture rather than against it.
Mistake 8: Assigning references without explaining their roles
Why it fails: The model doesn't know whether @Image2 is a character, an environment, a style reference, or a prop. It will make its own interpretation, which may be wrong.
Why it works: Each reference has a clearly defined job. The model can execute each role precisely without guessing.
16. Seedance 2.0 vs. Veo 3.1 — Practical Comparison
Both Seedance 2.0 and Veo 3.1 are serious, professional-grade video generation models with native audio. But they make different choices, and understanding those choices helps you pick the right tool for the right job.
Feature Comparison
| Feature | Seedance 2.0 | Veo 3.1 |
|---|---|---|
| Max Duration | 15 seconds | 8 seconds |
| Native Audio | Yes (dual-channel stereo) | Yes |
| Max Resolution | 2K (2048px) | 4K |
| Image References | Up to 9 | Up to 3 |
| Video References | Up to 3 | Not available |
| Audio References | Up to 3 | Not available |
| Multi-Shot in Single Gen | Yes — via "lens switch" / "scene cuts" | Via timestamp prompting (more limited) |
| Physics Accuracy | Excellent | Good |
| Usable Output Rate | 90%+ (reported) | High |
When to Choose Seedance 2.0
Seedance 2.0 is the better choice when:
- You need longer than 8 seconds — there's simply no other option here; Seedance's 15-second ceiling is a genuine competitive advantage
- You want to tell a story with multiple shots in a single generation — the multi-shot capability via "lens switch" is unique and powerful
- You have existing assets (character photos, style references, audio tracks, motion reference clips) — the 9-image + 3-video + 3-audio reference system is unmatched
- Audio complexity matters — dialogue with multiple speakers, specific foley details, ASMR precision, music-synchronized cuts
- You're doing product commercial work where brand asset references (logo, style guide, product photo) need to be faithfully reproduced
- Character consistency across cuts is required — the model maintains identity without additional prompting
- You want the physics-heavy shot: liquid pouring, fabric in wind, practical lighting behavior, complex particle effects
- ASMR or tactile content — Seedance 2.0's audio engine is particularly well-tuned for this use case
When to Choose Veo 3.1
Veo 3.1 is the better choice when:
- 4K resolution is a hard requirement — Veo's 4K ceiling versus Seedance's 2K matters for large-screen or broadcast work
- Your shot concept is contained within 8 seconds and doesn't require multi-shot editing
- You're working primarily with text-to-video (no reference assets to manage) and want to iterate quickly
- Google ecosystem integration matters — if you're inside the Google/YouTube workflow
- The prompt calls for a single, clean, continuous camera move with a simple scene — both models do this well, but Veo's workflow is sometimes simpler for straightforward single-shot prompts
- You're comparing outputs between models for a client deliverable and want to see two perspectives on the same prompt
Practical Decision Framework
Do you need > 8 seconds?
└─ YES → Seedance 2.0 (no alternative)
└─ NO → Continue
Do you have reference assets (images, video clips, audio)?
└─ YES (many) → Seedance 2.0 (much more powerful reference system)
└─ YES (1-3 images) → Either model works; Seedance gives you more flexibility
└─ NO → Continue
Do you need 4K resolution?
└─ YES → Veo 3.1
└─ NO → Continue
Is multi-shot editing within a single generation important?
└─ YES → Seedance 2.0
└─ NO → Either model; evaluate based on current output quality for your specific scene typeThe Same Prompt, Two Models — What to Expect
Prompt: "A woman stands at the edge of a cliff at golden hour, the ocean below. She opens her arms. The camera pulls back slowly to reveal the vast coastline."
- Seedance 2.0: Will likely produce accurate physics (hair and clothing movement in the coastal wind, realistic light behavior on the cliff face), audio (wind, waves, ambient sound), and execute the pull-back with cinematic weight. 2K resolution, up to 15 seconds.
- Veo 3.1: May produce higher-resolution output (4K), and may handle the visual composition with excellent fidelity. 8-second maximum means the reveal will be compressed.
Neither is definitively better — the right choice depends on your resolution needs and how much time the scene requires to breathe.
17. Prompt Building Workshop
This section walks you through building three prompts from scratch — starting with a basic idea and layering in specificity until you have something production-ready. It's the most direct way to internalize the Subject → Action → Camera → Style structure.
Workshop Exercise 1: "The Old Lighthouse"
Starting idea: A lighthouse at night.
Step 1 — Add a subject with specificity:
"A tall stone lighthouse, 19th-century construction, white with a red stripe, standing on a rocky headland."
Step 2 — Add action:
Step 3 — Add camera:
Step 4 — Add style:
Step 5 — Add audio:
Step 6 — Add technical specs:
Complete prompt:
A tall stone lighthouse, 19th-century construction, white with a red stripe, stands on a rocky headland. The lamp room rotates slowly, its beam sweeping through low-lying sea fog. Wide locked-off shot from the waterline below, looking up. The camera holds completely still — the only movement is the rotating beam.
Audio: A low foghorn in the distance, waves on rocks below, wind. The beam passes twice in the duration of the shot. Near silence otherwise.
16:9, 12 seconds. Deep blue-black night, warm amber light within the lamp room, cold fog cutting the beam into a visible shaft.Workshop Exercise 2: "The Announcement"
Starting idea: Someone getting exciting news.
Step 1 — Add specificity to the subject and situation:
Step 2 — Add action with emotional specificity:
Step 3 — Add camera that serves the emotion:
Step 4 — Add style:
Step 5 — Add audio:
Complete prompt:
A man in his early 30s sits at a kitchen table with a laptop open. His morning coffee steams beside him. He's been waiting. He reads something on screen. His expression moves — disbelief, then the realization it's real, then a laugh that starts as a breath through his nose and builds to something full.
Camera: Medium shot, locked-off at table height. The camera doesn't move. His face is the only thing that changes.
Audio: Morning kitchen quiet — refrigerator hum, distant birds, morning street sounds. His quiet. Then the laugh. No music.
16:9, 12 seconds. Real morning kitchen light — slightly overexposed by the window, warm and honest.Workshop Exercise 3: "The Training Montage (Product)"
Starting idea: A gym shoe commercial.
Step 1 — Add reference structure:
Step 2 — Build a shot sequence:
Step 3 — Build the audio:
Step 4 — Add technical specs and pace notes:
Complete prompt:
@Image1 is the product: a high-performance trail running shoe in electric blue and white. Show it faithfully throughout.
@Image2 is the visual style: high-contrast, gritty athletic photography. Real conditions, real effort — not a studio.
Shot sequence:
— Close-up of the shoe alone on wet gravel at early morning. A single raindrop falls on the upper. Locked-off. (4 seconds)
— Wide tracking shot at ankle-height: the shoe in motion, pushing off mud on a forest trail. We see only from the ankle down — the ground rushing past. (4 seconds)
— Wide shot: a runner reaches the crest of a hill at dawn. The city is visible in the valley below, the sky just beginning to lighten. She stops. Looks out. (7 seconds)
Audio: No dialogue. No music. Rain on the first shot. Mud and exertion sounds on the trail shot. Then wind, her breathing, and silence at the top.
16:9, 15 seconds total.18. Cheat Sheet
A quick-reference summary of everything in this guide. Print it, save it, or paste it at the top of your generation workflow.
The Formula
Subject → Action → Camera → Style → Audio → SpecsTechnical Limits (Always Keep in Mind)
| Parameter | Limit |
|---|---|
| Max duration | 15 seconds |
| Image references | 9 max — tag as @Image1, @Image2, etc. |
| Video references | 3 max (15s total) — tag as @Video1, etc. |
| Audio references | 3 max (15s total) — tag as @Audio1, etc. |
| Resolution | 2K (2048px) |
| Aspect ratios | 16:9, 4:3, 1:1, 3:4, 9:16 |
| Negative prompts | NOT supported — use positive descriptions |
| Actions per shot | One dominant action — use multi-shot for more |
Reference Roles — Quick Assignment Guide
| What you have | Tag it as | Say in prompt |
|---|---|---|
| Character photo | @Image1 | "Use the face from @Image1 exactly" |
| Costume/outfit photo | @Image2 | "@Image2 is the outfit — replicate precisely" |
| Location photo | @Image3 | "@Image3 is the setting" |
| Style/mood reference | @Image4 | "@Image4 is the visual color grade reference" |
| Storyboard sketch | @Image1 | "Refer to the storyboard in @Image1" |
| Camera move reference | @Video1 | "Use the camera movement from @Video1" |
| Performance reference | @Video1 | "Transfer the motion from @Video1 to the character below" |
| Music track | @Audio1 | "@Audio1 is the background music — begin at opening frame" |
| Voiceover | @Audio1 | "The narration in @Audio1 says: [script]. Generate visuals that illustrate this." |
Camera Movements — Quick Reference
| Want to... | Use this |
|---|---|
| Move toward the subject | Dolly-in / Push-in |
| Move away from subject | Pull-back |
| Follow a moving subject | Tracking shot |
| Show what's next door | Pan left / Pan right |
| Rise or fall dramatically | Crane up / Crane down |
| Circle around something | Orbit |
| Feel raw and immediate | Handheld |
| Feel smooth and polished | Gimbal / Steadicam |
| Create unease | Dutch angle |
| See through a character's eyes | POV shot |
Multi-Shot Trigger Words
Use these phrases to signal a cut within a single generation:
"Scene cuts to...""Lens switch to...""Cut to:""Switches to ultra-slow motion""Normal speed resumes""Hard cut to""Cross-dissolve into"
Audio Direction — Quick Templates
Audio: [List specific sounds]. [Music or no music]. [Emotional character of sound].| Situation | Template |
|---|---|
| Silence / near-silence | "Audio: Near silence — only [one specific sound]. No music." |
| ASMR | "Audio: ASMR style — [specific tactile sounds]. No voice. Room ambient only." |
| Dialogue | "She says with [tone], '[exact words].'" |
| Action SFX | "Audio: [impact sound], [material sound], [environment sound]. No music." |
| Music-driven | "Audio: [describe genre/tempo/mood] music begins at [point]. [Does it fade? End hard?]" |
The Non-Negative Prompt Cheat Sheet
| Don't write | Write instead |
|---|---|
| "No people" | "Empty, deserted, unpopulated" |
| "No music" | "Audio: ambient only. No score." |
| "No harsh shadows" | "Soft, diffused light from overcast sky" |
| "No clutter" | "Minimal, clean surfaces, organized space" |
| "No fast movement" | "Slow, deliberate, unhurried movement throughout" |
| "No dialogue" | "Audio: no voiceover, no speech. [What sounds instead?]" |
Choose Your Model
| If you need... | Use |
|---|---|
| More than 8 seconds | Seedance 2.0 |
| 4K resolution | Veo 3.1 |
| More than 3 image references | Seedance 2.0 |
| Video or audio references | Seedance 2.0 |
| Multi-shot in one generation | Seedance 2.0 |
| Simple, single-shot, text-only | Either |
The Three Questions Before You Generate
- Have I specified the aspect ratio? (16:9, 9:16, 1:1, 4:3, 3:4)
- Have I directed the audio explicitly? (Even "near silence, no music" is better than nothing)
- Have I tagged every reference file with its role? ("@Image1 is the character" — not just "@Image1")
If yes to all three: generate with confidence.
Guide expanded from original technical spec. Sources: ByteDance Seed Official Blog · ZenCreator Guide · Forbes · fal.ai · Atlabs AI · Freepik Blog · Imagine.Art · TechCrunch
