How to Sync Seedance 2.0 with External Voice Over Audio (Step by Step)
Two proven ways to sync Seedance 2.0 with external voice over audio: feed your VO as an @Audio reference so the model matches it, or generate first and sync in an editor. Step-by-step, with credit costs and troubleshooting.

How to Sync Seedance 2.0 with External Voice Over Audio (Step by Step)
Last week I recorded the perfect voice over. Clean room tone, right pacing, the whole script delivered in one take. Then I opened the generator, typed my prompt, and the video came back with the character talking on a completely different beat — with music drowning out my narration.
It's the most common question we get from creators using Seedance 2.0: how do you sync the model's output with voice over audio you already have?
The short answer: Seedance 2.0 gives you two legitimate paths, and both work.
Path A — Feed the voice over into the model. Seedance 2.0 generates native audio (dialogue, ambient sound, music) in the same pass as the video, and its reference system accepts audio files tagged with @Audio. Upload your VO as an audio reference, and the model generates video matched to it.
Path B — Generate first, sync in an editor. Generate the video (with or without its native audio), then align your external VO on the timeline in any editor, using waveform matching and speed adjustments.
This guide walks through both paths step by step, when to use each, and how to fix the four most common sync failures. It's based on hands-on testing of the generator at seedance2pro.io as of September 2026 — plan features, credit costs, and reference-system behavior included. By the end, you'll be able to take any recorded voice over and get a Seedance 2.0 video that lands on the beat — without guessing.
Before We Start: What You Need to Know
The generator supports text-to-video, image-to-video, multi-shot sequences, and native audio with lip-sync. Its reference system accepts up to 12 reference files — images, videos, and audio — each tagged with @Image, @Video, or @Audio.
One note on labeling: some details below (exact credit costs, plan features) are published on the pricing page; anything that isn't officially published is marked as such. If a feature changes, the pricing page is the source of truth.
How Audio Conditioning Actually Works in Seedance 2.0
Before the steps, it helps to understand what the model is doing — because most sync failures come from a wrong mental model.
Seedance 2.0 is not a video generator with audio bolted on. It produces dialogue, ambient sound, and music in the same generation pass as the visuals. That means the audio isn't a separate track you can swap — it's baked into the video's timeline, which is why "just replace the audio" doesn't always work cleanly.
When you attach an @Audio reference, you're not asking the model to "play" that file. You're giving it a conditioning signal: the model analyzes the audio's content, pacing, and emotional shape, then generates visuals that fit it. Think of it like giving a director a reference track — the director doesn't play the track in the scene, but every performance is timed to it.
This is the key insight that makes both paths work:
- Path A works when you want the model to respect your VO — its timing, its words, its emotional arc.
- Path B works when you want full control over the final mix — because the model's native audio is a starting point, not a contract.
One more thing to know: the model's native audio includes lip-sync for characters who speak. If your external VO features a character talking on screen, Path A is the only path that gives you a chance at matching mouth movement — Path B will require you to either accept mismatched lips or hide them.
With that mental model in place, choosing a path becomes one question — answered in the table below.
Which Path Should You Use? A Decision Framework
| Your project | Path | Why |
|---|---|---|
| Character speaks on camera, VO is the script | Path A | Only path where seedance 2.0 lip sync can match your audio |
| Narration over b-roll, no visible speaker | Path B | Full control over the final mix, no lip-sync constraint |
| Music video / ambient scene with VO on top | Path B | You'll re-mix anyway; generate visuals freely |
| Short ad with a spokesperson | Path A | Timing and mouth movement both matter |
| Multi-shot sequence with one VO across shots | Path A (per shot) | Each shot can reference the same audio file |
| You need pixel-perfect audio timing | Path B | Model matching is close, not frame-exact |
Rule of Thumb: if a mouth is visible and speaking, use Path A. If no one's lips are on screen, use Path B. That single question resolves 90% of the choice.
Path A: Feed Your Voice Over as an @Audio Reference
If the framework pointed you to Path A — a visible speaker whose timing matters — here's how to run it. This is the direct route: your external VO becomes the model's audio reference, and the video is generated to match it.
Step 1: Prepare the Audio File
The model reads audio references as conditioning signals, so the file quality matters more than the file size.
- Format: WAV or MP3. Both work; WAV preserves more detail if your VO is a high-quality recording.
- Sample rate: 44.1 kHz or higher. Lower sample rates compress the timing information the model uses.
- Clean room tone: Remove background hum, fan noise, and reverb before uploading. The model treats everything in the file as signal — a noisy reference produces noisy results.
- Trim the silence: Cut long pauses at the start and end. Keep natural pauses between sentences, but remove dead air at the file edges.
- Length: Keep the reference close to the length of the shot you're generating. A 15-second VO reference for a 5-second shot will confuse the timing.
Step 2: Upload and Tag the File
Open the generator, upload your audio file, and tag it @Audio — that's the tag that marks it as the seedance 2.0 audio reference. The system uses @Image, @Video, and @Audio tags to tell the model what each file is. You can mix reference types in one generation (up to 12 files total), so a character reference image plus your VO audio works in the same prompt.
Step 3: Write the Prompt with Timing Cues
The prompt is where you tell the model how to use the audio. Be explicit:
- Describe who is speaking and what they're doing while speaking.
- Mention the emotional tone of the VO ("calm narration", "urgent warning").
- If the VO has a distinct rhythm — a beat, a pause, a punchline — describe it: "the speaker pauses for two seconds before the final line."
The model won't transcribe your VO word-for-word into the video (that's not officially published behavior, so treat it as a limitation, not a feature). What it will do is match pacing and mood.
Step 4: Generate and Check the Sync
Pick your settings and generate. The first check is simple: does the character's speech land on the same beats as your VO? Play them side by side — the model's native audio should feel like it's answering your reference, not ignoring it.
If the sync is close but not perfect, regenerate with a tighter prompt. If it's completely off, the problem is usually the reference file (noise, wrong length, wrong pacing) — fix the file before burning more credits.
Step 5: Export and Mix
All plans export MP4, so you get a single file with the model's native audio baked in. If you want your original VO on top of the model's visuals, you'll still do a light mix in an editor — but because the model matched your reference, the alignment work is minimal.
Path B: Generate First, Sync in an Editor
If no lips are visible in your shot, this is the control route: generate the video, then align your external VO yourself. It's the right choice when you need the final mix to be exactly your recording — no model interpretation.
Step 1: Generate the Video
Generate your video with native audio enabled (or with dialogue in the prompt if you want lip-sync to work with). You're not trying to match anything yet — you're creating footage with a usable timeline.
Step 2: Align on the Timeline
Drop the video and your VO into any editor (CapCut, Premiere, DaVinci Resolve — all work). The fastest alignment method is waveform matching:
- Mute the video's audio track.
- Zoom into the waveform of your VO.
- Find the first strong syllable — the first word's onset.
- Slide the VO track until that onset lands on the moment the character's mouth first moves.
This gets you within a frame or two in under a minute. For longer videos, repeat the process at the midpoint and the end — drift accumulates, and one anchor point isn't enough.
Step 3: Fix Drift with Speed Adjustments
If the VO drifts from the video's pacing — the character finishes speaking before your VO does — you have two options:
- Speed-ramp the video (0.5–2% speed changes are invisible to the eye) to stretch or compress the footage to match your VO.
- Cut the VO to match the video's natural beats, if the video's pacing is the part you want to keep.
Rule of Thumb: if the drift is under 5% of the clip length, speed-ramp the video. If it's more than that, re-cut the VO — stretching footage more than ~5% starts to look unnatural.
Step 4: Handle Lip-Sync Honestly
This is the hard truth of Path B: if your character is visibly speaking, the model's mouth movements were generated for the model's own audio, not your VO. You have three options:
- Accept it — for narration over b-roll, wide shots, or characters facing away, mismatched lips are invisible.
- Hide it — cut to other shots during speech, or keep the speaker off-camera.
- Regenerate with Path A — if lip-sync matters, Path A is the only path that gives the model a chance to match mouth movement to your audio.
Rule of Thumb: if the speaker's mouth is visible for more than 3 seconds, use Path A. If not, Path B is faster and gives you full mix control.
Verify Before You Commit: Use the 5 Free Generations
Before you spend a single credit on either path, there's a free way to find out which one works for your project: the generator offers 5 free generations with no card required. Use them deliberately:
- Generation 1–2: Test Path A with your actual VO file. Check sync quality on a 5-second clip.
- Generation 3: Test Path B — generate the same scene without the reference, and time how long waveform alignment takes in your editor.
- Generation 4–5: Test the edge cases — a noisy VO file, and a VO with unusual pacing — so you know your file's limits before you pay.
This is the lowest-friction way to learn your own failure modes. If Path A sync is close enough in your test, you've saved yourself an entire editing pass. If it isn't, you've learned that for a few minutes and zero dollars.
Troubleshooting: The 4 Most Common Sync Failures
Even with the right path, sync can still fail — here are the four failures we see most often, and how to fix each one.
1. The Voice Over Drifts Mid-Clip
- Symptom: The first 2 seconds match, then the character falls behind your VO.
- Root cause: The reference file's pacing doesn't match the shot length — usually a VO that's longer than the generated clip, or a prompt that didn't describe the timing.
- Resolution: Re-trim the VO to the shot length, tighten the prompt with explicit timing cues ("speaks steadily for 5 seconds, pauses, then delivers the final line"), and regenerate. If you're on Path B, apply the speed-ramp fix from Step 3.
2. The Model Ignores the Audio Reference Entirely
- Symptom: The generated video's audio has nothing to do with your VO — different mood, different pacing, no relation.
- Root cause: The file wasn't tagged correctly, or the reference was too noisy for the model to extract a signal.
- Resolution: Confirm the
@Audiotag is applied (not@Imageor@Video), re-export the file at 44.1 kHz+ with clean room tone, and retry. If it still fails, test with a different, cleaner recording to isolate whether the problem is the file or the model.
3. Lip-Sync Is Off
- Symptom: Mouth movements don't match the words, even on Path A.
- Root cause: Seedance 2.0's lip-sync matches the model's generated audio, which follows your reference's pacing and mood — not a phoneme-perfect transcription of your VO. (Phoneme-level matching to external audio is not officially published behavior.)
- Resolution: Accept close-but-not-perfect sync, or switch to Path B and hide the mouth (cutaways, off-camera speaker). If you need exact lip-sync, plan shots that don't show the mouth during speech.
4. Audio Quality Loss After Export
- Symptom: The final MP4 sounds thinner than your original VO.
- Root cause: You're hearing the model's native audio, not your VO — or your VO was compressed twice (upload → generation → export).
- Resolution: On Path A, keep your original VO file and mix it over the export in your editor, using the model's audio as a timing guide. On Path B, mute the model's audio entirely and use your original file. Keep your master VO at 44.1 kHz+ WAV so you're never working from a compressed copy.
FAQ: Real Questions About Seedance 2.0 and Voice Over
Does Seedance 2.0 accept audio files as references?
Yes. The reference system accepts up to 12 files — images, videos, and audio — and audio files are tagged with @Audio. This is the seedance 2.0 audio reference feature.
Can I upload my own voice over to Seedance 2.0?
Yes. Upload your VO file, tag it @Audio, and the model will use it as a conditioning signal for the generation. It won't play your file verbatim — it generates video and native audio matched to it.
How do I make the model follow my VO's timing?
Three things: prepare a clean file (WAV/MP3, 44.1 kHz+, trimmed), tag it @Audio, and describe the timing in your prompt — pacing, pauses, emotional tone. The model matches rhythm and mood, not word-for-word transcription.
What's the best audio format for a Seedance 2.0 voice over reference? WAV at 44.1 kHz or higher with clean room tone. MP3 works too, but WAV preserves more timing detail. Avoid heavily compressed files and noisy recordings.
Why is my lip-sync off? Seedance 2.0's native audio includes lip-sync, but it syncs to the model's generated audio, which follows your reference's pacing — not a phoneme-perfect match to your external VO. If lips are visible and speaking, Path A is your best shot; otherwise, plan shots that don't show the mouth.
Can I remove the generated audio and use only my VO? Yes. Export the MP4, mute the model's audio track in your editor, and place your VO on the timeline. All plans export MP4, so this works on every tier.
Do I need the Pro plan for audio references? No. Audio references and native audio work on all plans. The differences between plans are credits, speed, and extras: Basic is 19.99/mo for 530 credits (commercial license, 16-bit HDR & EXR exports), and Max is $29.9/mo for 790 credits (fastest generation, permanent storage). See the pricing page for details.
Use It Responsibly: Rights and Guardrails
Only Upload Voice Over You Have Rights To
If you recorded it, you're fine. If you licensed it, keep the license. If you cloned or sampled someone's voice, stop — uploading a voice you don't own is a rights problem, not a technical one.
Don't Use Other People's Voices Without Consent
This includes celebrity voices, friends' voices, and AI-cloned voices. Consent is the line.
Check the Commercial License Before Monetizing
The Pro plan includes a commercial license; if you're producing client work or monetized content, confirm your plan covers it before publishing.
Label AI-Generated Content Where Required
Platform rules on AI video disclosure vary — check the platform you're publishing on before posting.
Core Summary
Syncing Seedance 2.0 with external voice over audio comes down to one decision: does the model need to respect your audio, or do you need to control the final mix?
- Visible character speaking → Path A. Feed your VO as an
@Audioreference — clean file, correct tag, timing cues in the prompt. - No lips on screen → Path B. Generate freely, then align in your editor with waveform matching and the 5% speed-ramp rule.
- Not sure yet → spend your 5 free generations testing both paths before spending a single credit.
And the rule to remember: visible mouth → Path A, no visible mouth → Path B.
Start With a 5-Second Test Clip
Open the Seedance 2.0 generator, upload your cleanest VO file, tag it @Audio, and generate one 5-second test at 720p. That's 90 credits on the Basic plan — and it's free if you haven't used your 5 free generations yet. If you want to see how the credit system scales before you commit, the API pricing breakdown covers it in detail.
More Posts

How to Create Handheld TikTok-Like Video in Seedance 2.0 (Step by Step)
Learn how to create handheld TikTok-like video in Seedance 2.0: camera-motion prompts, the 9:16 settings recipe, @Video references, 5 ready-to-use prompts, and troubleshooting.

How to Use Seedance 2.0 in Browser: Step-by-Step Guide (No Install, No GPU)
Learn how to use Seedance 2.0 in browser on seedance2pro.io — no install, no GPU, no app. Step-by-step walkthrough: account setup, model choice, prompts, references, queue, and MP4 download.

How Does Seedance 2.0 Find What a Face Is? (Face Detection & Identity Explained)
How Seedance 2.0 detects faces, compresses them into identity embeddings, and keeps the same face consistent across every frame — explained in plain English with a simple pipeline.
Newsletter
Join the community
Subscribe to our newsletter for the latest news and updates