All posts
guides

Lip Sync Video Creation Guide 2026

Lip sync takes a video and the audio that face should be speaking. The output is as long as the audio, which changes the order you should do everything in.

By Flixly TeamJune 15, 2026
Lip Sync Video Creation Guide 2026

TL;DR

Lip sync takes a video containing a face and the audio it should speak. It is a separate job from video generation: the models are LatentSync, Sync LipSync and Volcengine Lip Sync, not Seedance, Kling or Veo. Output length follows the audio, and billing is on the measured audio duration, so trim before submitting. Finish the video, finalise the audio, then sync, then caption, because anything that changes the audio invalidates the sync.

Lip sync takes two files: a video with a face in it, and the audio that face should be speaking.

It is not a setting on a video model. Seedance, Kling and Veo generate video; they do not sync mouths to a track. The models that do this are LatentSync, Sync LipSync and Volcengine Lip Sync, and they exist as a separate job for a reason.

Understanding that split fixes most of what goes wrong here, because the order you do things in turns out to matter more than any parameter.

The rule that governs everything

The output is as long as the audio.

Not as long as your video, not a duration you pick. The audio decides, and you are billed on its measured length rather than an estimate or a ceiling.

Two consequences worth internalising:

Trim the audio before you submit, not after. Silence at the end is output you paid for and will cut off anyway.

If your video is shorter than your audio, you have a problem to solve before running the job, not after.

Order of operations

This is where most of the wasted work happens. Lip sync matches a mouth to an existing track, so anything that changes the audio invalidates the sync.

The correct order:

  1. Finish the video. Whatever generation or editing it needs, done.
  2. Finalise the audio. Script locked, voice chosen, trimmed to length.
  3. Lip sync. Now the mouth matches a track that will not change.
  4. Captions last, once picture and audio are both final.

Doing lip sync before the audio is settled means doing it twice. Doing it before the video is settled means doing it twice. It belongs near the end.

Getting the audio right

Sync quality follows audio quality closely, and this is the input you have most control over.

Clean speech, one speaker. Overlapping voices give the model no single mouth shape to target.

Trimmed. Leading and trailing silence become dead frames in the output.

Generated audio is usually easier than recorded. Text to Speech produces clean, evenly-paced speech with no room tone or plosives. For a series where the voice must stay consistent across episodes, Voice Cloning is the tool.

If you are dubbing into another language, Volcengine Lip Sync supports multilingual dubbing specifically. Video Translator handles the translation side when the source already has speech.

Getting the video right

The face has to be workable.

Visible and reasonably large in frame. A face occupying a small part of a wide shot gives the model little to work with.

Facing roughly toward camera. Profiles have half a mouth visible, and results degrade accordingly.

One face. Multiple faces are ambiguous about which should be speaking.

Not already speaking different words, ideally. Replacing existing mouth movement is harder than animating a neutral or closed mouth.

If you are generating the video anyway, generate it with this in mind: a medium shot, face to camera, mouth closed or neutral, is far easier to sync than a wide dynamic shot.

Picking a model

Three, and the differences are about tier rather than a benchmark:

LatentSync — standard tier, latent-space sync.

Sync LipSync — premium tier, the higher-quality option.

Volcengine Lip Sync — standard tier, video-to-video, with multilingual dubbing support. This is the one to reach for when the audio is in a different language from the original performance.

You will find articles quoting phoneme counts, facial landmark counts and accuracy percentages per model. Those numbers are invented. No such figures are published, and the useful comparison is tier and dubbing support.

When it still looks wrong

Mouth movement lags or leads. Almost always an audio problem: leading silence, or a track that does not start where you think it does. Trim and retry.

Mouth moves but the face is stiff. Expected. Lip sync animates the mouth, not a performance. If you need the whole face to act, that is a generation problem — start from a clip where the person is already animated.

Result is uncanny on a still image. A photograph has no natural head movement, so a moving mouth on a motionless face reads oddly. Animate the still first with image to video, then sync the result.

Sync degrades over a long clip. Cut into shorter segments, sync each, and rejoin. Short jobs also fail more cheaply.

Where it sits in a real workflow

For a talking-head short: generate or shoot the video, write and generate the voice, lip sync, then cut vertical versions in the Shorts Generator and caption with Auto Captions.

For a dubbed piece: get the translated audio, lip sync with Volcengine, caption in the target language.

For a character series: lock the character, keep the same cloned voice across episodes, and sync each episode separately.

Costs are based on the measured audio length and quoted before the job runs. Pack prices are on the pricing page, and the catalog is at Models.

The single highest-value habit here is simple: settle the audio completely before you sync anything.

Frequently Asked Questions

Which models actually do lip sync?

LatentSync at standard tier, Sync LipSync at premium tier, and Volcengine Lip Sync at standard tier with multilingual dubbing support. Seedance, Kling and Veo are video generation models and do not sync mouths to an audio track. Articles quoting phoneme counts, facial landmark counts or accuracy percentages per model invented those figures.

How long will the output be?

As long as the audio. Not as long as your video and not a duration you choose. Billing is on the measured audio length, so trim leading and trailing silence before submitting rather than after, since silence becomes output you paid for.

Should I lip sync before or after editing?

After. Lip sync matches a mouth to an existing audio track, so anything that changes the audio invalidates the sync. Finish the video, finalise the audio, then sync, then caption. Doing it earlier means doing it twice.

Why does the mouth lag behind the audio?

Almost always an audio problem rather than a model one, usually leading silence or a track that does not start where you think it does. Trim the audio precisely and run it again.

Why does the result look uncanny on a photo?

A still image has no natural head movement, so a moving mouth on a motionless face reads oddly. Animate the still first with image to video, then lip sync the resulting clip. Lip sync animates the mouth, not a whole performance.

What makes a good source video for lip sync?

A face that is visible and reasonably large in frame, facing roughly toward camera, and only one face in shot. A neutral or closed mouth is easier to work with than someone already speaking different words. If you are generating the video, a medium shot facing camera syncs far better than a wide dynamic one.

Can I dub a video into another language?

Yes. Volcengine Lip Sync supports multilingual dubbing, so you can sync a face to audio in a different language from the original performance. Video Translator handles the translation itself when the source already has speech.

Tools mentioned in this post

guideslip-syncvideodubbing

Ready to create with guides?

Jump straight into Flixly's AI studio and try guides with 50+ models — free to start.