Home/Glossary/What Is Lip Sync?
GLOSSARY

What Is Lip Sync?

Lip sync is the alignment between visible mouth movement and spoken or sung audio. In AI video, software may generate new mouth motion from audio, retime an existing performance for translated speech, or drive an avatar live from sound or camera tracking.

Try it freeSee how it works ↓
LIVE · REAL-TIME
✦ AI live
// glossary · live in browser

Good lip sync is easy to ignore because the sound and picture feel like one event. Bad lip sync is immediate: a word arrives before the mouth moves, a closed mouth continues speaking, or every sound produces the same open-and-close animation. The term covers several techniques that solve different production problems, so “AI lip sync” alone does not tell you whether a tool is live, batch-generated, translation-focused, or avatar-based.

What lip sync means

Lip sync connects phonetic speech with visible articulation. The system does not merely detect volume. It tries to choose mouth shapes that plausibly match sounds — for example, closed lips for sounds such as m, b, and p, and different open shapes for vowels.

In animation and avatar software, those reusable mouth shapes are often called visemes: visual counterparts to units of speech. A simple system maps audio to a small set of visemes. A generative system may synthesize the entire lower face and surrounding expression rather than switching between prepared shapes.

Four common kinds of lip sync

Manual animation lip sync

An animator listens to a recorded performance and assigns mouth poses frame by frame or through keyframes. It offers the most deliberate control and can be stylized to match a character, but it takes time.

Audio-driven avatar lip sync

Software analyzes live or recorded audio and drives a rigged 2D/3D avatar. Basic versions respond mainly to volume; better versions infer phonemes or visemes. Tools such as conventional VTuber applications may combine audio lip sync with webcam tracking.

Generated talking-photo or avatar video

The user provides a photo or avatar plus text or audio, and an AI model generates a finished speaking clip. This is the workflow behind many talking-avatar products. It is useful when no performer needs to be on camera, but it usually involves processing rather than an open-ended live session.

Video translation and redubbing

The original speech is translated or replaced, then the visible mouth region is adjusted to match the new language. This keeps an existing recorded performance while changing what is spoken. It is a different goal from making a still image talk.

Live lip sync vs generated lip sync

DimensionLive lip syncGenerated lip sync
InputMicrophone and/or camera in the momentRecorded audio, script, or existing video
Output timingContinuousProduced after processing
Best forStreaming, calls, live avatarsAds, explainers, translation, social clips
ControlPerformer controls timing liveScript, voice track, and editor control timing
Failure modeTracking or latency changes during the sessionArtifacts appear in the rendered clip and require regeneration

Neither category is automatically better. A streamer needs low delay and robust tracking. A marketing team needs repeatable narration, language versions, and editable scripts.

How LiveGen relates to lip sync

LiveGen is not an audio-only lip-sync generator. Its character view is driven by a live camera performance through Live Morph. The person's visible speaking, expressions, gaze, and movement carry into the generated character frame by frame. That means the human performer remains responsible for the words and timing.

This is useful for live reactions because the mouth is part of a complete performance rather than a separate post-production pass. It is not the right workflow for typing a script, selecting a synthetic voice, and producing an unattended talking presenter. For that batch category, see the independent D-ID alternative and DreamFace alternative comparisons.

What makes lip sync look convincing?

Common lip-sync problems

FAQ

Frequently asked questions

What is AI lip sync?
+

It is software-generated alignment between speech audio and visible mouth motion, applied to an avatar, photo, or existing video.

Can AI lip sync work in real time?
+

Yes. Some avatar systems infer mouth shapes from live audio or camera tracking. Other tools process uploaded audio and return a finished clip later.

What is the difference between lip sync and face tracking?
+

Lip sync focuses on speech and mouth motion. Face tracking measures broader movement such as head pose, gaze, blinks, brows, and expressions; it may include mouth tracking.

Does LiveGen lip-sync uploaded audio to a photo?
+

No. LiveGen's character is driven by the visible live camera performance. Use an audio-driven talking-avatar tool for uploaded narration.

Why is my avatar's mouth delayed in OBS?
+

The camera, tracking or generation step, capture path, audio interface, and encoder may each add buffering. Record a clap test and apply a measured sync offset to the faster source.

What is a viseme?
+

A viseme is a visible mouth shape associated with one or more speech sounds. Rigged avatars often map detected phonemes to a prepared viseme set.

Is lip sync the same as dubbing?
+

No. Dubbing replaces the audio performance. Lip-sync technology may then adjust the picture so that it matches the new audio.

Play the world. In real time.

Open your camera and become any character — free to start.

Start generating free