what sign up
nixm nixm

← learn

AI lip sync video — one clip where it works, one where it doesn't

September 27, 2026

AI video can now generate a spoken line and the mouth movement together, in one pass. But it is not reliable yet. Below are two real clips from nixm, made the same week with the same character. In the first, a short line lands in sync. In the second, a longer line after a cut does not: the lips stop moving while the words carry on. Short lines, one line per shot, the face toward the camera and no cut right before the line give the best odds. On nixm, lip sync is still experimental.

What "lip sync" means in AI video

There are two different things people call AI lip sync.

Generated with the clip. The first is when the video model produces the voice and the mouth movement together, from a line written in the prompt. Nothing is recorded; the model invents both. This is how Veo-class models work, and how nixm renders dialogue. How that audio is made is in text to video with sound.

Dubbing. The second is a separate category of tools that take existing footage and existing audio, such as your own recorded voice, and re-animate the mouth to match. If you need a specific voice or a long script, that is the right kind of tool, and it is a different job from generating a scene. The distinction is covered in AI video with voiceover.

This page is about the first kind, shown honestly.

Clip one: it works

A reality-TV-style street interview. The character, Lilly, was cast once in nixm and dressed from a reference. An off-camera voice asks "are you a nixm character?" and she turns to camera and answers, "yeah, of course i am." The line is five words, delivered in the first shot, with her face turned toward the lens.

Episode one, as rendered: a 15-second CAST scene with sound generated in the same pass. The short line lands in sync. Nothing was edited afterwards.

Clip two: it doesn't

Same character, a new outfit and a new place. Lilly pulls an espresso shot behind a café counter. The scene then cuts to a confessional setup, and she says to the lens: "last week, a night street. this week, barista. never bored." The line is nine words and takes about four seconds to say, and it starts right after the cut.

Episode two, as rendered. In the confessional, the mouth moves for roughly the first two seconds, then settles into a closed hold while the line carries on. Nothing was edited afterwards.

On listening, the mismatch is obvious: the words keep going after the mouth has stopped.

What was different

Two clips are not a study, so these are observations rather than rules. Still, the differences line up with what creators commonly report.

  • Line length. Five words fit comfortably in the time the model animates a mouth; nine words over four seconds did not.
  • The cut. In clip two the line starts in a brand-new shot, right after a cut. In clip one it is spoken inside a shot that is already established.
  • What the shot has to do. Clip two spends its first half on a different action at the counter, then asks for a speech. Clip one is built around the moment she answers.

Both clips have the face toward the camera, so framing alone does not explain it.

How to get better AI lip sync

  • Keep lines short. A handful of words, not a paragraph.
  • One line per shot, and not right after a cut. Let the shot settle before the character speaks.
  • Face toward the camera, head fairly still. A turning head or a profile makes the mouth harder to read, and harder to animate.
  • Make the line the point of the shot. Do not stack a separate action and a speech into the same few seconds.
  • For long speech, use a different tool. Record the voice and dub it, or lay a voice-over across quieter shots.

The other half: the same character twice

These two clips were also a test of keeping one character across episodes. The styling held: the same shoulder-length blonde hair and heavy smoky eye makeup, and in episode two a new outfit made from a new still.

Lilly's reference still: a young woman with shoulder-length blonde hair and heavy smoky eye makeup, looking toward the camera.
The "Lilly look" still that episode one was cast from.
The same character in a green barista apron over a black tee at a café counter.
The "Lilly barista" still made for episode two's costume change.

The face was less steady. Episode two matches the stills closely, while episode one's Lilly reads slightly fuller-faced. That may be because the prompt asked her to hold a wide grin throughout, which pulls the face away from its neutral reference. Why that happens, and how to limit it, is in why AI video drops your references.

What it cost

Each clip is a 15-second CAST scene with sound, at 3 video credits. The costume change in episode two needed one new still first, for 2 chat credits. Renders took about 9 and 12 minutes. At the 30-credit pack price that works out to roughly CA$10 per scene, retakes not included. The full workflow for longer pieces is in making a short film with AI.

Try a line yourself

A new nixm account includes one free HD clip with sound. Open the cinematic studio, keep the line short, and give it a shot of its own.

frequently asked

Can AI video generate lip sync?

Yes. Current models can generate a spoken line and the mouth movement together in one pass, but results vary: short lines in a settled, front-facing shot work best.

Why is my AI video lip sync off?

Common causes are lines that are too long for the shot, a line that starts right after a cut, a turning or side-on face, and asking one short clip to do an action and a speech at once.

How long can a spoken line be in an AI video clip?

Keep it to a handful of words. In our test, a five-word line synced well; a nine-word line lasting about four seconds did not.

Is nixm lip sync reliable?

Not yet. nixm labels lip sync as experimental. Short lines in a shot of their own give the best odds, as the two clips on this page show.

What is the difference between AI lip sync and dubbing?

Generated lip sync makes the voice and the mouth together from a written line. Dubbing tools re-animate the mouth in existing footage to match recorded audio, which suits long scripts and specific voices.

Does the same character stay consistent across AI video clips?

Styling usually holds when the character is cast from a reference. Faces can drift, especially when the prompt asks for a strong expression throughout.

Can I try it free?

Yes. A new nixm account includes one free HD video with sound, no card required.

try it now — nixm makes cinematic ai video, with sound. no subscription. first video free.

open the studio →

Enter your email and receive a link to sign in. No passwords, just a secure signal to connect.

why connect

  • 20 free credits to start
  • your signals, remembered
  • buy more credits anytime
  • image & video generation features
or
+