NEXINFINITY META · Blog
AI Video

Why AI Talking Videos Look Fake — and the Lip-Sync Fixes That Work in Tamil

S. Veera Kumar 8 min read

Most AI talking videos look fake for the same four reasons: the face was animated without the voice and synced afterwards, so the mouth moves but says nothing; the voice on the soundtrack is not the voice the face was performing to; two faces share a shot, so both mouths move; and the eyes are half-closed or dead while the lips are busy. None of these is fixed by a better prompt. They are fixed by the order you do things in. We learned each of them the expensive way on an eight-minute Tamil tribute film with two lip-synced speakers, and this is what we would tell anyone about to make a spoken AI video in Tamil or any Indian language.

Key takeaways

  • Animate the face to the line's own recording first, then lip-sync; a face generated silently and synced afterwards has a mouth that moves but says nothing.
  • Never keep the model's re-synthesised audio — lay the original recording back and sync to it.
  • One speaker per shot: two faces in a speaking shot means two moving mouths.
  • Eyes open, face front-on while speaking; sync fails in profile and half-closed eyes read as absent.
  • In Tamil, record real people; text-to-speech still gets names and spoken forms wrong.

1. The generic talking mouth: animate to the voice, not before it

The most common way to make an AI presenter is to generate a clip of a face 'talking', with no audio, and then run a lip-sync tool over it with the real voice. It is fast and it looks almost right. The lips open and close in time. But the face around them — the jaw, the cheeks, the eyebrows, the small nods on the stressed words — was never performing that sentence. People feel the mismatch even when they cannot name it.

We measured this directly. A speaking shot that a family reviewer called unreal had been generated silently and then lip-synced. We ran a second lip-sync pass on it with the same recording: the mouth region changed by an average of under 4 levels on a 0–255 scale — nothing anyone could see. The sync was never the problem. The performance was.

The fix is to reverse the order. Give the video model the line's own recording as a reference, so the whole face is performed to the rhythm of that sentence, and only then run the lip-sync pass to fit the mouth to the exact recording. The re-made shot was approved at its first review.

If a synced shot still looks wrong

Do not buy another sync pass on the same clip. Re-make the performance with the voice driving it, then sync. A second sync over a silent-generated face changes almost nothing.

2. The voice on the soundtrack must be the recording, not the model's copy of it

When you give a video model an audio reference, many models do not pass that audio through. They re-synthesise the speech in the reference voice — and in doing so they drift. On our film, a line that ended at 5.6 seconds in the recording ended at 5.0 seconds in the generated shot, and the pronunciation of Tamil words shifted enough that a listener who knew the original heard it at once.

So the model's audio is never kept. The clean recording is laid back over the picture, and the lip-sync pass is run against that recording, not the model's. That way the timing, the pronunciation and the voice are all the original, and the mouth is fitted to them.

3. One speaker per shot

Put two people in a frame and give the video model a line of speech as its reference, and both mouths move. It does not matter which face you name in the prompt; the speech reference animates every face it can find.

You can generate the two-shot silently and let the lip-sync pass animate only the speaker — that works, but it brings back problem one: a face that was never performing the line. The reliable fix is editorial. Every spoken line gets its own single shot, and a conversation becomes a cut between singles. A two-shot can still appear, as long as nobody in it speaks. Films were cut this way long before AI, and it is still the cleanest way to have two AI people talk to each other.

4. The eyes, and the face turning away

A face that talks with its eyes half-closed reads as absent, even when every syllable is in sync. It also fools any automatic check you might run: blinking adds movement, so a motion score says the face is lively exactly when it is not. The only reliable check is to look at the frames. In the prompt, ask for eyes open and one or two natural blinks, and say it in plain positive words.

Lip-sync also fails when the face turns to profile — the mouth is no longer where the tool expects it, and on our film one such shot simply could not be synced. Keep speaking shots facing the camera. If the character must turn, let them turn after the last word, not during it.

5. Tamil needs real voices

Text-to-speech in Tamil has improved, but it still stumbles over exactly the words that matter in a family or a local brand: names, kinship terms, spoken rather than written forms, the difference between a mother saying 'ம்மா' and 'டா'. We tried a cloned voice reading new lines and a native listener caught the wrong Tamil within a sentence.

What worked was recording the lines with real people, and — where a different voice colour was needed — converting the voice while keeping the speaker's own timing and pronunciation. Check every recording to its last syllable before it goes near a face: on our film one short call had been cut off mid-word in the source file itself, and no amount of lip-sync would have fixed it.

Measure it, every time you rebuild

Sync is a number, so treat it as one. On our film every build measured the delay between voice and lips on sampled shots against a limit of 120 milliseconds, and checked that every line was heard over its own picture rather than the shot beside it. Those checks ran before anyone watched, which is how a voice that had slipped after an edit was caught in minutes instead of in the hall.

SymptomCauseFix
The mouth moves but says nothingFace generated silently, synced afterwardsDrive the performance with the line's recording, then sync
The voice drifts off the lipsModel re-synthesised the reference audioKeep the original recording; sync to it, not to the model's audio
Two mouths move at onceTwo faces in a speaking shotOne speaker per shot; cut between singles
The face looks absentEyes half-closed while talkingAsk for eyes open and a blink or two; check the frames
Sync fails on a shotFace turned to profile while speakingKeep speaking shots front-on; turn after the last word
Tamil words sound wrongText-to-speech or a cloned voice reading new textRecord real people; convert the voice only if needed

What makes an AI talking video look fake, and the fix

The sync is rarely the problem. The performance is, and the performance is decided before the sync ever runs.

If you need a spoken AI video

We make lip-synced films in Tamil and English — a presenter for a product, a founder speaking to camera, or a family film like the one this was learned on. Our rate card is public: ₹6,000 per 30 seconds of finished film, with a ₹40,000 project minimum. See the work, including a 73-second excerpt of the tribute film, on our AI Video page.

Frequently asked questions

Why do AI talking-head videos look fake?

Usually because the face was animated without the voice and lip-synced afterwards, so the mouth moves but the face is not performing the sentence. Other common causes are a soundtrack that is not the voice the face performed to, two faces in a speaking shot, and eyes half-closed while talking.

Does a second lip-sync pass fix a bad AI talking shot?

Rarely. On a shot generated silently and then synced, a second pass with the same audio changed the mouth by almost nothing measurable. Re-make the performance with the line's recording as the reference, then sync once.

Can AI lip-sync work in Tamil?

Yes. Lip-sync fits the mouth to the audio regardless of language. The weak point in Tamil is the voice, not the sync: text-to-speech still mispronounces names and spoken forms, so record real people and sync to those recordings.

How do you make two AI characters talk to each other?

Give each spoken line its own single shot and cut between them. When two faces share a shot generated with a speech reference, the model moves both mouths; generating it silently and syncing only the speaker works, but the face then is not performing the line.

How accurate should lip-sync be?

We hold every build to a delay of 120 milliseconds or less between voice and lips on sampled shots; the worst shot in our last film measured 60 milliseconds.

Have a project in mind?

We design, build, and ship software end-to-end — with a fixed, written quote after a free scoping call.

More from the blog

Keep reading