Avatar video — a person on screen speaking your script — is the feature people ask about most and understand least. The confusion is not about quality settings or prompts. It is that two completely different technologies are both sold as lip-sync, and choosing the wrong one cannot be fixed afterwards.
This is worth understanding before spending anything, because the difference is visible to every viewer even when they cannot name it.
Two classes, and only one of them looks right
How to tell them apart
| Mouth-repainting | Audio-driven performance | |
|---|---|---|
| What it needs | An existing video clip | One photo or a short clip, plus audio |
| What moves | The mouth region only | Head, jaw, cheeks, expression |
| Reads as | A mouth pasted on a still face | Someone speaking |
| Typical cost | Very cheap per clip | Priced per second, much dearer |
| Good for | Small corrections to real footage | Anything that must look like speech |
Test before you build anything on it
The mistake that costs money is designing a whole content plan around a class of tool without checking that it clears the quality bar. Prompts and settings do not rescue a class mismatch, so the test has to come first.
- Generate one clip of about ten seconds with your own script and your own voice.
- Watch it with the sound off. If the face is still and only the mouth moves, it is mouth-repainting.
- Watch the jaw line and the throat specifically. Real speech moves both.
- Show it to someone who does not know it is generated and ask what they notice.
- Only then decide whether to build a format around it.
The honest alternative
For most Gujarati creators, a cloned voice over matched B-roll outperforms a mediocre avatar. It costs far less, it produces no uncanny quality for a viewer to react to, and the recognition that an avatar was supposed to provide comes from the voice instead.
An avatar earns its cost when a face is genuinely the point — a personal brand, a presenter format, a spokesperson. It is a poor trade for news, devotional, listicles or explainers, where a face adds nothing the visuals were not already doing.
Permission and disclosure
- Never generate a real person — a celebrity, a politician, a neighbour — speaking words they did not say. This is the clearest line in the whole subject.
- Get explicit written permission before generating anyone other than yourself, including family.
- Disclose when a person on screen is generated where a viewer could mistake it for real footage.
- Do not use a generated person to give advice that a real qualified person should be giving.
- Keep children out of it entirely.
What it costs, honestly
Mouth-repainting is cheap enough to be effectively free per clip. Audio-driven performance models are priced per second of output, which makes them an order of magnitude dearer and makes a daily format expensive quickly.
That price gap is the whole reason people try the cheap class first and are disappointed. Knowing that the disappointment is structural rather than a settings problem saves both the money and the week spent trying to prompt around it.
The short version
- Two classes: mouth-repainting and audio-driven performance. Only the second looks like speech.
- If a tool wants a video rather than a photo, it is almost certainly the cheap class.
- Watch the jaw and throat with the sound off. That is the tell.
- No prompt moves one class into the other. Choose correctly first.
- For most creators, a cloned voice over B-roll beats a mediocre avatar.
- Never generate a real person saying something they did not say.