AI video dubbing is four steps, and the first one decides the rest
By Andrey ChmerevI build Kekoso, which does step one of the four and none of the others, so treat the last section as interested. Prices are from each vendor's own pricing page, read on 5 September 2026; the text-expansion figures are the IBM table published by the W3C. ElevenLabs blocks requests from my network, so nothing here is quoted from their documentation.

Ask an AI dubbing tool for a video “in Spanish” and you are asking for four different things to happen in a row:
- Transcription — the speech becomes text.
- Translation — the text becomes Spanish text.
- Voice synthesis — the Spanish text is spoken, often in a clone of the original voice. (Your Mac can already do a plain version of this step, free and offline.)
- Timing — the new audio is fitted against the picture, sometimes with the lips adjusted to match.
Every step takes the one before it as fact. That is the whole story of why dubbing goes wrong, and it is worth understanding before you pick a tool, because the tools differ mostly in whether they let you interrupt the chain.
The first step is the one that matters
A mistake in transcription does not stay a transcription mistake.
Say the speaker mentions a product called Kestrel and the model hears “castrol”. Step two faithfully translates “castrol”. Step three reads it aloud in a synthetic copy of the speaker’s own voice, with their intonation. Step four syncs it to their lips.
What comes out is a person who appears to say a word they never said, in their own voice, in a language you may not speak well enough to notice. There is no stage in that chain where anything looks wrong — each step did its job correctly on the input it was given.
Compare that with the same error in subtitles: it sits there as text, in a file you can open, and anyone who watches with sound on will catch it. Dubbing removes every opportunity to notice the error and adds a false witness to it.
That is the practical reason to care about the accuracy of step one specifically, and to look for tools that let you read the transcript before the rest of the chain runs.
The second problem is arithmetic
Translated text is longer than the original. This is not a stylistic tendency, it is measurable, and the W3C publishes IBM’s table of average expansion from English into European languages:
| Characters in the English source | Average expansion |
|---|---|
| Up to 10 | 200–300% |
| 11–20 | 180–200% |
| 21–30 | 160–180% |
| 31–50 | 140–160% |
| 51–70 | 151–170% |
| Over 70 | 130% |
A dubbed line has to fit the slot the original occupied, because the picture does not wait. So roughly 30% more text has to be spoken in the same number of seconds, and something has to absorb the difference:
The voice speeds up. The most common choice, and the reason dubbed audio often sounds slightly hurried without you being able to say why.
The translation gets shortened. A human localiser does this deliberately and well — it is a craft with a name, adaptation. A machine does it by dropping whatever fit least.
The timing drifts. Lines start creeping past the shots they belong to.
Subtitles hit the same wall, which is why broadcast standards cap reading speed — around 17 characters a second in the guidelines most captioning follows. But subtitles can be trimmed without lying, because the viewer still hears the original. A shortened dub is the only version the viewer gets.
Where the tools differ
They all do the same four steps. The interesting question is whether you can get inside the chain.
Fully automatic is the default on the newest offerings: upload, choose a language, download. A dubbing AI running this way takes the machine transcript as given. Fast, and you have no way to correct step one before it propagates.
Editable transcript is the option worth looking for. Several services accept an SRT upload or let you edit the transcription and translation before generating audio — Rask lists “SRT download & upload” among its editing features, and its plans include “video translation into any language, emotion-preserving voice cloning”.
That upload path is the useful one: it means you can produce the transcript separately, check it, fix names and terms, and hand the corrected version to the dubbing service. The step you can actually control is the one you do outside their pipeline.
What it costs
HeyGen advertises “175+ languages with realistic AI voices, lip sync, and subtitles”. Its free plan gives 3 videos a month, up to 1 minute each, with 30+ languages. Voice cloning and the full language list start on Creator at $29 a month; Pro is $49 with 4K export.
Rask includes translation and voice cloning on every plan.
Both read on 5 September 2026, and this part of the market changes faster than any other section of this article.
Dub, or subtitle?
Subtitles are cheaper, faster, editable after publication, and they keep the original performance — the actual voice, the actual timing, the actual person. They also fail politely: a bad subtitle is visibly a bad subtitle.
Dubbing earns its cost when the viewer cannot read along: driving, cooking, working with their hands, watching on a phone in daylight, or reading the written language less comfortably than they hear it. For children’s content it is not optional. For a conference talk or a software tutorial, subtitles usually do the job and the dubbing budget buys nothing — and if all you needed was to translate video to English for your own understanding, neither route applies: that is a transcript and a translator, in two steps.
The honest default is subtitles first, dubbing only for the material that has earned an audience which cannot read them.
And the consent part, briefly
A cloned voice belongs to the person it was cloned from. Reputable services require permission for custom voice clones, and several jurisdictions now regulate synthetic likeness directly. A dubbed video of someone saying words they never said is precisely the case those rules exist for. If the voice is not yours, get it in writing before, not after.
Where my part of this ends
Kekoso, which I build, does step one and nothing else: it transcribes the audio on your own Mac and exports the text as TXT, SRT or VTT. It does not translate, does not synthesise speech and does not dub — and here is why that is deliberate.
What it is good for in this workflow is the correctable step: get an accurate transcript first, with names spelled the way you spell them, then hand that to whichever dubbing service you chose. The subtitle export is the file those services accept.
If dubbing is not what you needed after all — if you wanted the words on screen rather than in another voice — that is a shorter path with fewer places to go wrong.
Prices and product claims here come from each vendor’s own pages, read on 5 September 2026. The expansion figures are IBM’s, as published by the W3C.
Questions people ask
How does AI video dubbing actually work?
Four steps in a chain. The audio is transcribed into text, the text is translated, the translation is spoken by a synthetic voice — often cloned from the original speaker — and the result is fitted back against the picture. Some services also adjust lip movement. Each step takes the previous step's output as fact, which is why an error at the start is not caught later.
Why does dubbed audio sound rushed or clipped?
Because translated text is longer than the original and the slot it has to fit is the same length. The IBM expansion table published by the W3C puts average growth from English at around 130% for strings over 70 characters, and far higher for short ones. Something has to give: either the synthetic voice speeds up, or the translation is shortened, or the timing drifts against the picture.
Can I fix the transcript before the video is dubbed?
With some tools, and it is worth looking for the option. Fully automatic dubbing takes the machine transcript as given. Services that let you supply or edit the transcript — several accept an SRT upload — put the one correctable step back in your hands before an error gets spoken aloud in a synthetic voice.
What does AI dubbing cost?
HeyGen's free plan allows 3 videos a month of up to 1 minute each with 30+ languages; voice cloning and the full 175+ languages start on Creator at $29 a month, with Pro at $49. Rask includes video translation and voice cloning on every plan. Prices read from each vendor on 5 September 2026.
Should I dub or just add subtitles?
Subtitles are cheaper, faster, correctable after publishing and keep the original performance. Dubbing wins when the viewer cannot read along — driving, cooking, small phone screens, low literacy in the written language, or accessibility needs of a different kind. For most explainer and course content, subtitles first and dubbing only for the material that earns it.
Is it legal to clone a speaker's voice for dubbing?
Only with that person's permission, and the reputable services require it for custom voice clones. Beyond the vendor's rules, several jurisdictions now regulate synthetic likeness and voice, and a dubbed video of someone saying words they never said is exactly the case those rules are about. If the voice is not yours, get consent in writing.