AI Video Dubbing Guide: Keep the Voice, Change the Language
How AI video dubbing works: upload, speaker detection, translation, synthesis in the original voice, timing and lip alignment, per-minute cost, plus a sample.
Listen to this article
AI narration · about 2 min
The promise of AI dubbing is "change the language, keep the person": viewers still hear the original speaker, now in English, Spanish or Japanese. This is the full workflow, the usual traps, and what it costs.
Hear the result first
One voice, Chinese then English:
The workflow
- Upload the video to video dubbing — MP4 or MOV, or paste a YouTube or TikTok link.
- Speakers are detected and transcribed; in multi-speaker videos each keeps their own voice.
- The transcript is translated; edit the translation before export.
- The foreign-language track is synthesised, each line aligned to its original timing, with pace adjusted to fit the mouth.
- Export the video with the new track, or just the track and an SRT.
Three traps
- The music gets "translated" too: turn on voice isolation so music and effects stay as they were.
- Translations run long and miss the lips: choose "fit to duration" and the translation is tightened.
- Proper nouns come out wrong: add them to the pronunciation dictionary before batch processing.
Cost
Billed per minute of video, about 1,000 credits a minute. A three-minute product video into three languages is roughly 9,000 credits — a fraction of three human sessions.
Who it is for
Cross-border product videos, courses going abroad, short drama exports, multilingual YouTube channels. The two routes for drama are compared in short drama dubbing; choosing accents is covered in the multilingual accent guide.