Captioning Shorts Properly: What Character-Level Timing Means
Most tools estimate subtitle timing from average reading speed, so captions drift. How character-level timestamps make subtitles frame-accurate, plus the full script-to-captioned-video flow.
Listen to this article
AI narration · about 2 minWhen captions lag the voice by half a beat, viewers can't name the problem — but your completion rate reports it. The cause: most tools estimate subtitle timing by spreading text evenly at an average reading speed, while real speech speeds up, slows down and pauses.
What character-level timing is
When VoiceSmiths generates a voiceover, the API also returns the start and end time of every character, to the millisecond:
The [0.08s→0.14s] rain [0.16s→0.34s] had [0.36s→0.50s] ...
Subtitles stop being estimates — they're derived from the actual audio. Line breaks and disappearance timing follow real speech.
The full workflow
- Generate: paste your script in the Voice Studio, pick a voice — timing comes included
- Download SRT: one click in the history panel; cue-splitting by punctuation and duration is automatic
- Import: CapCut / Premiere / Resolve take the SRT with zero adjustment
Three common variants:
- Have a recording and its script: use forced alignment — more precise than transcription, zero typos
- Only have audio/video: use speech to text — auto language detection with word timestamps
- Multi-language distribution: generate each translated script separately; every language is frame-accurate on its own
One detail
Emotion tags like [excited] shape delivery but are never spoken — and we strip them at export, so subtitles stay clean.
Try a free line on the homepage and hear it yourself.