Subtitle Generator: Timings That Line Up
Generate SRT subtitles from audio or video. Timings are measured per word during transcription, not estimated afterwards, so they align with no second pass.
Last updated: 2026-08-26
Subtitles drift because the timings were guessed
The usual automatic-caption pipeline recognises the words, then distributes the time evenly across them based on speaking rate. Real speech has pauses, held vowels and filler — so even distribution drifts. The first half of a line disappears early; the second half arrives half a second late.
The timings here are not distributed. They are measured per word during transcription, to 0.01 seconds.
How precise that actually is
Real word timings from a Mandarin storytelling clip:
The hard part of subtitles is timing: two speakers alternating, pauses between lines, and every cue has to land on the exact in and out point. Listen to this exchange, then see how the transcript splits into cues:
上[0.42→0.72] 回[0.72→0.88] 书[0.88→1.06] 说[1.06→1.48]
到[1.48→1.84] ,[1.84→2.62] 那[2.62→2.96] 少[2.96→3.18]
年[3.18→3.40] 提[3.40→3.64] 剑[3.64→3.96] 立[3.96→4.22]
Look at the sixth entry — the comma. It occupies 0.78 seconds, longer than any word before it. That is the storyteller's pause.
Distribute time evenly and those 0.78 seconds get smeared across the surrounding words, and every cue after that point drifts. Measured timings record the pause as it happened, so everything downstream stays put.
Two paths — take the right one
| Where your audio came from | Path | Why |
|---|---|---|
| Generated here | Use the caption track from synthesis | Timings are exact — synthesis knows when each word is emitted |
| Recordings, outside material | Upload and transcribe, export SRT | Timings are measured by recognition, word level |
The first row is easy to miss: if the voiceover was generated here, do not run recognition over it. That swaps exact values for estimated ones and can add transcription errors on top. The caption track already exists.
Correcting words never moves the timings
Recognition is never perfect — homophones are a hard limit. But the words and the timings are two separate pieces of data.
Fix a word in the text and its start and end times are unchanged. So the right proofreading order is bulk-correct the text, then export the SRT — not click through cues one at a time inside an editor.
That is the practical difference from editor auto-captions, where text and timing are welded together.
How to use it
- Open the voice studio → transcribe
- Upload audio or video (200MB total per upload)
- Read it through, concentrating on names, jargon and code-switching
- Export the SRT and drop it into CapCut, Premiere or Final Cut
One ordering trap: lock the edit before transcribing. Re-cutting afterwards leaves the timings misaligned against the new picture.
FAQ
Different from my editor's auto-captions? There, text and timings are locked together and corrected one cue at a time. Here you bulk-edit text and export; timings are untouched.
How precise? Word level, 0.01s. Breaks land on word boundaries, not estimated midpoints.
Can I split long cues? Yes — split at any word and read the exact seconds.
Do I transcribe my own generated voiceover? No. Synthesis already emitted exact timings.
Start here
Cost: 20 credits per file (up to 200MB per upload), refunded automatically if it fails. The $3.99 trial pack has 150 credits, enough for 7 files, and they never expire.
Upload material and generate subtitles →
Related: MP4 to text · MP3 to text · meeting transcription · subtitle timing guide