Subtitle Generator: Timings That Line Up

Generate SRT subtitles from audio or video. Timings are measured per word during transcription, not estimated afterwards, so they align with no second pass.

Last updated: 2026-08-26

Subtitles drift because the timings were guessed

The usual automatic-caption pipeline recognises the words, then distributes the time evenly across them based on speaking rate. Real speech has pauses, held vowels and filler — so even distribution drifts. The first half of a line disappears early; the second half arrives half a second late.

The timings here are not distributed. They are measured per word during transcription, to 0.01 seconds.

How precise that actually is

Real word timings from a Mandarin storytelling clip:

Voice: 老者说书人Chinese17 s
Transcript:上回书说到,那少年提剑立于城门之下,身后是三千追兵,身前是万丈深渊。他回头笑了一声——诸位,且听下回分解。

The hard part of subtitles is timing: two speakers alternating, pauses between lines, and every cue has to land on the exact in and out point. Listen to this exchange, then see how the transcript splits into cues:

Voice: 霸道总裁 / 冷艳女王Chinese16 s
Transcript:你知道我为什么让你留下来吗? 知道。因为整个公司,只有我不怕你。 错。是因为只有你,敢在会上说我错了。
上[0.42→0.72]  回[0.72→0.88]  书[0.88→1.06]  说[1.06→1.48]
到[1.48→1.84]  ,[1.84→2.62]  那[2.62→2.96]  少[2.96→3.18]
年[3.18→3.40]  提[3.40→3.64]  剑[3.64→3.96]  立[3.96→4.22]

Look at the sixth entry — the comma. It occupies 0.78 seconds, longer than any word before it. That is the storyteller's pause.

Distribute time evenly and those 0.78 seconds get smeared across the surrounding words, and every cue after that point drifts. Measured timings record the pause as it happened, so everything downstream stays put.

Two paths — take the right one

Where your audio came fromPathWhy
Generated hereUse the caption track from synthesisTimings are exact — synthesis knows when each word is emitted
Recordings, outside materialUpload and transcribe, export SRTTimings are measured by recognition, word level

The first row is easy to miss: if the voiceover was generated here, do not run recognition over it. That swaps exact values for estimated ones and can add transcription errors on top. The caption track already exists.

Correcting words never moves the timings

Recognition is never perfect — homophones are a hard limit. But the words and the timings are two separate pieces of data.

Fix a word in the text and its start and end times are unchanged. So the right proofreading order is bulk-correct the text, then export the SRT — not click through cues one at a time inside an editor.

That is the practical difference from editor auto-captions, where text and timing are welded together.

How to use it

  1. Open the voice studio → transcribe
  2. Upload audio or video (200MB total per upload)
  3. Read it through, concentrating on names, jargon and code-switching
  4. Export the SRT and drop it into CapCut, Premiere or Final Cut

One ordering trap: lock the edit before transcribing. Re-cutting afterwards leaves the timings misaligned against the new picture.

FAQ

Different from my editor's auto-captions? There, text and timings are locked together and corrected one cue at a time. Here you bulk-edit text and export; timings are untouched.

How precise? Word level, 0.01s. Breaks land on word boundaries, not estimated midpoints.

Can I split long cues? Yes — split at any word and read the exact seconds.

Do I transcribe my own generated voiceover? No. Synthesis already emitted exact timings.

Start here

Cost: 20 credits per file (up to 200MB per upload), refunded automatically if it fails. The $3.99 trial pack has 150 credits, enough for 7 files, and they never expire.

Upload material and generate subtitles →

Related: MP4 to text · MP3 to text · meeting transcription · subtitle timing guide