Generate SRT Subtitles from Audio: Frame-Accurate in Minutes
Turn an MP3, a recording or a voiceover into SRT: upload, speaker separation, word-level timestamps, edit text without breaking timing, export anywhere.
Listen to this article
AI narration · about 2 min
Timing subtitles by hand takes ten minutes per minute of audio and still drifts. Speech-to-text writes the SRT directly, with timing from word-level timestamps rather than guesses. Here is the full flow and the three places it goes wrong.
The flow
- Upload the audio to the subtitle generator — MP3, WAV, M4A; video files have the audio track extracted automatically.
- Pick the language or let it detect, and turn on speaker separation for conversations.
- Receive the SRT plus word timestamps; fixing a word does not move the timing.
- Import into CapCut, Premiere or DaVinci.
Source audio in, cues out
Two speakers alternating, pauses between lines, every cue landing on the exact in and out point:
Cues break per speaker and at any pause longer than 0.8 seconds.
Three things to watch
- Filler words: "um" and "you know" are kept by default and can be stripped in one click before export.
- Number formatting: "twenty twenty-six" is normalised to "2026"; switch it off in settings if you prefer words.
- Proper nouns: add names and brands to the recognition list in the pronunciation dictionary; accuracy improves noticeably.
Subtitles straight from voiceover
If the audio is AI-generated in the first place, skip transcription: text to speech exports the SRT alongside the audio, aligned per word with zero drift.
Related
Why timing drifts is explained in the subtitle timing guide; meeting recordings are covered in meeting transcription; video files go through MP4 to text.