MP3 to Text - Upload Once, Get Word-Level Timing
Convert MP3 audio to text online: automatic language detection across 90+ languages, a transcript with word-level timestamps, and an SRT file. 200MB per upload.
Last updated: 2026-08-26
One upload, three things back
A transcript, word-level timestamps, and an SRT file. Not three operations — three outputs of a single upload.
Most transcription tools give you only the first. The second is where the time actually goes: with per-word timing, subtitles need no alignment pass, editing needs no scrubbing to find the cut, and long audio becomes searchable by word.
Real output, errors included
Tested on a Mandarin storytelling clip from this site. Below is the unedited response:
Interviews are the most common MP3 that gets transcribed. Long sentences, pauses and spoken word order all test sentence segmentation:
Detected language: zho (confidence 88.7%)
Transcript:
上回书说到,那少年提剑立于城门之西,身后是三千追兵,
身前是莫张深渊。他回头笑了一声:“诸位,且听下回分解。
Word timings (first 12):
上[0.42→0.72] 回[0.72→0.88] 书[0.88→1.06] 说[1.06→1.48]
到[1.48→1.84] ,[1.84→2.62] 那[2.62→2.96] 少[2.96→3.18]
年[3.18→3.40] 提[3.40→3.64] 剑[3.64→3.96] 立[3.96→4.22]
It got two things wrong. The source reads 城门之下 and 万丈深渊; the engine heard 城门之西 and 莫张深渊 — homophone substitutions.
Homophones are a hard limit for every speech engine. We leave the errors on the page because that is more useful than a hand-picked perfect sample: you now know what to expect. Most words are right; proper nouns and set phrases need one pass of your eyes.
The important part: fixing text does not move the timings. Correct the word in the studio and its start and end times are unchanged, so the subtitles still line up.
What word-level timing is actually for
| Task | Without timings | With them |
|---|---|---|
| Producing subtitles | Re-transcribe inside the editor, then align by hand | Export the SRT, drop it in, done |
| Finding a specific line | Scrub the playhead back and forth | Search the text, click, jump |
| Cutting a passage | Audition repeatedly for the in and out points | Select the words, read the exact seconds |
| Clipping for social | Cut by feel | Cut on sentence boundaries, never mid-word |
The first row saves the most time. The SRT is computed during transcription, not transcribed again afterwards — so you never end up with one set of typos in the subtitles and a different set in the transcript.
Speaker separation is on by default
If more than one person is speaking, the system separates them automatically and labels who said what, up to 32 speakers. Interviews and two-host podcasts come back already shaped as dialogue.
See meeting transcription for how that performs and where it struggles.
How to use it
- Open the voice studio → transcribe
- Upload the file (MP3/WAV/M4A/AAC/OGG/FLAC, 200MB total per upload)
- Wait — an hour of audio is typically one to two minutes
- Fix any wrong words, then export the transcript or the SRT
No language to select, and speaker separation and audio-event tagging are on by default.
FAQ
Which formats, and how large? MP3, WAV, M4A, AAC, OGG, FLAC — up to 200MB per upload.
How accurate? Above 95% on clean single-speaker audio. Homophones are the hard limit — the sample above has two.
Do I pick the language? No. Detected automatically with a confidence score; mixed-language audio works.
Can I get subtitles? Yes — word timings export straight to SRT, no alignment pass.
Start here
Cost: 20 credits per file (up to 200MB per upload), refunded automatically if it fails. The $3.99 trial pack has 150 credits, enough for 7 files, and they never expire.
Related: video to text · meeting transcription · subtitle generator · speech to text overview