MP3 to Text - Upload Once, Get Word-Level Timing

Convert MP3 audio to text online: automatic language detection across 90+ languages, a transcript with word-level timestamps, and an SRT file. 200MB per upload.

Last updated: 2026-08-26

One upload, three things back

A transcript, word-level timestamps, and an SRT file. Not three operations — three outputs of a single upload.

Most transcription tools give you only the first. The second is where the time actually goes: with per-word timing, subtitles need no alignment pass, editing needs no scrubbing to find the cut, and long audio becomes searchable by word.

Real output, errors included

Tested on a Mandarin storytelling clip from this site. Below is the unedited response:

Voice: 老者说书人Chinese17 s
Transcript:上回书说到,那少年提剑立于城门之下,身后是三千追兵,身前是万丈深渊。他回头笑了一声——诸位,且听下回分解。

Interviews are the most common MP3 that gets transcribed. Long sentences, pauses and spoken word order all test sentence segmentation:

Voice: 播音腔大叔Chinese17 s
Transcript:很多人问我,做了三十年播音,最难的是什么。不是发音,也不是气息,是在你根本不想说话的那天,还能把每一个字送到听众耳朵里。
Detected language: zho  (confidence 88.7%)

Transcript:
上回书说到,那少年提剑立于城门之西,身后是三千追兵,
身前是莫张深渊。他回头笑了一声:“诸位,且听下回分解。

Word timings (first 12):
上[0.42→0.72]  回[0.72→0.88]  书[0.88→1.06]  说[1.06→1.48]
到[1.48→1.84]  ,[1.84→2.62]  那[2.62→2.96]  少[2.96→3.18]
年[3.18→3.40]  提[3.40→3.64]  剑[3.64→3.96]  立[3.96→4.22]

It got two things wrong. The source reads 城门之下 and 万丈深渊; the engine heard 城门之西 and 莫张深渊 — homophone substitutions.

Homophones are a hard limit for every speech engine. We leave the errors on the page because that is more useful than a hand-picked perfect sample: you now know what to expect. Most words are right; proper nouns and set phrases need one pass of your eyes.

The important part: fixing text does not move the timings. Correct the word in the studio and its start and end times are unchanged, so the subtitles still line up.

What word-level timing is actually for

TaskWithout timingsWith them
Producing subtitlesRe-transcribe inside the editor, then align by handExport the SRT, drop it in, done
Finding a specific lineScrub the playhead back and forthSearch the text, click, jump
Cutting a passageAudition repeatedly for the in and out pointsSelect the words, read the exact seconds
Clipping for socialCut by feelCut on sentence boundaries, never mid-word

The first row saves the most time. The SRT is computed during transcription, not transcribed again afterwards — so you never end up with one set of typos in the subtitles and a different set in the transcript.

Speaker separation is on by default

If more than one person is speaking, the system separates them automatically and labels who said what, up to 32 speakers. Interviews and two-host podcasts come back already shaped as dialogue.

See meeting transcription for how that performs and where it struggles.

How to use it

  1. Open the voice studio → transcribe
  2. Upload the file (MP3/WAV/M4A/AAC/OGG/FLAC, 200MB total per upload)
  3. Wait — an hour of audio is typically one to two minutes
  4. Fix any wrong words, then export the transcript or the SRT

No language to select, and speaker separation and audio-event tagging are on by default.

FAQ

Which formats, and how large? MP3, WAV, M4A, AAC, OGG, FLAC — up to 200MB per upload.

How accurate? Above 95% on clean single-speaker audio. Homophones are the hard limit — the sample above has two.

Do I pick the language? No. Detected automatically with a confidence score; mixed-language audio works.

Can I get subtitles? Yes — word timings export straight to SRT, no alignment pass.

Start here

Cost: 20 credits per file (up to 200MB per upload), refunded automatically if it fails. The $3.99 trial pack has 150 credits, enough for 7 files, and they never expire.

Upload an MP3 →

Related: video to text · meeting transcription · subtitle generator · speech to text overview