Speech to Text - 90+ Languages, Word Timestamps
AI speech to text with 90+ auto-detected languages, up to 32 separated speakers and word-level timestamps. Turn audio and video into transcripts and SRT subtitles.
Last updated: 2026-08-23
More than "audio in, words out"
Speech to text turns audio or video into a transcript. This engine adds three things on top:
- 90+ languages, auto-detected — no language picker, mixed-language content handled
- Speaker diarization — up to 32 speakers, every word labeled with who said it
- Word-level timestamps — exact start/end per word, one click to SRT subtitles
A real transcription
We ran the Chinese homepage sample through it. Below is the raw, unedited output:
Detected: zho (confidence 98.3%)
Text: 三步做出爆款口播,念帖文案,选择音色,一键到处带字幕的成品。
Word timestamps (first 6):
三 [0.20s→0.44s] 步 [0.44s→0.68s] 做 [0.68s→0.88s]
出 [0.88s→1.06s] 爆 [1.06s→1.28s] 款 [1.28s→1.40s]
Full disclosure: two homophones were misheard in this clip ("粘贴"→"念帖", "导出"→"到处") — homophones remain ASR's frontier. That's exactly why we show raw output: you can edit the text in the studio and the timestamps stay intact.
Typical uses
- Subtitling: transcribe video → download SRT → drop into your editor
- Interviews & meetings: diarization outputs "who said what" directly
- Podcast show notes: one episode, one click, one SEO-friendly transcript
- Voiceover QA: transcribe generated audio back to verify script fidelity
How to use it
Open Voice Studio → Transcribe, upload audio or video (up to 200MB) — diarization is on by default. Then copy the text or download SRT from the history panel.
Already have the script and just need exact timing? Forced alignment is faster and more precise.