Meeting Transcription - It Knows Who Said What
Transcribe meeting recordings with automatic speaker separation for up to 32 people. Every line carries a speaker label and timing, across 90+ languages.
Last updated: 2026-08-26
Ninety percent of writing minutes is working out who said it
Turning an hour of audio into text is the easy part. The hard part is the wall of subject-less sentences that comes back — you have to re-listen to attribute anything.
Speaker diarisation solves exactly that step: attribution happens during transcription, so the output arrives shaped as a conversation rather than one undifferentiated block.
Real output
A two-person English exchange, uploaded with no configuration. Unedited response:
A second recording that sounds like a real stand-up: two people trading turns, cutting in, then splitting the work. The transcript below is exactly what transcription has to recover.
Detected language: eng (confidence 99.3%)
Speakers detected: 2
speaker_0 @0.16s So where did we land on the migration date?
speaker_1 @3.56s I'd push it a week. QA hasn't signed off on
the rollback path yet.
speaker_0 @8.54s Fine, a week, but I want the rollback tested
by Thursday
Three turns, all attributed correctly, each with a start time. Nothing was configured — the speaker count was not supplied, and no labelling was done.
Look at the third line: when speaker_0 talks again, the system recognises the same person rather than adding a third. That is the difference between diarisation and simply splitting on pauses.
One honest detail: the source script reads "Fine. A week. But I want…" and came back as "Fine, a week, but I want…". Punctuation is inferred from pauses and is the least stable part of any transcript. It does not matter for subtitles; it matters for a formal record.
Anonymous labels, no voiceprint identification
You get speaker_0 and speaker_1, not names. The guarantee is that one label means one person — the system never asks who that person is.
That is a deliberate limit. Matching voices against identities is biometric processing and should not be switched on by default in a transcription tool. Replacing the labels with names takes ten seconds and applies across the whole transcript.
Where it fits
| Setting | What separation buys you | Watch out for |
|---|---|---|
| Customer interviews | A clean Q&A record; quotes never misattributed | State the count when there are many participants |
| Two-host podcasts | Dialogue-shaped output, ready for show notes | The category it handles best |
| Project standups | Who committed to what, with a timestamp to check | Re-listen to any cross-talk |
| Formal records | Turn-taking is separated very accurately | Confirm compliance for sensitive material |
Take that last row seriously: the audio is processed by a speech recognition service. For highly sensitive meetings, confirm that fits your organisation's requirements before uploading.
Where it breaks
Overlapping speech is the known hard problem. Turn-taking is highly accurate; two people speaking at once for more than a second or two can be misattributed.
Two practical mitigations: separate microphones where possible, or re-listen to the disputed passage — with timestamps that is one click, not a hunt along the playhead.
We put this on the page rather than letting you discover it after uploading.
How to use it
- Open the voice studio → transcribe
- Upload the recording — audio or video, 200MB total per upload
- Wait; an hour typically takes one to two minutes
- Replace speaker_0 with real names, then export
Speaker separation needs no configuration, and the language is detected automatically across 90+ languages.
FAQ
How many speakers? Declare the count? Up to 32, inferred automatically. Stating a known count improves separation.
Can it name people? No, and it should not — anonymous labels only, no voiceprint identification.
Cross-talk? Separated, but overlaps beyond a second or two can misattribute. Use separate mics or re-listen.
Confidential meetings? Results live under your account and are deletable, but audio is processed by a recognition service — confirm compliance first.
Start here
Cost: 20 credits per file (up to 200MB per upload), refunded automatically if it fails. The $3.99 trial pack has 150 credits, enough for 7 files, and they never expire.
Related: MP3 to text · video to text · subtitle generator · multi-speaker dialogue