Meeting Transcription - It Knows Who Said What

Transcribe meeting recordings with automatic speaker separation for up to 32 people. Every line carries a speaker label and timing, across 90+ languages.

Last updated: 2026-08-26

Ninety percent of writing minutes is working out who said it

Turning an hour of audio into text is the easy part. The hard part is the wall of subject-less sentences that comes back — you have to re-listen to attribute anything.

Speaker diarisation solves exactly that step: attribution happens during transcription, so the output arrives shaped as a conversation rather than one undifferentiated block.

Real output

A two-person English exchange, uploaded with no configuration. Unedited response:

English12 s
Transcript:So where did we land on the migration date? I'd push it a week. QA hasn't signed off on the rollback path yet. Fine, a week, but I want the rollback tested by Thursday

A second recording that sounds like a real stand-up: two people trading turns, cutting in, then splitting the work. The transcript below is exactly what transcription has to recover.

Voice: 科技大佬 / 知性御姐Chinese43 s
Transcript:好,我们开始。先同步一下上周的进度:支付这块,Stripe 和支付宝都已经在线上跑通了,第一笔真实订单是昨天下午。 补充一点,退款流程我们也测过了,资金原路退回大概五到十天。剩下的问题是提现,还差一份补充材料。 那这周的目标就两个:把补充材料交上去,然后把素材页的音频全部补齐。谁负责? 素材我来,材料你交。周三下午三点再对一次。
Detected language: eng  (confidence 99.3%)
Speakers detected: 2

speaker_0  @0.16s   So where did we land on the migration date?
speaker_1  @3.56s   I'd push it a week. QA hasn't signed off on
                    the rollback path yet.
speaker_0  @8.54s   Fine, a week, but I want the rollback tested
                    by Thursday

Three turns, all attributed correctly, each with a start time. Nothing was configured — the speaker count was not supplied, and no labelling was done.

Look at the third line: when speaker_0 talks again, the system recognises the same person rather than adding a third. That is the difference between diarisation and simply splitting on pauses.

One honest detail: the source script reads "Fine. A week. But I want…" and came back as "Fine, a week, but I want…". Punctuation is inferred from pauses and is the least stable part of any transcript. It does not matter for subtitles; it matters for a formal record.

Anonymous labels, no voiceprint identification

You get speaker_0 and speaker_1, not names. The guarantee is that one label means one person — the system never asks who that person is.

That is a deliberate limit. Matching voices against identities is biometric processing and should not be switched on by default in a transcription tool. Replacing the labels with names takes ten seconds and applies across the whole transcript.

Where it fits

SettingWhat separation buys youWatch out for
Customer interviewsA clean Q&A record; quotes never misattributedState the count when there are many participants
Two-host podcastsDialogue-shaped output, ready for show notesThe category it handles best
Project standupsWho committed to what, with a timestamp to checkRe-listen to any cross-talk
Formal recordsTurn-taking is separated very accuratelyConfirm compliance for sensitive material

Take that last row seriously: the audio is processed by a speech recognition service. For highly sensitive meetings, confirm that fits your organisation's requirements before uploading.

Where it breaks

Overlapping speech is the known hard problem. Turn-taking is highly accurate; two people speaking at once for more than a second or two can be misattributed.

Two practical mitigations: separate microphones where possible, or re-listen to the disputed passage — with timestamps that is one click, not a hunt along the playhead.

We put this on the page rather than letting you discover it after uploading.

How to use it

  1. Open the voice studio → transcribe
  2. Upload the recording — audio or video, 200MB total per upload
  3. Wait; an hour typically takes one to two minutes
  4. Replace speaker_0 with real names, then export

Speaker separation needs no configuration, and the language is detected automatically across 90+ languages.

FAQ

How many speakers? Declare the count? Up to 32, inferred automatically. Stating a known count improves separation.

Can it name people? No, and it should not — anonymous labels only, no voiceprint identification.

Cross-talk? Separated, but overlaps beyond a second or two can misattribute. Use separate mics or re-listen.

Confidential meetings? Results live under your account and are deletable, but audio is processed by a recognition service — confirm compliance first.

Start here

Cost: 20 credits per file (up to 200MB per upload), refunded automatically if it fails. The $3.99 trial pack has 150 credits, enough for 7 files, and they never expire.

Upload a meeting recording →

Related: MP3 to text · video to text · subtitle generator · multi-speaker dialogue