See it before you pay: paste a story and get its full cast and voices free. No account needed.

Multi-Speaker Dialogue - A Whole Scene, One Call

Multi-speaker text to speech: one voice per line, emotion tags like [nervous] and [angry], one seamless track with timestamps. Podcasts, drama, ads; from $3.99.

Last updated: 2026-10-08

Multi-speaker dialogue generates a whole scene in one call: every line gets its own voice and optional emotion tag, and you get one continuous track with timestamps. It runs on ElevenLabs' text-to-dialogue model (Eleven v3), takes up to 2,000 characters per generation, and costs the same as single-voice speech, 7 credits per 100 characters. Have a whole story rather than a scene? The audio drama studio finds the characters and casts them for you.

At a glance

Multi-speaker dialogue
InputLines of script, a voice per line, optional tag per line
ModelElevenLabs text-to-dialogue, Eleven v3 (official API)
LengthUp to 2,000 characters per generation (about 2 min 20 s of English)
VoicesAny of the 54 curated voices, 5,000+ community voices, or your cloned voices
Emotion tags[whispers] [excited] [sighs] [laughs] [nervous] [angry]…
OutputOne MP3 for the whole scene, with timestamps
Price7 credits per 100 characters; a full 2,000-character scene ≈ $1.11–1.63 on monthly plans
Plans$3.99 one-time for 150 credits, or from $7 a month; card or Alipay
Whole chaptersAudio drama studio: automatic casting, up to 8,000 characters per render on Eleven v4

Prices as of October 2026.

One scene, one generation

Ordinary TTS reads one voice at a time. Multi-speaker dialogue takes the whole script in a single call: each line gets its own voice and optional emotion tag, and you get back one continuous conversation — pacing between speakers handled by the model, no manual stitching.

Hear a real result first

This three-line exchange was generated by this studio in one call — two characters, two emotion directions ([nervous], [angry]), zero editing:

Chinese9 s
Transcript:你到底把芯片藏在哪儿了? 我可以解释,但不是现在。 现在就说。

Three characters on one track, emotion tags included — this is what multi-speaker dialogue looks like in full:

Voice: George / Alice / JessicaEnglish17 s
Transcript:I've read the contract. The terms are negotiable. The people are not. How charming. You think this is a negotiation? It is a notice. Um — maybe some tea first? That rain isn't stopping any time soon. [laughs] Fine. Nobody leaves until it does.

The same scene in Chinese:

Voice: 霸道总裁 / 冷艳女王 / 甜美治愈Chinese30 s
Transcript:合同我看过了,条件可以谈,但人,一个都不能动。 有意思。你以为这是在谈判?这是通知。 两位……要不先喝口茶?外面的雨,一时半会儿停不了。 [laughs] 也好。雨停之前,谁也别想走。
A: 你到底把芯片藏在哪了? (Where did you hide the chip?)
B: [nervous] 我可以解释……但不是现在。 (I can explain… just not now.)
A: [angry] 现在就说! (Say it now!)

Emotion tags shape delivery only — they're never spoken aloud, and never leak into subtitles.

What it's for

  • Short drama & audio drama: whole scenes in one pass, stable voice per character
  • Podcast dialogues: two-host scripts generated directly
  • Scenario ads: 15-second conversational spots, one per market language
  • Game cutscenes: batch NPC exchanges

How to use it well

  1. In Voice Studio → Dialogue, enter lines and assign voices
  2. Tag any line: [whispers] [excited] [sighs] [laughs] [nervous] [angry]
  3. Keep each request under 2,000 characters total; split longer scenes
  4. Clone or design character voices first, then reuse the same IDs across the whole story

Combined with text to speech timestamps, dialogue audio exports precise subtitles too.

FAQ

Can AI text to speech do several speakers in one file? Yes. In multi-speaker dialogue you give each line its own voice and optional emotion tag, and the whole scene is generated in one call as one continuous track, with the timing between speakers handled by the model. It uses ElevenLabs text-to-dialogue (Eleven v3) through the official API.

How long can one dialogue be? Up to 2,000 characters per generation, about 2 minutes 20 seconds of English. For a whole chapter, the audio drama studio splits the text, casts every character and renders up to 8,000 characters at a time on Eleven v4.

How much does a dialogue cost? The same as single-voice speech: 7 credits per 100 characters. A full 2,000-character scene is 140 credits, about $1.11–1.63 on the monthly plans. Failed generations are refunded. See pricing.

Is it free to try? Multi-speaker generation uses credits, from $3.99 one-time. Free without an account: the samples on this page, two single-voice generations a day on the homepage, and finding the characters in a story in the audio drama studio.

Which emotion tags work? Tags such as [whispers], [excited], [sighs], [laughs], [nervous] and [angry], written in square brackets before the words. They shape the delivery and are never read aloud or written into subtitles. More in the emotion tags cheat sheet.

Can I make a two-host podcast with it? Yes. Write the script as alternating lines, pick a voice for each host, and generate. Keep each request under 2,000 characters and generate a long episode in segments.

Generate my first scene →