Multi-Speaker Dialogue - A Whole Scene, One Call
Multi-speaker text to speech: one voice per line, emotion tags like [nervous] and [angry], one seamless track with timestamps. Podcasts, drama, ads; from $3.99.
Last updated: 2026-10-08
Multi-speaker dialogue generates a whole scene in one call: every line gets its own voice and optional emotion tag, and you get one continuous track with timestamps. It runs on ElevenLabs' text-to-dialogue model (Eleven v3), takes up to 2,000 characters per generation, and costs the same as single-voice speech, 7 credits per 100 characters. Have a whole story rather than a scene? The audio drama studio finds the characters and casts them for you.
At a glance
| Multi-speaker dialogue | |
|---|---|
| Input | Lines of script, a voice per line, optional tag per line |
| Model | ElevenLabs text-to-dialogue, Eleven v3 (official API) |
| Length | Up to 2,000 characters per generation (about 2 min 20 s of English) |
| Voices | Any of the 54 curated voices, 5,000+ community voices, or your cloned voices |
| Emotion tags | [whispers] [excited] [sighs] [laughs] [nervous] [angry]… |
| Output | One MP3 for the whole scene, with timestamps |
| Price | 7 credits per 100 characters; a full 2,000-character scene ≈ $1.11–1.63 on monthly plans |
| Plans | $3.99 one-time for 150 credits, or from $7 a month; card or Alipay |
| Whole chapters | Audio drama studio: automatic casting, up to 8,000 characters per render on Eleven v4 |
Prices as of October 2026.
One scene, one generation
Ordinary TTS reads one voice at a time. Multi-speaker dialogue takes the whole script in a single call: each line gets its own voice and optional emotion tag, and you get back one continuous conversation — pacing between speakers handled by the model, no manual stitching.
Hear a real result first
This three-line exchange was generated by this studio in one call — two characters, two emotion directions ([nervous], [angry]), zero editing:
Three characters on one track, emotion tags included — this is what multi-speaker dialogue looks like in full:
[laughs] Fine. Nobody leaves until it does.The same scene in Chinese:
[laughs] 也好。雨停之前,谁也别想走。A: 你到底把芯片藏在哪了? (Where did you hide the chip?)
B: [nervous] 我可以解释……但不是现在。 (I can explain… just not now.)
A: [angry] 现在就说! (Say it now!)
Emotion tags shape delivery only — they're never spoken aloud, and never leak into subtitles.
What it's for
- Short drama & audio drama: whole scenes in one pass, stable voice per character
- Podcast dialogues: two-host scripts generated directly
- Scenario ads: 15-second conversational spots, one per market language
- Game cutscenes: batch NPC exchanges
How to use it well
- In Voice Studio → Dialogue, enter lines and assign voices
- Tag any line:
[whispers][excited][sighs][laughs][nervous][angry] - Keep each request under 2,000 characters total; split longer scenes
- Clone or design character voices first, then reuse the same IDs across the whole story
Combined with text to speech timestamps, dialogue audio exports precise subtitles too.
FAQ
Can AI text to speech do several speakers in one file? Yes. In multi-speaker dialogue you give each line its own voice and optional emotion tag, and the whole scene is generated in one call as one continuous track, with the timing between speakers handled by the model. It uses ElevenLabs text-to-dialogue (Eleven v3) through the official API.
How long can one dialogue be? Up to 2,000 characters per generation, about 2 minutes 20 seconds of English. For a whole chapter, the audio drama studio splits the text, casts every character and renders up to 8,000 characters at a time on Eleven v4.
How much does a dialogue cost? The same as single-voice speech: 7 credits per 100 characters. A full 2,000-character scene is 140 credits, about $1.11–1.63 on the monthly plans. Failed generations are refunded. See pricing.
Is it free to try? Multi-speaker generation uses credits, from $3.99 one-time. Free without an account: the samples on this page, two single-voice generations a day on the homepage, and finding the characters in a story in the audio drama studio.
Which emotion tags work? Tags such as [whispers], [excited], [sighs], [laughs], [nervous] and [angry], written in square brackets before the words. They shape the delivery and are never read aloud or written into subtitles. More in the emotion tags cheat sheet.
Can I make a two-host podcast with it? Yes. Write the script as alternating lines, pick a voice for each host, and generate. Keep each request under 2,000 characters and generate a long episode in segments.