Multi-Speaker Dialogue - A Whole Scene, One Call
AI multi-speaker dialogue: assign a voice per line, direct delivery with emotion tags like [nervous] and [angry], and get one seamless conversation audio with timestamps.
Last updated: 2026-08-23
One scene, one generation
Ordinary TTS reads one voice at a time. Multi-speaker dialogue takes the whole script in a single call: each line gets its own voice and optional emotion tag, and you get back one continuous conversation — pacing between speakers handled by the model, no manual stitching.
Hear a real result first
This three-line exchange was generated by this studio in one call — two characters, two emotion directions ([nervous], [angry]), zero editing:
A: 你到底把芯片藏在哪了? (Where did you hide the chip?)
B: [nervous] 我可以解释……但不是现在。 (I can explain… just not now.)
A: [angry] 现在就说! (Say it now!)
Emotion tags shape delivery only — they're never spoken aloud, and never leak into subtitles.
What it's for
- Short drama & audio drama: whole scenes in one pass, stable voice per character
- Podcast dialogues: two-host scripts generated directly
- Scenario ads: 15-second conversational spots, one per market language
- Game cutscenes: batch NPC exchanges
How to use it well
- In Voice Studio → Dialogue, enter lines and assign voices
- Tag any line:
[whispers][excited][sighs][laughs][nervous][angry] - Keep each request under 2,000 characters total; split longer scenes
- Clone or design character voices first, then reuse the same IDs across the whole story
Combined with text to speech timestamps, dialogue audio exports precise subtitles too.