Multi-Speaker Dialogue - A Whole Scene, One Call

AI multi-speaker dialogue: assign a voice per line, direct delivery with emotion tags like [nervous] and [angry], and get one seamless conversation audio with timestamps.

Last updated: 2026-08-23

One scene, one generation

Ordinary TTS reads one voice at a time. Multi-speaker dialogue takes the whole script in a single call: each line gets its own voice and optional emotion tag, and you get back one continuous conversation — pacing between speakers handled by the model, no manual stitching.

Hear a real result first

This three-line exchange was generated by this studio in one call — two characters, two emotion directions ([nervous], [angry]), zero editing:

A: 你到底把芯片藏在哪了? (Where did you hide the chip?)
B: [nervous] 我可以解释……但不是现在。 (I can explain… just not now.)
A: [angry] 现在就说! (Say it now!)

Emotion tags shape delivery only — they're never spoken aloud, and never leak into subtitles.

What it's for

  • Short drama & audio drama: whole scenes in one pass, stable voice per character
  • Podcast dialogues: two-host scripts generated directly
  • Scenario ads: 15-second conversational spots, one per market language
  • Game cutscenes: batch NPC exchanges

How to use it well

  1. In Voice Studio → Dialogue, enter lines and assign voices
  2. Tag any line: [whispers] [excited] [sighs] [laughs] [nervous] [angry]
  3. Keep each request under 2,000 characters total; split longer scenes
  4. Clone or design character voices first, then reuse the same IDs across the whole story

Combined with text to speech timestamps, dialogue audio exports precise subtitles too.

Generate my first scene →