Back to blog

AI Voiceover Emotion Tags Cheatsheet: Make TTS Act

[whispers] [excited] [sad] [angry] [laughs] [sighs] — how to write emotion tags for the flagship model, where to put them, which combos work, with real A/B samples.

Aug 23, 2026VoiceSmiths TeamVoiceSmiths Team

Listen to this article

AI narration · about 2 min
0:00 / –:––

Same line, four tags, four scenes. That's the mechanism behind the emotion samples on the homepage — the flagship model treats bracketed words as performance direction: never read aloud, only acted.

Common tags

| Category | Tags | Effect | | --- | --- | --- | | Volume & breath | [whispers] [shouting] [sighs] [exhales] | whisper, shout, sigh, exhale | | Emotion | [excited] [sad] [angry] [nervous] [curious] [sarcastic] | as named | | Reactions | [laughs] [laughs harder] [crying] [gasps] | laugh, big laugh, cry, gasp | | Pacing | ellipses ... / CAPITALS | pause / emphasis (punctuation, not tags) |

Writing them well

  1. Place tags at the start of a sentence or at a turn: [whispers] They said the signal was gone. [excited] But look — it's still transmitting!
  2. Two or three tags per line at most — more cancel each other out
  3. Drop stability to ~0.3 so the model has room to act; above 0.8 tags barely register
  4. Only the flagship model honors tags; standard/flash ignore them (and won't read them either)
  5. Tags never leak into subtitles — they're stripped at export

It depends on the voice

Per the docs, tag response depends on a voice's training data. In our tests [whispers] barely moved some premade voices while [shouting] was obvious on most. Try the same tag on two or three voices before rewriting the script. Voices tagged characters_animation in the library usually respond best.

A real comparison

The four Chinese clips under the homepage Emotions tab use one voice (Yun, Mandarin anchor), one line, different tags. Listen, then open the studio, pick the flagship model and click the emotion chips to insert tags.