Back to blog

AI Voiceover Emotion Tags Cheatsheet: Make TTS Act

[whispers] [excited] [laughs] [sighs]: how to write emotion tags for Eleven v3 and v4, where to put them and which combos work, with real A/B samples.

Aug 23, 2026VoiceSmiths TeamVoiceSmiths Team

Listen to this article

AI narration · about 2 min
0:00 / –:––
AI Voiceover Emotion Tags Cheatsheet: Make TTS Act

Same line, four tags, four scenes. That's the mechanism behind the emotion samples on the homepage — the flagship model treats bracketed words as performance direction: never read aloud, only acted.

Common tags

CategoryTagsEffect
Volume & breath[whispers] [shouting] [sighs] [exhales]whisper, shout, sigh, exhale
Emotion[excited] [sad] [angry] [nervous] [curious] [sarcastic]as named
Reactions[laughs] [laughs harder] [crying] [gasps]laugh, big laugh, cry, gasp
Pacingellipses ... / CAPITALSpause / emphasis (punctuation, not tags)

Hear what a tag does

[whispers]:

Voice: 知性御姐Chinese3 s
Transcript:[whispers] 别出声,他还在门外。

[laughs]:

Voice: 俏皮甜妹Chinese4 s
Transcript:[laughs] 你居然真的信了?我随口说的啊。

[sighs]:

Voice: 深夜电台主播Chinese5 s
Transcript:[sighs] 算了,这件事就到这里吧。

The lines themselves are neutral — every difference above comes from the tag.

AI Voiceover Emotion Tags Cheatsheet: Make TTS Act

Writing them well

  1. Place tags at the start of a sentence or at a turn: [whispers] They said the signal was gone. [excited] But look — it's still transmitting!
  2. Two or three tags per line at most — more cancel each other out
  3. Drop stability to ~0.3 so the model has room to act; above 0.8 tags barely register
  4. Only the flagship model honors tags; standard/flash ignore them (and won't read them either)
  5. Tags never leak into subtitles — they're stripped at export

It depends on the voice

Per the docs, tag response depends on a voice's training data. In our tests [whispers] barely moved some premade voices while [shouting] was obvious on most. Try the same tag on two or three voices before rewriting the script. Voices tagged characters_animation in the library usually respond best.

A real comparison

The four Chinese clips under the homepage Emotions tab use one voice (Yun, Mandarin anchor), one line, different tags. Listen, then open the studio, pick the flagship model and click the emotion chips to insert tags.

The same tags work on ElevenLabs' newer Eleven v4 model. Hear v4 perform them, or compare it with v3 on identical scripts: in our test neither model read a tag aloud in 42 generations.