AI Voiceover Emotion Tags Cheatsheet: Make TTS Act
[whispers] [excited] [laughs] [sighs]: how to write emotion tags for Eleven v3 and v4, where to put them and which combos work, with real A/B samples.
Listen to this article
AI narration · about 2 min
Same line, four tags, four scenes. That's the mechanism behind the emotion samples on the homepage — the flagship model treats bracketed words as performance direction: never read aloud, only acted.
Common tags
| Category | Tags | Effect |
|---|---|---|
| Volume & breath | [whispers] [shouting] [sighs] [exhales] | whisper, shout, sigh, exhale |
| Emotion | [excited] [sad] [angry] [nervous] [curious] [sarcastic] | as named |
| Reactions | [laughs] [laughs harder] [crying] [gasps] | laugh, big laugh, cry, gasp |
| Pacing | ellipses ... / CAPITALS | pause / emphasis (punctuation, not tags) |
Hear what a tag does
[whispers]:
[whispers] 别出声,他还在门外。[laughs]:
[laughs] 你居然真的信了?我随口说的啊。[sighs]:
[sighs] 算了,这件事就到这里吧。The lines themselves are neutral — every difference above comes from the tag.
Writing them well
- Place tags at the start of a sentence or at a turn:
[whispers] They said the signal was gone. [excited] But look — it's still transmitting! - Two or three tags per line at most — more cancel each other out
- Drop stability to ~0.3 so the model has room to act; above 0.8 tags barely register
- Only the flagship model honors tags; standard/flash ignore them (and won't read them either)
- Tags never leak into subtitles — they're stripped at export
It depends on the voice
Per the docs, tag response depends on a voice's training data. In our tests [whispers] barely moved some premade voices while [shouting] was obvious on most. Try the same tag on two or three voices before rewriting the script. Voices tagged characters_animation in the library usually respond best.
A real comparison
The four Chinese clips under the homepage Emotions tab use one voice (Yun, Mandarin anchor), one line, different tags. Listen, then open the studio, pick the flagship model and click the emotion chips to insert tags.
The same tags work on ElevenLabs' newer Eleven v4 model. Hear v4 perform them, or compare it with v3 on identical scripts: in our test neither model read a tag aloud in 42 generations.