Picking the Right TTS Model: Flagship vs Standard vs Flash
Same script, 2x cost difference and 10x latency difference across models. What belongs on flagship, what belongs on flash — explained.
Listen to this article
AI narration · about 2 minOpen the model dropdown in the Voice Studio and you'll see three families. A wrong pick won't break anything — it just costs more or performs less.
What each family is for
Flagship: most expressive, honors emotion tags like [whispers] and [excited], 70+ languages, ~5,000 chars per request. For drama, ads, audiobooks — anything that's performed rather than read.
Multilingual standard: stable and balanced, 29 languages, 10,000 chars. The default for voiceovers, explainers and courses.
Flash: lowest latency, half the cost, 32 languages, up to 40,000 chars. The economical pick for bulk long-form and informational content.
A one-question rule
Ask: does this line need to be acted?
- Yes (emotion, characters, drama) → flagship
- No (just deliver the information) → standard, switch to flash at volume
Details worth knowing
- Only flagship honors emotion tags; standard/flash skip them
- All three return character-level timestamps — subtitles work everywhere
- The same voice sounds slightly different across models; A/B them once for important projects
Per-operation credit prices are on the pricing page; failed generations refund automatically.