Picking the Right TTS Model: Flagship vs Standard vs Flash
Same script, 2x cost difference and 10x latency difference across models. What belongs on flagship, what belongs on flash — explained.
Listen to this article
AI narration · about 2 min
Open the model dropdown in the Voice Studio and you'll see three families. A wrong pick won't break anything — it just costs more or performs less.
What each family is for
Flagship: most expressive, honors emotion tags like [whispers] and [excited], 70+ languages, ~5,000 chars per request. For drama, ads, audiobooks — anything that's performed rather than read.
Multilingual standard: stable and balanced, 29 languages, 10,000 chars. The default for voiceovers, explainers and courses.
Flash: lowest latency, half the cost, 32 languages, up to 40,000 chars. The economical pick for bulk long-form and informational content.
A one-question rule
Ask: does this line need to be acted?
- Yes (emotion, characters, drama) → flagship
- No (just deliver the information) → standard, switch to flash at volume
One sentence, three models
What the spec table cannot tell you, your ears can. The same Chinese line, the same voice, rendered on the standard, fast and flagship tiers — the flagship take carries two emotion tags, because only that tier supports them:
[curious] 同样一句话,三个模型读出来,你能听出差别吗?[excited] 稳定、自然、还是更有戏。Details worth knowing
- Only flagship honors emotion tags; standard/flash skip them
- All three return character-level timestamps — subtitles work everywhere
- The same voice sounds slightly different across models; A/B them once for important projects
Per-operation credit prices are on the pricing page; failed generations refund automatically.