Back to blog

Picking the Right TTS Model: Flagship vs Standard vs Flash

Same script, 2x cost difference and 10x latency difference across models. What belongs on flagship, what belongs on flash — explained.

Aug 20, 2026VoiceSmiths TeamVoiceSmiths Team

Listen to this article

AI narration · about 2 min
0:00 / –:––

Open the model dropdown in the Voice Studio and you'll see three families. A wrong pick won't break anything — it just costs more or performs less.

What each family is for

Flagship: most expressive, honors emotion tags like [whispers] and [excited], 70+ languages, ~5,000 chars per request. For drama, ads, audiobooks — anything that's performed rather than read.

Multilingual standard: stable and balanced, 29 languages, 10,000 chars. The default for voiceovers, explainers and courses.

Flash: lowest latency, half the cost, 32 languages, up to 40,000 chars. The economical pick for bulk long-form and informational content.

A one-question rule

Ask: does this line need to be acted?

  • Yes (emotion, characters, drama) → flagship
  • No (just deliver the information) → standard, switch to flash at volume

Details worth knowing

  • Only flagship honors emotion tags; standard/flash skip them
  • All three return character-level timestamps — subtitles work everywhere
  • The same voice sounds slightly different across models; A/B them once for important projects

Per-operation credit prices are on the pricing page; failed generations refund automatically.