Text to Speech - Online TTS with Subtitle Timing
Online text to speech with native delivery in 70+ languages, inline emotion tags, and character-level timestamps that export straight to SRT subtitles.
Last updated: 2026-08-23
Text to speech that ships finished content
Most TTS tools stop at "it reads the words." Creators need something they can publish. VoiceSmiths's text to speech differs in three ways:
1. Character-level timing, frame-accurate subtitles
Every generation returns the start and end time of each character, to the millisecond. SRT subtitles aren't sliced by estimated reading speed — they're strictly aligned to the audio. Captioned shorts, chaptered podcasts, line-by-line course highlights: one generation covers all of it.
2. Direct the performance
The flagship model takes inline emotion direction, written straight into the script:
[whispers] I shouldn't be here.
[excited] But look at this number — it tripled!
Tags like [whispers], [excited], [sighs] and [laughs] shape delivery only. They're never read aloud, and never leak into your subtitles.
3. Native delivery in 70+ languages
One studio generates English, Chinese, Spanish, Japanese, Portuguese and 65+ more. Note: TTS does not translate — provide the script in the target language. For transcribe + translate + re-voice in one job, use video dubbing.
Picking a model
| Model | Strength | Best for | | --- | --- | --- | | Flagship | Most expressive, emotion tags | Drama, ads, audiobooks | | Multilingual standard | Stable and balanced | Voiceovers, explainers, courses | | Flash | Lowest latency, half price | Bulk and long-form |
Switch anytime in the Voice Studio — the credit price sits right on the button, and failures refund automatically.
Typical workflow
- Write (or translate) the script in the target language
- Pick a voice — or clone your own, or design one from a text description
- Generate → preview → download MP3 + SRT
FAQ
How long can one generation be? 5,000–40,000 characters per request depending on the model; split longer scripts into sections.
Is SSML supported? The flagship model uses emotion tags and punctuation instead — ellipses add pauses, capitals add emphasis.
How is it priced? Credits per operation, shown on the button. See pricing for plans.