Text to Speech - Online TTS with Subtitle Timing
Online text to speech on Eleven v4: 90+ languages, emotion tags, SRT subtitles timed to the character. Try it free with no sign-up; from $3.99, no subscription.
Last updated: 2026-10-08
VoiceSmiths is an online text-to-speech studio built on ElevenLabs' newest models, Eleven v4 included. Paste a script, pick one of 54 curated voices or 5,000+ community voices, and download an MP3 plus SRT subtitles timed to the character. Try it below for free with no account; after that, credits start at $3.99 one-time, with no subscription.
Try Eleven v4 now: free, no sign-up
2 free generations a day, up to 200 characters each. Tags in [brackets] are performed, not read aloud.
Text to speech at a glance
| VoiceSmiths text to speech | |
|---|---|
| Models | Eleven v4, Eleven v3, Multilingual v2, Flash v2.5 (official ElevenLabs API) |
| Languages | 90+ on Eleven v4, 70+ on v3 |
| Voices | 54 curated (24 original characters, 30 hand-picked), 5,000+ community voices, cloning and voice design |
| Length | 5,000 characters per generation; 50,000 in one long-form job |
| Output | MP3, plus character-level timestamps and SRT subtitles |
| Emotion control | Inline tags such as [whispers], [excited], [laughs], never read aloud |
| Free | 2 generations a day, 200 characters each, no account |
| Price | 7 credits per 100 characters (Flash: 4). About $0.47–0.70 per minute of English speech on monthly plans |
| Plans | $3.99 one-time for 150 credits (never expire), or from $7 a month for 600 credits |
| Payment | Card (Stripe) or Alipay |
| Commercial use | Yes, under the Terms of Service |
Prices as of October 2026; the current price is always printed on the generate button.
Text to speech that ships finished content
Most TTS tools stop at "it reads the words." Creators need something they can publish. VoiceSmiths's text to speech differs in three ways:
1. Character-level timing, frame-accurate subtitles
Every generation returns the start and end time of each character, to the millisecond. SRT subtitles aren't sliced by estimated reading speed — they're strictly aligned to the audio. Captioned shorts, chaptered podcasts, line-by-line course highlights: one generation covers all of it.
2. Direct the performance
Eleven v4 and v3 take inline emotion direction, written straight into the script:
[whispers] I shouldn't be here.
[excited] But look at this number — it tripled!
Tags like [whispers], [excited], [sighs] and [laughs] shape delivery only. They're never read aloud, and never leak into your subtitles. The full list is in the emotion tags cheat sheet.
The same line, plain versus tagged:
No tag · calm female voice:
[excited] · livestream host:
[excited] 最后三十秒!库存只剩最后两百件,拍下立减一百,手慢无!Broadcast news read:
The same workflow reading English news:
3. Native delivery in 90+ languages
One studio generates English, Chinese, Spanish, Japanese, Portuguese and dozens more. Note: TTS does not translate — provide the script in the target language. For transcribe + translate + re-voice in one job, use video dubbing.
Picking a model
| Model | Studio name | Languages | Strength | Best for |
|---|---|---|---|---|
| Eleven v4 | Eleven v4 | 90+ | Newest, most expressive, tags | Drama, ads, audiobooks |
| Eleven v3 | Resona Ultra | 70+ | Expressive, tags | Character work |
| Multilingual v2 | Resona Multi | 29 | Stable and balanced | Voiceovers, explainers, courses |
| Flash v2.5 | Resona Flash | 32 | Lowest latency, about half the price | Bulk and long-form |
Switch anytime in the Voice Studio — the credit price sits right on the button, and failures refund automatically. Not sure which one? Our v4 vs v3 test ran the same scripts through both, and Google's Gemini 3.8 TTS has a free demo for comparison.
Typical workflow
- Write (or translate) the script in the target language
- Pick a voice — or clone your own, or design one from a text description
- Generate → preview → download MP3 + SRT
FAQ
Is VoiceSmiths text to speech free? You can try it free without an account: two generations a day, up to 200 characters each, on this page or the homepage. Every voice sample on the site is free to play. Beyond that, generation uses credits; there is no free monthly allowance.
How much does text to speech cost? 7 credits per 100 characters on the standard models (Eleven v4, v3, Multilingual v2) and 4 on Flash. 1,000 English characters, about 70 seconds of speech, cost 70 credits: about $0.55–0.82 on the monthly plans, or $1.86 from the one-time $3.99 pack. Failed generations are refunded automatically. See pricing.
Which TTS models can I use? Eleven v4 (released September 28, 2026), Eleven v3, Multilingual v2 and Flash v2.5, all through the official ElevenLabs API. In the studio, v3, Multilingual v2 and Flash are labelled Resona Ultra, Resona Multi and Resona Flash.
How many languages does it support? 90+ on Eleven v4, 70+ on Eleven v3, 32 on Flash v2.5 and 29 on Multilingual v2. Text to speech does not translate, so write the script in the language you want to hear.
How long can one generation be? Up to 5,000 characters per generation in the text-to-speech tool, and up to 50,000 characters in one long-form job, which keeps one voice and one continuous subtitle track.
Can I get subtitles? Yes. Every generation returns the start and end time of each character, so the SRT export is aligned to the audio instead of estimated from reading speed.
Is SSML supported? Eleven v4 and v3 use emotion tags and punctuation instead — ellipses add pauses, capitals add emphasis. For brand names and jargon, use the pronunciation dictionary.
Can I use the audio commercially? Yes. Audio you generate is yours and may be used commercially, including ads, YouTube, podcasts and apps, under the Terms of Service. Cloning a real person's voice requires their consent.
Is VoiceSmiths affiliated with ElevenLabs? No. VoiceSmiths is an independent studio built on the official ElevenLabs API, with its own voices, workflow, pricing and payment options (card or Alipay).