Gemini 3.8 TTS Online - Free Gemini TTS Demo
Try Gemini 3.8 Flash TTS free, no sign-up: pick a voice, describe the delivery, hear it. Real samples, all 30 voices, cost per minute, an Eleven v4 test.
Last updated: 2026-10-08
Gemini 3.8 Flash TTS is Google's newest text-to-speech model, released on September 23, 2026. Instead of picking from emotion presets, you tell it how to sound in plain words, such as "whispering nervously" or "excited sports commentator", and drop sounds into the text with tags like <laugh> and <sigh>. It handles 100+ languages and two speakers per request, and costs about $0.017 per minute of audio at Google's paid rate. Try it below:
Try Gemini 3.8 TTS now: free, no sign-up
2 free generations a day, up to 200 characters. Describe the delivery in plain words; tags like <laugh> are performed, not read.
Runs on Google's free tier: Google may use what you type here to improve its products, so don't enter anything private.
Gemini 3.8 TTS at a glance
| Gemini 3.8 Flash TTS | |
|---|---|
| Made by | |
| Released | September 23, 2026, alongside Gemini 3.8 Flash-Lite TTS |
| Model IDs | gemini-3.8-flash-tts, gemini-3.8-flash-lite-tts |
| Voices | 30 classic prebuilt voices; the voices API lists 2,089 across 30 locales (no Mandarin); plus voice design from a description and voice cloning from a 10–30 s sample |
| Languages | 100+ languages and dialects, detected automatically (the API docs say 130+ for Flash) |
| Directing | A plain-language style per turn, plus inline tags such as <laugh>, <sigh> and <short pause> |
| Speakers | Up to 2 per request |
| Output | WAV, 24 kHz mono 16-bit |
| Price | Free tier with a small daily limit; paid: $0.50 per 1M text tokens in, $9 per 1M audio tokens out (doubles on January 1, 2027) |
| Cost per minute (measured) | About $0.017 now, $0.035 from 2027, at 32 audio tokens per second |
| Watermark | SynthID on every clip |
| Where | Gemini API, Google AI Studio, Gemini Notebook; Flash-Lite in Google Vids |
Listen: real Gemini 3.8 TTS output
Every clip was generated with gemini-3.8-flash-tts on October 5, 2026: one take each, no editing. The transcript is the exact text sent; the style we asked for is written above each player. These are the same scripts we ran through Eleven v4, so you can compare them directly.
English narration. Voice Charon, style "warm, unhurried British storyteller":
English with tags. Voice Leda, style "whispering nervously at first, then relieved and amused":
<laugh> Relax, it's just the cat. <sigh> You scared me half to death.Mandarin narration. Voice Charon, style "calm, warm Mandarin news anchor":
Mandarin with tags. Voice Kore, style "cold and mocking, dropping to a threatening whisper":
<laugh> 你以为我会就这么算了?合同在我手里,三天之内,你会求着来找我。<sigh> 可惜啊,我们本来可以是朋友。Japanese (Leda, "bright, cheerful anime girl") and Brazilian Portuguese (Puck, "friendly Brazilian Portuguese host"):
Two speakers in one request. Algenib as the old master, Fenrir as the Monkey King, with a style per line:
<laugh> 师父莫急,俺老孙不过是借了他三根毫毛!
<sigh> 借?只怕人家此刻正满山找你呢。
找便找,俺一个筋斗,早到了十万八千里外!How to direct Gemini 3.8 TTS
Style describes the whole turn. Short, concrete descriptions work best: who is speaking, how they feel, how fast. For example, "calm news anchor", "sarcastic and dry", "out of breath", or "warm bedtime storyteller". Google's 3.8 models treat the text itself strictly as a transcript, so put directions in the style, not in the script.
Inline tags mark a moment inside the text:
| Tag | What it does |
|---|---|
<laugh> | A laugh at that point |
<sigh> | A sigh |
<short pause> / <long pause> | A beat of silence |
<breath> | An audible breath |
<gasp> | A sharp intake of breath |
<cough> / <throat-clearing> | Exactly that |
Two speakers: name each speaker, give each one a voice, and give each line its own style. That's how the Monkey King scene above was made.
The 30 classic Gemini TTS voices
The classic prebuilt voices, with Google's one-word description of each. The voices API lists 2,089 prebuilt voices in all, across 30 locales, but none of them is Mandarin; for a native Mandarin Gemini voice, design one.
| Voice | Character | Voice | Character | Voice | Character |
|---|---|---|---|---|---|
| Zephyr | Bright | Puck | Upbeat | Charon | Informative |
| Kore | Firm | Fenrir | Excitable | Leda | Youthful |
| Orus | Firm | Aoede | Breezy | Callirrhoe | Easy-going |
| Autonoe | Bright | Enceladus | Breathy | Iapetus | Clear |
| Umbriel | Easy-going | Algieba | Smooth | Despina | Smooth |
| Erinome | Clear | Algenib | Gravelly | Rasalgethi | Informative |
| Laomedeia | Upbeat | Achernar | Soft | Alnilam | Firm |
| Schedar | Even | Gacrux | Mature | Pulcherrima | Forward |
| Achird | Friendly | Zubenelgenubi | Casual | Vindemiatrix | Gentle |
| Sadachbia | Lively | Sadaltager | Knowledgeable | Sulafat | Warm |
Every voice speaks every supported language; the language comes from the text.
Gemini 3.8 TTS vs Eleven v4: the same scripts
We ran the seven scripts above through both models: Gemini with one take each (its free tier allows only a few requests a day), and Eleven v4 with three takes each. Both outputs were scored with the same speech-to-text against the script.
| Gemini 3.8 Flash TTS | Eleven v4 | |
|---|---|---|
| Seconds to generate 1 s of audio (median) | 0.57 | 0.23 |
| No words added or dropped | 5 of 7 takes | 21 of 21 takes |
| Words added that weren't in the script | "I, I…", "Uh", "Oh" (English with tags); "嗯" (dialogue) | None |
| 眼镜 (glasses) heard as 眼睛 (eyes) | Yes | No |
| Tags performed, not read aloud | Yes | Yes |
| Same scripts, total audio length | 110 s (slower, more pauses) | 88 s |
| Cost per minute of audio | ~$0.017 (paid list price) | Billed per character: roughly $0.15–0.20 on ElevenLabs subscription plans (our estimate, at ~900 characters per minute of English) |
In short, Gemini is the cheaper engine by a wide margin, and directing it in plain words is easy. Eleven v4 is faster and sticks to the exact script. That matters for ads, subtitles, and anything a client signs off line by line. Our full-cast audio drama studio and theater run on Eleven v4 for that reason. More detail on v4 is in Eleven v4 vs v3.
Using the Gemini 3.8 TTS API
The 3.8 models use the Gemini API's interactions endpoint and return WAV:
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.8-flash-tts",
"input": [{"type": "user_input", "content": [{
"type": "text",
"text": "Wait... <short pause> did you hear that? <sigh>",
"annotations": [{"type": "speech_metadata", "style": "whispering nervously"}]
}]}],
"response_format": {"type": "audio"},
"generation_config": {"speech_config": [{"voice": "Kore"}]}
}'
The audio comes back base64-encoded in steps[].content[].data. Two things to know before you build on the free tier: Google may use free-tier inputs to improve its products, and apps for users in the EEA, the UK or Switzerland must use the paid tier.
FAQ
What is Gemini 3.8 TTS? Google's text-to-speech model released on September 23, 2026, with a cheaper sibling, Gemini 3.8 Flash-Lite TTS. You direct delivery in plain words and with inline tags such as <laugh> and <sigh>, and it can voice two speakers in one request.
Is Gemini 3.8 TTS free? Google offers a free tier with a small daily request limit; on it, Google may use your inputs to improve its products, and it may not be offered to users in the EEA, the UK or Switzerland. On this page you can generate two clips a day without an account while the site's daily free allowance lasts.
How much does Gemini 3.8 TTS cost? On the paid tier, $0.50 per million input text tokens and $9 per million output audio tokens through December 31, 2026, rising to $1 and $18 from January 1, 2027. We measured 32 audio tokens per second of speech, which is about $0.017 per minute of audio now and $0.035 from 2027.
How many voices does Gemini 3.8 TTS have? The 30 classic voices, such as Kore (firm), Puck (upbeat) and Charon (informative), are the ones most guides list. The voices API actually returned 2,089 prebuilt voices across 30 locales when we listed it on October 8, 2026, none of them Mandarin. You can also design a voice from a description or replicate one from a 10–30 second sample with a spoken consent phrase.
Does Gemini 3.8 TTS support Chinese? Yes, it detects the language automatically. In our test it read a Mandarin narration paragraph with one character off (眼镜 heard as 眼睛) and performed <laugh> and <sigh> in a Mandarin line.
Is Gemini 3.8 TTS better than ElevenLabs Eleven v4? It is far cheaper per minute. On our seven test scripts, Eleven v4 generated audio about 2.4 times as fast and never added or dropped a word; Gemini added filler words that were not in the script in two of seven takes.
What is the difference between Gemini 3.8 Flash TTS and Flash-Lite TTS? Flash-Lite is the cheaper, high-volume model: $6 instead of $9 per million audio tokens through 2026. Google positions Flash for fidelity and acting, and Flash-Lite for bulk production.
Is VoiceSmiths affiliated with Google? No. VoiceSmiths is an independent studio. This demo calls the public Gemini API; Gemini is a trademark of Google LLC.