Gemini 3.8 TTS Online - Free Gemini TTS Demo

Try Gemini 3.8 Flash TTS free, no sign-up: pick a voice, describe the delivery, hear it. Real samples, all 30 voices, cost per minute, an Eleven v4 test.

Last updated: 2026-10-08

Gemini 3.8 Flash TTS is Google's newest text-to-speech model, released on September 23, 2026. Instead of picking from emotion presets, you tell it how to sound in plain words, such as "whispering nervously" or "excited sports commentator", and drop sounds into the text with tags like <laugh> and <sigh>. It handles 100+ languages and two speakers per request, and costs about $0.017 per minute of audio at Google's paid rate. Try it below:

Try Gemini 3.8 TTS now: free, no sign-up

2 free generations a day, up to 200 characters. Describe the delivery in plain words; tags like <laugh> are performed, not read.

Voice
Tags100 / 200
Examples
2 free left today

Runs on Google's free tier: Google may use what you type here to improve its products, so don't enter anything private.

Gemini 3.8 TTS at a glance

Gemini 3.8 Flash TTS
Made byGoogle
ReleasedSeptember 23, 2026, alongside Gemini 3.8 Flash-Lite TTS
Model IDsgemini-3.8-flash-tts, gemini-3.8-flash-lite-tts
Voices30 classic prebuilt voices; the voices API lists 2,089 across 30 locales (no Mandarin); plus voice design from a description and voice cloning from a 10–30 s sample
Languages100+ languages and dialects, detected automatically (the API docs say 130+ for Flash)
DirectingA plain-language style per turn, plus inline tags such as <laugh>, <sigh> and <short pause>
SpeakersUp to 2 per request
OutputWAV, 24 kHz mono 16-bit
PriceFree tier with a small daily limit; paid: $0.50 per 1M text tokens in, $9 per 1M audio tokens out (doubles on January 1, 2027)
Cost per minute (measured)About $0.017 now, $0.035 from 2027, at 32 audio tokens per second
WatermarkSynthID on every clip
WhereGemini API, Google AI Studio, Gemini Notebook; Flash-Lite in Google Vids

Listen: real Gemini 3.8 TTS output

Every clip was generated with gemini-3.8-flash-tts on October 5, 2026: one take each, no editing. The transcript is the exact text sent; the style we asked for is written above each player. These are the same scripts we ran through Eleven v4, so you can compare them directly.

English narration. Voice Charon, style "warm, unhurried British storyteller":

Voice: CharonEnglishModel: gemini-3.8-flash-tts18 s
Transcript:The house had been empty for years, yet every clock in it still told the right time. Margaret noticed it the moment she stepped inside: the soft, patient ticking, as if someone had been winding them every night, waiting for her to come home.

English with tags. Voice Leda, style "whispering nervously at first, then relieved and amused":

Voice: LedaEnglishModel: gemini-3.8-flash-tts15 s
Transcript:Okay, don't move. I think it's still in the kitchen. What if it heard us? <laugh> Relax, it's just the cat. <sigh> You scared me half to death.

Mandarin narration. Voice Charon, style "calm, warm Mandarin news anchor":

Voice: CharonChineseModel: gemini-3.8-flash-tts14 s
Transcript:天还没亮,菜市场已经亮起了灯。卖豆腐的老陈把第一板豆腐抬上案子,热气一下子冒起来,模糊了他的眼镜。这么多年,他每天都是第一个到。

Mandarin with tags. Voice Kore, style "cold and mocking, dropping to a threatening whisper":

Voice: KoreChineseModel: gemini-3.8-flash-tts20 s
Transcript:<laugh> 你以为我会就这么算了?合同在我手里,三天之内,你会求着来找我。<sigh> 可惜啊,我们本来可以是朋友。

Japanese (Leda, "bright, cheerful anime girl") and Brazilian Portuguese (Puck, "friendly Brazilian Portuguese host"):

Voice: LedaJapaneseModel: gemini-3.8-flash-tts10 s
Transcript:おはようございます!今日は朝から雨ですが、新しい傘を買ったので、ちょっとだけ楽しみです。駅まで一緒に歩きませんか?
Voice: PuckPortugueseModel: gemini-3.8-flash-tts9 s
Transcript:Bom dia! Hoje vamos aprender a fazer um café coado perfeito, sem pressa e sem segredo. Primeiro, aqueça a água sem deixar ferver.

Two speakers in one request. Algenib as the old master, Fenrir as the Monkey King, with a style per line:

Voice: Algenib / FenrirChineseModel: gemini-3.8-flash-tts24 s
Transcript:你这猴头,又从哪里闯了祸回来? <laugh> 师父莫急,俺老孙不过是借了他三根毫毛! <sigh> 借?只怕人家此刻正满山找你呢。 找便找,俺一个筋斗,早到了十万八千里外!

How to direct Gemini 3.8 TTS

Style describes the whole turn. Short, concrete descriptions work best: who is speaking, how they feel, how fast. For example, "calm news anchor", "sarcastic and dry", "out of breath", or "warm bedtime storyteller". Google's 3.8 models treat the text itself strictly as a transcript, so put directions in the style, not in the script.

Inline tags mark a moment inside the text:

TagWhat it does
<laugh>A laugh at that point
<sigh>A sigh
<short pause> / <long pause>A beat of silence
<breath>An audible breath
<gasp>A sharp intake of breath
<cough> / <throat-clearing>Exactly that

Two speakers: name each speaker, give each one a voice, and give each line its own style. That's how the Monkey King scene above was made.

The 30 classic Gemini TTS voices

The classic prebuilt voices, with Google's one-word description of each. The voices API lists 2,089 prebuilt voices in all, across 30 locales, but none of them is Mandarin; for a native Mandarin Gemini voice, design one.

VoiceCharacterVoiceCharacterVoiceCharacter
ZephyrBrightPuckUpbeatCharonInformative
KoreFirmFenrirExcitableLedaYouthful
OrusFirmAoedeBreezyCallirrhoeEasy-going
AutonoeBrightEnceladusBreathyIapetusClear
UmbrielEasy-goingAlgiebaSmoothDespinaSmooth
ErinomeClearAlgenibGravellyRasalgethiInformative
LaomedeiaUpbeatAchernarSoftAlnilamFirm
SchedarEvenGacruxMaturePulcherrimaForward
AchirdFriendlyZubenelgenubiCasualVindemiatrixGentle
SadachbiaLivelySadaltagerKnowledgeableSulafatWarm

Every voice speaks every supported language; the language comes from the text.

Gemini 3.8 TTS vs Eleven v4: the same scripts

We ran the seven scripts above through both models: Gemini with one take each (its free tier allows only a few requests a day), and Eleven v4 with three takes each. Both outputs were scored with the same speech-to-text against the script.

Gemini 3.8 Flash TTSEleven v4
Seconds to generate 1 s of audio (median)0.570.23
No words added or dropped5 of 7 takes21 of 21 takes
Words added that weren't in the script"I, I…", "Uh", "Oh" (English with tags); "嗯" (dialogue)None
眼镜 (glasses) heard as 眼睛 (eyes)YesNo
Tags performed, not read aloudYesYes
Same scripts, total audio length110 s (slower, more pauses)88 s
Cost per minute of audio~$0.017 (paid list price)Billed per character: roughly $0.15–0.20 on ElevenLabs subscription plans (our estimate, at ~900 characters per minute of English)

In short, Gemini is the cheaper engine by a wide margin, and directing it in plain words is easy. Eleven v4 is faster and sticks to the exact script. That matters for ads, subtitles, and anything a client signs off line by line. Our full-cast audio drama studio and theater run on Eleven v4 for that reason. More detail on v4 is in Eleven v4 vs v3.

Using the Gemini 3.8 TTS API

The 3.8 models use the Gemini API's interactions endpoint and return WAV:

curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3.8-flash-tts",
    "input": [{"type": "user_input", "content": [{
      "type": "text",
      "text": "Wait... <short pause> did you hear that? <sigh>",
      "annotations": [{"type": "speech_metadata", "style": "whispering nervously"}]
    }]}],
    "response_format": {"type": "audio"},
    "generation_config": {"speech_config": [{"voice": "Kore"}]}
  }'

The audio comes back base64-encoded in steps[].content[].data. Two things to know before you build on the free tier: Google may use free-tier inputs to improve its products, and apps for users in the EEA, the UK or Switzerland must use the paid tier.

FAQ

What is Gemini 3.8 TTS? Google's text-to-speech model released on September 23, 2026, with a cheaper sibling, Gemini 3.8 Flash-Lite TTS. You direct delivery in plain words and with inline tags such as <laugh> and <sigh>, and it can voice two speakers in one request.

Is Gemini 3.8 TTS free? Google offers a free tier with a small daily request limit; on it, Google may use your inputs to improve its products, and it may not be offered to users in the EEA, the UK or Switzerland. On this page you can generate two clips a day without an account while the site's daily free allowance lasts.

How much does Gemini 3.8 TTS cost? On the paid tier, $0.50 per million input text tokens and $9 per million output audio tokens through December 31, 2026, rising to $1 and $18 from January 1, 2027. We measured 32 audio tokens per second of speech, which is about $0.017 per minute of audio now and $0.035 from 2027.

How many voices does Gemini 3.8 TTS have? The 30 classic voices, such as Kore (firm), Puck (upbeat) and Charon (informative), are the ones most guides list. The voices API actually returned 2,089 prebuilt voices across 30 locales when we listed it on October 8, 2026, none of them Mandarin. You can also design a voice from a description or replicate one from a 10–30 second sample with a spoken consent phrase.

Does Gemini 3.8 TTS support Chinese? Yes, it detects the language automatically. In our test it read a Mandarin narration paragraph with one character off (眼镜 heard as 眼睛) and performed <laugh> and <sigh> in a Mandarin line.

Is Gemini 3.8 TTS better than ElevenLabs Eleven v4? It is far cheaper per minute. On our seven test scripts, Eleven v4 generated audio about 2.4 times as fast and never added or dropped a word; Gemini added filler words that were not in the script in two of seven takes.

What is the difference between Gemini 3.8 Flash TTS and Flash-Lite TTS? Flash-Lite is the cheaper, high-volume model: $6 instead of $9 per million audio tokens through 2026. Google positions Flash for fidelity and acting, and Flash-Lite for bulk production.

Is VoiceSmiths affiliated with Google? No. VoiceSmiths is an independent studio. This demo calls the public Gemini API; Gemini is a trademark of Google LLC.