Gemini 3.8 Voice Design: Free Online Demo
Try Gemini 3.8 voice design free, no sign-up: describe a voice, hear it improvise in character. What the safety filter blocks, the API, an ElevenLabs test.
Last updated: 2026-10-08
Gemini voice design turns a one-sentence description into a new voice. It is part of Gemini 3.8 TTS, released by Google on September 23, 2026: you describe age, accent, timbre and attitude, and Gemini creates a stored voice and answers with a short line it improvises in character. Try it below for free, with no account.
Design a voice with Gemini 3.8: free, no sign-up
2 free designs a day. Describe age, accent, timbre and attitude in a sentence or two. Gemini answers with a short line it improvises in character, in the language you write the description in.
Runs on Google's free tier: Google may use what you type here to improve its products, so don't enter anything private.
Gemini voice design at a glance
| Gemini voice design | |
|---|---|
| Made by | Google (Gemini 3.8 TTS) |
| Released | September 23, 2026 |
| Models | gemini-3.8-flash-tts, gemini-3.8-flash-lite-tts |
| Input | A description (Google suggests 1–2 sentences), gender, locale |
| Output | A stored voice (voice_…) plus a WAV preview improvised in character (7–49 s in our tests) |
| Storage | Must be stored; 200 voices per project, each kept for one year |
| Speed | 12–40 seconds per design on the free tier, in our tests |
| Cost | No separate price published; our 40-second preview was 1,002 audio tokens, about $0.009 at the 2026 paid rate |
| Safety | Some descriptions are refused; all Gemini audio carries a SynthID watermark |
| Where | Gemini API and Google AI Studio |
What we found testing it
We designed voices through the API on October 8, 2026, on Google's free tier, and kept every take.
1. The preview is improvised, not read. Asked for "a warm, gravel-voiced Scottish fisherman in his sixties who tells stories slowly, with a dry sense of humour", Gemini wrote and performed its own 40-second monologue about a stubborn boat:
Then the same voice read our line, which is what you would use in production:
2. The preview speaks the language of the description, not the locale. We gave the Monkey King description in English with the locale set to Mandarin (cmn-CN). The preview came back in English. A description written in Chinese came back in Chinese, complete with a Beijing storyteller's pauses:
The locale still matters when the voice reads your script: the English-described Monkey King read Chinese text cleanly (sample below).
3. The safety filter refuses some fiction. Our Ice Queen description, "An icy, menacing villainess queen: low, slow, every word a blade", came back as Voice prompt was blocked by safety policies. ElevenLabs Voice Design had accepted the same words. If you write villains, describe the sound (low, slow, cold, precise), not the menace.
4. There is no Mandarin in Google's prebuilt library. Listing the voices API returned 2,089 prebuilt voices across 30 locales. The biggest are US English (215), Indian English (120), British English (118), Korean (117), Japanese (115), Hindi (114) and Brazilian Portuguese (107). None is Mandarin or Cantonese, so for a native Mandarin Gemini voice, design is the only route.
5. Japanese design failed on the day. Two attempts to design our Anime Girl with the locale ja-JP returned 503 The service is currently unavailable, while English designs a minute before and after worked. That may be temporary, but test your locale before you build on it.
Same description, two engines
All 24 character voices on VoiceSmiths were made with ElevenLabs Voice Design. We fed Gemini the description from each voice's page, then had the Gemini voice read the same script that voice's page already plays.
Monkey King · "A Monkey King–style character: high-pitched, cheeky, slightly raspy with opera flair."
Gemini 3.8 (designed with locale cmn-CN):
ElevenLabs, the Monkey King voice:
And the preview Gemini improvised for it, in English, laughing at the gods:
[laughs] See? They thought they had a trick up their sleeve, eh? Did they? To trap me, the mighty master of mischief, in something so petty. [laughs] Nonsense. I'm faster than a thousand lightning bolts, and all that big mouth talk about sealing me away, [laughs] all the gods, I simply shrug. Look like a hundred sleepy owls lost in a fog. [laughs]Gemini read the 49-character ad in 10.4 seconds against 17.3 for ElevenLabs: brighter and faster, with less of the opera rasp.
Trailer Voice · "The classic Hollywood "In a world…" trailer voice: ultra-deep, gravelly, epic."
Gemini 3.8 (designed with locale en-US):
ElevenLabs, the Trailer Voice:
The preview Gemini wrote for itself, an epic opening about the course of human history:
British Butler · "A refined British butler in crisp RP: calm, dry, impeccably courteous."
Gemini 3.8 (designed with locale en-GB):
ElevenLabs, the British Butler:
Gemini's preview this time was a single 7-second line, setting aside a fine Bordeaux:
| Voice | Script | Gemini 3.8 | ElevenLabs |
|---|---|---|---|
| Monkey King | Chinese ad, 49 characters | 10.4 s | 17.3 s |
| Trailer Voice | English ad, 198 characters | 19.3 s | 22.9 s |
| British Butler | English narration, 218 characters | 16.1 s | 15.0 s |
| Ice Queen | — | Refused by the safety filter | Designed |
Anime Girl (ja-JP) | — | "Service unavailable" (503), twice | Designed |
Gemini read both ads faster than ElevenLabs and the narration at about the same pace. One caveat: each ElevenLabs voice is the one we picked for that character when we built the site, while each Gemini voice here is one description and one take. Read this as what each engine gives you first time, not as a ceiling.
How to write a description that works
- Fix the permanent traits: age, gender, timbre, accent, pace. Google's own example is "a crisp, energetic sports announcer in her 30s with a slight Midwestern accent".
- Keep moods for later. Situational emotion ("whispering", "furious") belongs in the style direction of each Gemini TTS request, not in the voice.
- Write it in the language you want to hear in the preview.
- Describe sound, not threat. "Cold, slow, precise" passes; "menacing, every word a blade" did not.
Using voice design through the API
Create the voice. The response carries id and sample_audio (base64 WAV):
curl -X POST "https://generativelanguage.googleapis.com/v1beta/voices" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"store": true,
"voice": {
"model": "gemini-3.8-flash-tts",
"type": "prompted",
"display_name": "Scottish fisherman",
"gender": "male",
"language_code": "en-GB",
"prompted": { "input": "A warm, gravel-voiced Scottish fisherman in his sixties who tells stories slowly." }
}
}'
Read a script with it by passing the ID as the voice:
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.8-flash-tts",
"input": [{ "type": "user_input", "content": [{ "type": "text", "text": "Right then. The tide turns at half past four." }] }],
"response_format": { "type": "audio" },
"generation_config": { "speech_config": [{ "voice": "voice_YOUR_ID" }] }
}'
Two things to know: "store": false is rejected for designed voices, and voices count against the 200-per-project limit until you delete them with DELETE /v1beta/voices/{id}. The demo on this page deletes each voice as soon as its preview is back.
Gemini vs ElevenLabs voice design
| Gemini 3.8 voice design | ElevenLabs Voice Design | |
|---|---|---|
| Preview | One, improvised in character | Several, reading a text you give (or it writes) |
| Where you use the voice | Gemini API, AI Studio | ElevenLabs apps and API, and studios built on it such as VoiceSmiths |
| Mandarin | Yes, via design (cmn-CN) | Yes |
| Villain-style descriptions | Some refused | Accepted in our test |
| Keeping voices | 200 per project, one year | Saved to your voice library |
On VoiceSmiths you can design a voice with ElevenLabs in the voice studio and keep it for every tool, or start from the 24 character voices already designed in the voice library. For Gemini's prebuilt voices and its style direction, see Gemini 3.8 TTS online; for copying a real voice, see Gemini voice cloning.
FAQ
What is Gemini voice design? Voice design is a Gemini 3.8 TTS feature, released by Google on September 23, 2026, that creates a new voice from a one- or two-sentence description of age, gender, accent, timbre and attitude. The voice is stored in your Google project and can then read any script through the Gemini API.
Can I try Gemini voice design for free? Yes. On this page you can design two voices a day without an account. Google AI Studio also lets developers try it. Through the API, Google publishes no separate price for designing a voice; in our test a design used about 1,000 audio tokens, under one US cent at the 2026 paid rate.
What does Gemini return when you design a voice? A voice ID (voice_…) and a WAV preview. The preview is not a fixed sentence: Gemini improvises a short line in character (7 to 49 seconds in our tests), in the language the description is written in. To hear your own script, send it to the interactions endpoint with that voice ID.
Can Gemini design a Mandarin Chinese voice? Yes, with the locale cmn-CN. That matters because Google's prebuilt library, 2,089 voices across 30 locales when we listed it on October 8, 2026, has no Mandarin voice at all. Write the description in Chinese if you want the preview in Chinese.
Why was my voice description blocked? Google's safety filter refuses some descriptions with "Voice prompt was blocked by safety policies." In our test it refused a villain description ("an icy, menacing villainess queen… every word a blade") that ElevenLabs Voice Design accepted. Describe how the voice sounds, such as low, slow, cold and precise, instead of what the character threatens.
How many designed voices can I keep? Designed voices must be stored, and a Google project holds at most 200 stored voices (designed and cloned together). Each expires one year after it is created, and you can delete one at any time.
Gemini voice design or ElevenLabs Voice Design? Gemini gives one improvised preview per request and only works through Google's API or AI Studio. ElevenLabs returns several previews of a text you choose and saves voices to a library you can use in its apps. All 24 character voices on VoiceSmiths were designed with ElevenLabs; the comparison on this page runs the same descriptions through Gemini.
Is VoiceSmiths affiliated with Google? No. VoiceSmiths is an independent studio. Gemini is a trademark of Google LLC; the demo on this page calls the public Gemini API.