Gemini 3.8 Voice Design: Free Online Demo

Try Gemini 3.8 voice design free, no sign-up: describe a voice, hear it improvise in character. What the safety filter blocks, the API, an ElevenLabs test.

Last updated: 2026-10-08

Gemini voice design turns a one-sentence description into a new voice. It is part of Gemini 3.8 TTS, released by Google on September 23, 2026: you describe age, accent, timbre and attitude, and Gemini creates a stored voice and answers with a short line it improvises in character. Try it below for free, with no account.

Design a voice with Gemini 3.8: free, no sign-up

2 free designs a day. Describe age, accent, timbre and attitude in a sentence or two. Gemini answers with a short line it improvises in character, in the language you write the description in.

Examples109 / 300
Gender
Locale
2 free left today

Runs on Google's free tier: Google may use what you type here to improve its products, so don't enter anything private.

Gemini voice design at a glance

Gemini voice design
Made byGoogle (Gemini 3.8 TTS)
ReleasedSeptember 23, 2026
Modelsgemini-3.8-flash-tts, gemini-3.8-flash-lite-tts
InputA description (Google suggests 1–2 sentences), gender, locale
OutputA stored voice (voice_…) plus a WAV preview improvised in character (7–49 s in our tests)
StorageMust be stored; 200 voices per project, each kept for one year
Speed12–40 seconds per design on the free tier, in our tests
CostNo separate price published; our 40-second preview was 1,002 audio tokens, about $0.009 at the 2026 paid rate
SafetySome descriptions are refused; all Gemini audio carries a SynthID watermark
WhereGemini API and Google AI Studio

What we found testing it

We designed voices through the API on October 8, 2026, on Google's free tier, and kept every take.

1. The preview is improvised, not read. Asked for "a warm, gravel-voiced Scottish fisherman in his sixties who tells stories slowly, with a dry sense of humour", Gemini wrote and performed its own 40-second monologue about a stubborn boat:

Voice: Gemini designed voice · Scottish fishermanEnglishModel: gemini-3.8-flash-tts40 s
Transcript:Ah, it's a, it's a stubborn lass, that one. Aye. I've been at it now, what, an hour? And it, so right she won't. It's not what's putting against her, she says. Ah, it's the, the things in her head. I can tell you when the wind picks up and, and the waves come. But this one, no, she's, she's just had enough for the day. Yeah, I suppose that's quite right, actually. After all this time, well, who am I to question it? Little so.

Then the same voice read our line, which is what you would use in production:

Voice: Gemini designed voice · Scottish fishermanEnglishModel: gemini-3.8-flash-tts6 s
Transcript:Right then. The tide turns at half past four, and if you're not on the boat by then, you're not coming.

2. The preview speaks the language of the description, not the locale. We gave the Monkey King description in English with the locale set to Mandarin (cmn-CN). The preview came back in English. A description written in Chinese came back in Chinese, complete with a Beijing storyteller's pauses:

Voice: Gemini 设计音色 · 北京说书先生ChineseModel: gemini-3.8-flash-tts21 s
Transcript:您说的没错。就那么一回呀,老吴提着灯笼,正打后院那片荒草地边上过,突然间,哎,就听见那么一声响,嘿,轻飘飘的,可愣是叫人听得头皮发麻呀。

The locale still matters when the voice reads your script: the English-described Monkey King read Chinese text cleanly (sample below).

3. The safety filter refuses some fiction. Our Ice Queen description, "An icy, menacing villainess queen: low, slow, every word a blade", came back as Voice prompt was blocked by safety policies. ElevenLabs Voice Design had accepted the same words. If you write villains, describe the sound (low, slow, cold, precise), not the menace.

4. There is no Mandarin in Google's prebuilt library. Listing the voices API returned 2,089 prebuilt voices across 30 locales. The biggest are US English (215), Indian English (120), British English (118), Korean (117), Japanese (115), Hindi (114) and Brazilian Portuguese (107). None is Mandarin or Cantonese, so for a native Mandarin Gemini voice, design is the only route.

5. Japanese design failed on the day. Two attempts to design our Anime Girl with the locale ja-JP returned 503 The service is currently unavailable, while English designs a minute before and after worked. That may be temporary, but test your locale before you build on it.

Same description, two engines

All 24 character voices on VoiceSmiths were made with ElevenLabs Voice Design. We fed Gemini the description from each voice's page, then had the Gemini voice read the same script that voice's page already plays.

Monkey King · "A Monkey King–style character: high-pitched, cheeky, slightly raspy with opera flair."

Gemini 3.8 (designed with locale cmn-CN):

Voice: Gemini designed voice · Monkey KingChineseModel: gemini-3.8-flash-tts10 s
Transcript:新一代降噪耳机,四十小时续航,三秒开盖即连。今天下单,立减两百,再送一年延保。数量有限,现在就来。

ElevenLabs, the Monkey King voice:

Voice: 齐天大圣·猴王 (ElevenLabs Voice Design)Chinese17 s
Transcript:新一代降噪耳机,四十小时续航,三秒开盖即连。今天下单,立减两百,再送一年延保。数量有限,现在就来。

And the preview Gemini improvised for it, in English, laughing at the gods:

Voice: Gemini designed voice · Monkey KingEnglishModel: gemini-3.8-flash-tts30 s
Transcript:[laughs] See? They thought they had a trick up their sleeve, eh? Did they? To trap me, the mighty master of mischief, in something so petty. [laughs] Nonsense. I'm faster than a thousand lightning bolts, and all that big mouth talk about sealing me away, [laughs] all the gods, I simply shrug. Look like a hundred sleepy owls lost in a fog. [laughs]

Gemini read the 49-character ad in 10.4 seconds against 17.3 for ElevenLabs: brighter and faster, with less of the opera rasp.

Trailer Voice · "The classic Hollywood "In a world…" trailer voice: ultra-deep, gravelly, epic."

Gemini 3.8 (designed with locale en-US):

Voice: Gemini designed voice · Trailer VoiceEnglishModel: gemini-3.8-flash-tts19 s
Transcript:The new noise-cancelling headphones: forty hours of battery, connected three seconds after you open the case. Order today, save two hundred, and get a year of extra warranty. Limited stock — go now.

ElevenLabs, the Trailer Voice:

Voice: Trailer Voice (ElevenLabs Voice Design)English23 s
Transcript:The new noise-cancelling headphones: forty hours of battery, connected three seconds after you open the case. Order today, save two hundred, and get a year of extra warranty. Limited stock — go now.

The preview Gemini wrote for itself, an epic opening about the course of human history:

Voice: Gemini designed voice · Trailer VoiceEnglishModel: gemini-3.8-flash-tts32 s
Transcript:In the vast tapestry of time, where echoes of ancient whispers persist among the sands, one thing remains clear: this is the course of human history. The world we built, the boundaries we pushed, and now we stand at a crossroads.

British Butler · "A refined British butler in crisp RP: calm, dry, impeccably courteous."

Gemini 3.8 (designed with locale en-GB):

Voice: Gemini designed voice · British ButlerEnglishModel: gemini-3.8-flash-tts16 s
Transcript:On the third morning the fog had still not lifted. She walked back along the embankment with dew on her shoes, and when a horn sounded from the distant ferry she remembered the letter, still in the desk drawer, unsent.

ElevenLabs, the British Butler:

Voice: British Butler (ElevenLabs Voice Design)English15 s
Transcript:On the third morning the fog had still not lifted. She walked back along the embankment with dew on her shoes, and when a horn sounded from the distant ferry she remembered the letter, still in the desk drawer, unsent.

Gemini's preview this time was a single 7-second line, setting aside a fine Bordeaux:

Voice: Gemini designed voice · British ButlerEnglishModel: gemini-3.8-flash-tts8 s
Transcript:I've taken the liberty of setting aside some very fine Bordeaux for your gathering, sir.
VoiceScriptGemini 3.8ElevenLabs
Monkey KingChinese ad, 49 characters10.4 s17.3 s
Trailer VoiceEnglish ad, 198 characters19.3 s22.9 s
British ButlerEnglish narration, 218 characters16.1 s15.0 s
Ice Queen—Refused by the safety filterDesigned
Anime Girl (ja-JP)—"Service unavailable" (503), twiceDesigned

Gemini read both ads faster than ElevenLabs and the narration at about the same pace. One caveat: each ElevenLabs voice is the one we picked for that character when we built the site, while each Gemini voice here is one description and one take. Read this as what each engine gives you first time, not as a ceiling.

How to write a description that works

  • Fix the permanent traits: age, gender, timbre, accent, pace. Google's own example is "a crisp, energetic sports announcer in her 30s with a slight Midwestern accent".
  • Keep moods for later. Situational emotion ("whispering", "furious") belongs in the style direction of each Gemini TTS request, not in the voice.
  • Write it in the language you want to hear in the preview.
  • Describe sound, not threat. "Cold, slow, precise" passes; "menacing, every word a blade" did not.

Using voice design through the API

Create the voice. The response carries id and sample_audio (base64 WAV):

curl -X POST "https://generativelanguage.googleapis.com/v1beta/voices" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "store": true,
    "voice": {
      "model": "gemini-3.8-flash-tts",
      "type": "prompted",
      "display_name": "Scottish fisherman",
      "gender": "male",
      "language_code": "en-GB",
      "prompted": { "input": "A warm, gravel-voiced Scottish fisherman in his sixties who tells stories slowly." }
    }
  }'

Read a script with it by passing the ID as the voice:

curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3.8-flash-tts",
    "input": [{ "type": "user_input", "content": [{ "type": "text", "text": "Right then. The tide turns at half past four." }] }],
    "response_format": { "type": "audio" },
    "generation_config": { "speech_config": [{ "voice": "voice_YOUR_ID" }] }
  }'

Two things to know: "store": false is rejected for designed voices, and voices count against the 200-per-project limit until you delete them with DELETE /v1beta/voices/{id}. The demo on this page deletes each voice as soon as its preview is back.

Gemini vs ElevenLabs voice design

Gemini 3.8 voice designElevenLabs Voice Design
PreviewOne, improvised in characterSeveral, reading a text you give (or it writes)
Where you use the voiceGemini API, AI StudioElevenLabs apps and API, and studios built on it such as VoiceSmiths
MandarinYes, via design (cmn-CN)Yes
Villain-style descriptionsSome refusedAccepted in our test
Keeping voices200 per project, one yearSaved to your voice library

On VoiceSmiths you can design a voice with ElevenLabs in the voice studio and keep it for every tool, or start from the 24 character voices already designed in the voice library. For Gemini's prebuilt voices and its style direction, see Gemini 3.8 TTS online; for copying a real voice, see Gemini voice cloning.

FAQ

What is Gemini voice design? Voice design is a Gemini 3.8 TTS feature, released by Google on September 23, 2026, that creates a new voice from a one- or two-sentence description of age, gender, accent, timbre and attitude. The voice is stored in your Google project and can then read any script through the Gemini API.

Can I try Gemini voice design for free? Yes. On this page you can design two voices a day without an account. Google AI Studio also lets developers try it. Through the API, Google publishes no separate price for designing a voice; in our test a design used about 1,000 audio tokens, under one US cent at the 2026 paid rate.

What does Gemini return when you design a voice? A voice ID (voice_…) and a WAV preview. The preview is not a fixed sentence: Gemini improvises a short line in character (7 to 49 seconds in our tests), in the language the description is written in. To hear your own script, send it to the interactions endpoint with that voice ID.

Can Gemini design a Mandarin Chinese voice? Yes, with the locale cmn-CN. That matters because Google's prebuilt library, 2,089 voices across 30 locales when we listed it on October 8, 2026, has no Mandarin voice at all. Write the description in Chinese if you want the preview in Chinese.

Why was my voice description blocked? Google's safety filter refuses some descriptions with "Voice prompt was blocked by safety policies." In our test it refused a villain description ("an icy, menacing villainess queen… every word a blade") that ElevenLabs Voice Design accepted. Describe how the voice sounds, such as low, slow, cold and precise, instead of what the character threatens.

How many designed voices can I keep? Designed voices must be stored, and a Google project holds at most 200 stored voices (designed and cloned together). Each expires one year after it is created, and you can delete one at any time.

Gemini voice design or ElevenLabs Voice Design? Gemini gives one improvised preview per request and only works through Google's API or AI Studio. ElevenLabs returns several previews of a text you choose and saves voices to a library you can use in its apps. All 24 character voices on VoiceSmiths were designed with ElevenLabs; the comparison on this page runs the same descriptions through Gemini.

Is VoiceSmiths affiliated with Google? No. VoiceSmiths is an independent studio. Gemini is a trademark of Google LLC; the demo on this page calls the public Gemini API.

Browse the 24 designed character voices →