Gemini Voice Cloning: How Replication Works

Gemini 3.8 voice cloning: a 10–30 s sample plus a spoken consent phrase (verbatim), where AI Studio blocks it, limits, the API call, vs ElevenLabs cloning.

Last updated: 2026-10-08

Gemini can clone a voice from a 10–30 second recording, as long as the speaker also records a consent phrase. Google calls it voice replication. It shipped with Gemini 3.8 TTS on September 23, 2026, runs in Google AI Studio and the Gemini API, and, according to Google's announcement, is not available in AI Studio in Illinois, Texas, the EEA, the UK, Switzerland or India. This page covers what it needs, the exact consent phrases, the API call and how it compares with ElevenLabs.

Gemini voice cloning at a glance

Gemini voice replication
Made byGoogle (Gemini 3.8 TTS)
ReleasedSeptember 23, 2026
Modelsgemini-3.8-flash-tts, gemini-3.8-flash-lite-tts
Reference sample10–30 seconds of one adult speaker; 24 kHz mono 16-bit WAV recommended
ConsentThe same speaker reads Google's exact phrase (30 languages)
ResultStored voice voice_… (1 year, 200 per project) or stateless voicekey_… (7 days)
WhereGoogle AI Studio and the Gemini API
Not in AI StudioIllinois, Texas, EEA, UK, Switzerland, India (Google's announcement)
WatermarkSynthID on all Gemini audio
PriceNo separate creation price; speech billed as Gemini TTS output ($9 per 1M audio tokens on Flash through 2026)

What Gemini needs from you

  1. A reference sample, 10–30 seconds. One adult speaker, talking naturally, with no music or other voices. A quiet room matters more than an expensive microphone.
  2. A consent recording. The same speaker reads Google's phrase word for word. Record it with the same microphone in the same room as the sample: Google checks that both recordings are the same person, and a change of setup is the easiest way to fail that check.
  3. A Google AI Studio account or a Gemini API key.

The consent phrases

Google publishes the phrase in 30 languages. Three of them, verbatim:

LanguagePhrase
English (US), en-USI am the owner of this voice and I consent to Google using this voice to create a synthetic voice model.
Chinese (Simplified), zh-CN我是此声音的拥有者并授权谷歌使用此声音创建语音合成模型
Japanese, ja-JP私はこの音声の所有者であり、Googleがこの音声を使用して音声合成モデルを作成することを承認します。

The full list (Arabic, Hindi, Korean, Portuguese, Spanish and more) is in Google's voice replication guide.

Cloning a voice through the API

Build the request from your two WAV files (base64 -i is macOS; on Linux use base64 -w0), then create the voice:

jq -n \
  --arg src "$(base64 -i reference.wav)" \
  --arg consent "$(base64 -i consent.wav)" \
  '{store: true, voice: {model: "gemini-3.8-flash-tts", type: "replicated", display_name: "My voice",
    replicated: {source_audio: {mime_type: "audio/wav", data: $src},
                 consent_audio: {mime_type: "audio/wav", data: $consent}}}}' > voice.json

curl -X POST "https://generativelanguage.googleapis.com/v1beta/voices" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H "Content-Type: application/json" \
  -d @voice.json

The response carries id (voice_…). Use it like any other Gemini voice:

curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3.8-flash-tts",
    "input": [{ "type": "user_input", "content": [{ "type": "text", "text": "This is my cloned voice." }] }],
    "response_format": { "type": "audio" },
    "generation_config": { "speech_config": [{ "voice": "voice_YOUR_ID" }] }
  }'

With "store": false you get a voicekey_… instead, which Google does not keep for you and which works for 7 days. Stored voices count toward the 200-per-project limit (shared with designed voices) until you delete them with DELETE /v1beta/voices/{id}.

Gemini vs ElevenLabs voice cloning

Gemini voice replicationElevenLabs Instant Voice CloningElevenLabs Professional Voice Cloning
Audio needed10–30 secondsShort samples, under two minutesMuch longer training audio
Speaker checkSpoken consent phrase, matched to the sampleNone described in ElevenLabs' docsVoice captcha
PlanGemini API or AI StudioMost plansCreator plan or above
Where the voice livesYour Google project (1 year)Your ElevenLabs voice libraryYour ElevenLabs voice library, shareable publicly
Typical useDevelopers building on GeminiFast clones for production workHighest-fidelity clone of your own voice

Sources: Google's Gemini API voice replication guide and announcement, and ElevenLabs' voice documentation, checked October 8, 2026.

Why there are no clone samples on this page

Google's consent check exists so that nobody's voice is copied without them. We hold ourselves to the same rule: we only publish a clone when its owner has agreed to it, and we have not yet recorded a volunteer for a side-by-side test. Every audio sample on our Gemini voice design page is from a voice that belongs to no one.

Cloning on VoiceSmiths

The VoiceSmiths voice studio clones voices with ElevenLabs instant cloning: upload 1–3 clean minutes, name the voice, and it works in text to speech, multi-speaker dialogue and long-form audiobooks, in 90+ languages on Eleven v4. A clone costs 30 credits; plans start at $3.99. The rule is the same as Google's: only your own voice, or one you have written permission to use.

No voice to clone? Design one with Gemini for free, or pick from 24 original character voices.

FAQ

Can Gemini clone a voice? Yes. Gemini 3.8 TTS, released by Google on September 23, 2026, can replicate a voice from a 10–30 second reference recording. Google calls the feature voice replication. It works in Google AI Studio and through the Gemini API with the gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts models.

What does Gemini need to clone a voice? Two recordings of the same adult speaker: a 10–30 second reference sample (Google recommends 24 kHz mono 16-bit WAV) and a consent recording in which the speaker reads Google's exact consent phrase. Google advises recording both with the same microphone in the same room so the speaker check passes.

What is the Gemini voice cloning consent phrase? In US English: "I am the owner of this voice and I consent to Google using this voice to create a synthetic voice model." In Simplified Chinese: "我是此声音的拥有者并授权谷歌使用此声音创建语音合成模型". Google lists verbatim phrases for 30 languages.

Where is Gemini voice cloning not available? Google's announcement says voice replication in AI Studio is not available in Illinois, Texas, the EEA, the UK, Switzerland and India. The Gemini API documentation does not list a regional limit as of October 8, 2026.

Can I clone someone else's voice with Gemini? Only with that person taking part: the consent recording has to come from the same speaker as the reference sample, reading Google's phrase. A clip of a celebrity or anyone who has not recorded the consent phrase will not pass.

How long does a Gemini cloned voice last? A stored voice (store: true) gets an ID like voice_… and lasts one year; a project can hold 200 stored voices, designed and cloned together. A stateless voice (store: false) returns a voicekey_… that you keep yourself, valid for 7 days.

Is Gemini voice cloning free? Google publishes no separate price for creating a voice. Speech made with it is billed like other Gemini 3.8 TTS output: $9 per million audio tokens on Flash TTS through 2026 ($18 from January 1, 2027), with a free tier for developers.

Gemini voice cloning or ElevenLabs voice cloning? Gemini needs only 10–30 seconds but requires the speaker to read a consent phrase, and the voice lives in your Google project. ElevenLabs Instant Voice Cloning works from short samples on most plans; Professional Voice Cloning trains on much longer audio, verifies the speaker with a voice captcha and needs the Creator plan or above.

Is VoiceSmiths affiliated with Google? No. VoiceSmiths is an independent studio built on the ElevenLabs API, with free demos of Google's Gemini TTS. Gemini is a trademark of Google LLC.

Design a Gemini voice free →