Google introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on September 23, 2026. The paired models let users describe a new voice, direct how individual lines are spoken and, with Flash TTS, replicate a voice from a 30-second sample.
The models are rolling out through the Gemini API and Google AI Studio, according to Google’s announcement. Features vary by product and region: enterprise API access and voice remixing are still forthcoming, and AI Studio voice replication has geographic restrictions.
For speech producers, a voice can become a reusable part of a project, while a script can carry its own performance directions. Google is also making consent verification and provenance tools part of the cloning workflow, though their effectiveness still needs independent scrutiny.
Voice Design Goes Beyond Choosing a Preset
Gemini 3.8 Flash TTS lets users create a voice by describing its role, accent and vocal characteristics in natural language. Google says it supports more than 100 languages and dialects and offers a separate library of more than 2,000 production-ready voices, including Mexican Spanish, Quebec French and Scots English varieties.
Creators can select a library voice, prompt a new character voice or replicate a specific person’s voice with permission. Google also says users can save custom voices for reuse across projects. Voice remixing, which would let users adjust characteristics such as timbre, pitch, pace and accent on a library voice, has been announced but is not yet available.
A game team, for example, could design a character voice and use it across scenes instead of settling for the closest preset. Whether that identity remains convincing through lengthy recording sessions is a separate question. Google claims its models can maintain character timbre across hours of audio, but the announcement is not an independent long-form test.
Language coverage says little on its own about performance in each dialect. A catalog may contain a regional variety without consistently getting its pronunciation, pacing or local phrasing right. Localization teams will need to listen to actual output, particularly where a voice is meant to represent a specific community.
The Script Can Direct the Performance
Both Flash TTS and Flash-Lite TTS support line-by-line control over delivery. Scripts can include directions for tone and pacing, along with vocal cues such as laughs, sighs and brief listening responses. Google says the models can stage two-speaker conversations while keeping the voices separate.
In the Gemini API speech-generation documentation, a developer supplies text and assigns a voice through speech configuration. Turn-level speech_metadata can specify a speaker and style; a conversational mode handles two-speaker exchanges. An application can therefore change delivery between lines without creating a separate voice for every emotional beat.
These text-to-speech models turn a supplied script into audio. Google’s documentation distinguishes them from the Live API, which is built for more open-ended, interactive audio exchanges. A scripted podcast scene and an assistant responding to unpredictable conversation may both need natural-sounding speech, but they place different demands on the system.
Line-level direction could reduce retakes in audiobooks, games and dubbed content, where a flat reading is often less useful than a consistent voice that responds to scene context. Editing remains necessary: a direction in a script is a request to the model, not a guarantee of the intended performance.
Flash-Lite Is the Scale Option, but Costs Need Checking
Google positions Gemini 3.8 Flash TTS for character design and detailed creative direction. Flash-Lite TTS is its cost-efficient, high-volume counterpart, aimed at work such as dubbing, audio content production and voice agents. Both have line-by-line delivery controls; Google specifically associates voice design from scratch and 30-second replication with Flash TTS.
“Cost-efficient” is a vendor description, not a calculated saving for a particular project. The announcement provides neither rates nor a measured throughput comparison from which to work out the difference. A localization business generating many hours of dialogue would need to check current API pricing, latency and the amount of editing each model’s output requires.
Cheaper generation may not produce a cheaper finished recording if it creates more correction work. A production that needs straightforward narration at volume, meanwhile, may have little reason to pay for the fullest character-design controls. Google has described the models’ intended roles; production economics depend on use-case testing.
Replication Requires More Than a Voice Sample
Flash TTS can build a vocal profile from a 30-second audio sample of a user’s voice or one they have the rights to use. Google says the user must also submit a verbal consent recording from the voice owner, and its system checks that the consent speaker matches the reference speaker before creating the replicated voice.
That second recording makes cloning more involved than uploading a short clip obtained elsewhere. It also addresses a production need: a performer who has authorized a project could provide a reusable voice for additional lines or localization without recording every variation.
A speaker match does not verify every term of permission. The consent process described by Google does not, in its announcement, explain how the service handles disputes over commercial rights, withdrawal of consent or attempts to submit a deceptive consent recording. Those are questions to test and document, not evidence that the check fails.
Voice actors and clients need to know what uses were authorized, who can generate new lines and what happens to an established voice if that authorization changes.
Watermarks and Credentials Address Provenance, Not Permission
Google says every clip generated by its Gemini Audio models carries an imperceptible SynthID watermark. Replicated voices also receive C2PA content credentials, it says. The watermark is embedded in generated audio to support detection, while content credentials are intended to communicate its provenance.
Neither measure makes an unauthorized recording harmless or prevents someone from sharing convincing synthetic speech. A marker on an output also cannot, by itself, establish that the speaker agreed to every use of their voice. Consent checking happens before replication; provenance measures help identify content after generation.
Google describes these safeguards as protections for developers and voice talent, but the release does not establish how reliably they hold up across common audio editing and distribution workflows. Independent testing would be more informative than treating a watermark or credential as proof that misuse has been solved.
Availability Varies by Product and Region
Both models began rolling out for developers in the Gemini API and Google AI Studio on September 23. Google also listed Flash TTS for Gemini Notebook and Flash-Lite TTS for Google Vids. API access through Gemini Enterprise is coming later, as is voice remixing. “Rolling out” does not promise immediate access to every capability for every account.
Frequently Asked Questions
5 questions
1Can Gemini 3.8 Flash TTS Clone a Voice From 30 Seconds of Audio?
Yes. Google says Gemini 3.8 Flash TTS can create a vocal profile from a 30-second sample of your voice or one you have the rights to use. Replication also requires a separate verbal consent recording from the voice owner, which Google says it checks against the reference speaker before creating the voice.
2
Sources
- Google’s announcementblog.google
- Gemini API speech-generation documentationai.google.dev
- Hacker News discussion of the announcementnews.ycombinator.com
- The Next Web’s coveragethenextweb.com






