Gemini 3.8 Flash TTS Turns Voice Design Into a Production Workflow
Natural-language voice design and two-speaker direction broaden production options, while consent, regional limits, consistency, and rights stay decisive.

Bottom line
Google released Gemini 3.8 Flash TTS and Flash-Lite TTS on September 23, 2026. Flash targets creative direction and custom character voices; Flash-Lite targets high-volume dubbing, content, and voice agents. Google advertises more than 100 languages and dialects, 2,000 production-ready voices, long-form and two-speaker control, SynthID watermarking, and voice replication from a 30-second sample with consent verification. Those are product claims, not proof that every voice, language, or long recording will meet a buyer's quality bar.
Editorial accountability
Who checked this guide
- Evaluation type
- Research-based verification
- Last materially checked
- Evidence
- 4 listed sources
Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.
Editorial basis
What this guidance is based on
- Editorial basis
- Source-led analysis
- Primary references
- 4
- Products covered
- 3
- Last checked
- 2026-09-26
Important limits
- • Vendor claims and demonstrations are not independent proof of outcomes.
- • Availability, pricing, policies, and behavior can change.
In this guide
Short answer
Google released Gemini 3.8 Flash TTS and Flash-Lite TTS on September 23, 2026. Flash targets creative direction and custom character voices; Flash-Lite targets high-volume dubbing, content, and voice agents. Google advertises more than 100 languages and dialects, 2,000 production-ready voices, long-form and two-speaker control, SynthID watermarking, and voice replication from a 30-second sample with consent verification. Those are product claims, not proof that every voice, language, or long recording will meet a buyer's quality bar.
Free creative AI buyer checklist
Measure approved assets, rights, and revision time.
Get a checklist for quality, credits, consent, licensing, and cost per approved asset—plus weekly creative-AI changes.
Flash versus Flash-Lite
Flash is the better candidate when performance direction, original character design, or a consistent brand voice matters. Flash-Lite is positioned for scale and cost efficiency. Teams should run the same approved script through both because the lower-cost model may be enough for routine narration while premium direction is reserved for demanding scenes.
The production opportunity
Natural-language instructions can move voice generation closer to a director's workflow: specify role, accent, pace, emotion, backchanneling, and line-level delivery rather than selecting one preset. Native two-speaker staging can simplify podcasts, learning content, games, localization, and scripted agent dialogue. Saved custom voices may reduce drift across a series.
Consent and regional boundaries
Google says replicated voices require a matching verbal consent recording and all generated Gemini Audio is watermarked with SynthID. Voice replication through AI Studio is not available in Illinois, Texas, the EEA, UK, Switzerland, or India at launch. Teams remain responsible for talent contracts, reuse scope, revocation, minors, unions, disclosure, and downstream distribution.
Evidence limits
Google cites top positions on Hume's Voice Design Benchmark and Voice Arena language evaluations. Buyers should treat those as useful vendor-selected evidence, not a substitute for their own scripts, listeners, accents, devices, and editorial standards. Long-form consistency, pronunciation, emotional appropriateness, and watermark durability need independent testing.
A credible buyer test
Create a 60-minute multilingual set with names, numbers, technical terms, emotional shifts, two speakers, long passages, and accessibility content. Blind-score naturalness, intelligibility, pronunciation, speaker separation, drift, edit time, consent records, rejected generations, latency, and cost per approved minute against a human recording and two alternatives.
Bottom line
Gemini 3.8 TTS expands Google from speech generation into directed voice production. Its value will come from approved audio per dollar, not the size of the voice library. Consent verification and watermarking are useful controls, but rights management and human editorial review still belong to the publisher.
Sources and verification
Product details and claims were checked against the following primary sources.
Frequently asked questions
What is Gemini 3.8 Flash TTS?
It is Google's speech-generation model for directed character voices, long-form audio, two-speaker scenes, and voice replication. Flash-Lite is the scale-oriented companion model.
Can Gemini 3.8 TTS clone a voice?
Google says Flash can replicate a voice from a 30-second sample after a matching verbal consent recording, subject to regional and account availability.
How many languages are supported?
Google advertises more than 100 languages and dialects. Production teams should test pronunciation, code-switching, cultural fit, and consistency in every target locale.
Is Gemini-generated speech watermarked?
Google says Gemini Audio output contains SynthID. Watermarking supports provenance but does not replace disclosure, rights clearance, or controls on redistribution.
Recommended tool
Use Google Gemini if this workflow fits your team
It has one of the clearest workflow fits in its category and is easier to recommend than tools that only look impressive in demos.
Tools mentioned in this article
Google Gemini
Google's deeply integrated AI assistant with unmatched access to Google's ecosystem
Gemini combines powerful AI with Google's vast data ecosystem — Search, Gmail, Docs, YouTube, and more — for a uniquely integrated experience.
ElevenLabs
A leading AI voice platform for text to speech, voice cloning, speech to text, dubbing, and conversational agents
ElevenLabs combines premium text to speech, voice cloning, multilingual audio generation, speech to text, developer APIs, and voice agents in one AI audio platform.
Fish Audio
Create expressive AI speech, cloned voices, dialogue, and transcription through a studio or API
Fish Audio is an AI voice platform for expressive text-to-speech, rapid voice cloning, multi-speaker dialogue, transcription, audio production, and developer integrations.
Read next
