GuideUpdated 2026-09-26

Gemini 3.8 Flash TTS Turns Voice Design Into a Production Workflow

Natural-language voice design and two-speaker direction broaden production options, while consent, regional limits, consistency, and rights stay decisive.

By DiscoverAI Editorial TeamReviewed by DiscoverAI Editorial Review2 min readVideo, Audio & CreativeHow we evaluate
Paper-cut editorial illustration of a voice director shaping multilingual character performances through consent, pronunciation, two-speaker, consistency, rights, and approval checkpoints
Original DiscoverAI editorial illustration. Editorial illustration: a voice director shaping multilingual character performances through consent, pronunciation, two-speaker, consistency, rights, and approval checkpoints.

Bottom line

Google released Gemini 3.8 Flash TTS and Flash-Lite TTS on September 23, 2026. Flash targets creative direction and custom character voices; Flash-Lite targets high-volume dubbing, content, and voice agents. Google advertises more than 100 languages and dialects, 2,000 production-ready voices, long-form and two-speaker control, SynthID watermarking, and voice replication from a 30-second sample with consent verification. Those are product claims, not proof that every voice, language, or long recording will meet a buyer's quality bar.

Editorial accountability

Who checked this guide

Meet the editorial team →
Evaluation type
Research-based verification
Last materially checked
Evidence
4 listed sources

Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.

Editorial basis

What this guidance is based on

Editorial basis
Source-led analysis
Primary references
4
Products covered
3
Last checked
2026-09-26

Important limits

  • • Vendor claims and demonstrations are not independent proof of outcomes.
  • • Availability, pricing, policies, and behavior can change.
In this guide
  1. Short answer
  2. Flash versus Flash-Lite
  3. The production opportunity
  4. Consent and regional boundaries
  5. Evidence limits
  6. A credible buyer test
  7. Bottom line

Short answer

Google released Gemini 3.8 Flash TTS and Flash-Lite TTS on September 23, 2026. Flash targets creative direction and custom character voices; Flash-Lite targets high-volume dubbing, content, and voice agents. Google advertises more than 100 languages and dialects, 2,000 production-ready voices, long-form and two-speaker control, SynthID watermarking, and voice replication from a 30-second sample with consent verification. Those are product claims, not proof that every voice, language, or long recording will meet a buyer's quality bar.

Free creative AI buyer checklist

Measure approved assets, rights, and revision time.

Get a checklist for quality, credits, consent, licensing, and cost per approved asset—plus weekly creative-AI changes.

Free · about 5 minutes · one email a week · unsubscribe any time

Free · one email a week · unsubscribe any timePreview the checklist →

Flash versus Flash-Lite

Flash is the better candidate when performance direction, original character design, or a consistent brand voice matters. Flash-Lite is positioned for scale and cost efficiency. Teams should run the same approved script through both because the lower-cost model may be enough for routine narration while premium direction is reserved for demanding scenes.

The production opportunity

Natural-language instructions can move voice generation closer to a director's workflow: specify role, accent, pace, emotion, backchanneling, and line-level delivery rather than selecting one preset. Native two-speaker staging can simplify podcasts, learning content, games, localization, and scripted agent dialogue. Saved custom voices may reduce drift across a series.

Google says replicated voices require a matching verbal consent recording and all generated Gemini Audio is watermarked with SynthID. Voice replication through AI Studio is not available in Illinois, Texas, the EEA, UK, Switzerland, or India at launch. Teams remain responsible for talent contracts, reuse scope, revocation, minors, unions, disclosure, and downstream distribution.

Evidence limits

Google cites top positions on Hume's Voice Design Benchmark and Voice Arena language evaluations. Buyers should treat those as useful vendor-selected evidence, not a substitute for their own scripts, listeners, accents, devices, and editorial standards. Long-form consistency, pronunciation, emotional appropriateness, and watermark durability need independent testing.

A credible buyer test

Create a 60-minute multilingual set with names, numbers, technical terms, emotional shifts, two speakers, long passages, and accessibility content. Blind-score naturalness, intelligibility, pronunciation, speaker separation, drift, edit time, consent records, rejected generations, latency, and cost per approved minute against a human recording and two alternatives.

Bottom line

Gemini 3.8 TTS expands Google from speech generation into directed voice production. Its value will come from approved audio per dollar, not the size of the voice library. Consent verification and watermarking are useful controls, but rights management and human editorial review still belong to the publisher.

Sources and verification

Product details and claims were checked against the following primary sources.

Frequently asked questions

What is Gemini 3.8 Flash TTS?

It is Google's speech-generation model for directed character voices, long-form audio, two-speaker scenes, and voice replication. Flash-Lite is the scale-oriented companion model.

Can Gemini 3.8 TTS clone a voice?

Google says Flash can replicate a voice from a 30-second sample after a matching verbal consent recording, subject to regional and account availability.

How many languages are supported?

Google advertises more than 100 languages and dialects. Production teams should test pronunciation, code-switching, cultural fit, and consistency in every target locale.

Is Gemini-generated speech watermarked?

Google says Gemini Audio output contains SynthID. Watermarking supports provenance but does not replace disclosure, rights clearance, or controls on redistribution.

Free creative AI buyer checklist

Measure approved assets, rights, and revision time.

Get a checklist for quality, credits, consent, licensing, and cost per approved asset—plus weekly creative-AI changes.

Free · one email a week · unsubscribe any timePreview the checklist →

Recommended tool

Use Google Gemini if this workflow fits your team

It has one of the clearest workflow fits in its category and is easier to recommend than tools that only look impressive in demos.

Tools mentioned in this article

Google Gemini

Google's deeply integrated AI assistant with unmatched access to Google's ecosystem

4.2

Gemini combines powerful AI with Google's vast data ecosystem — Search, Gmail, Docs, YouTube, and more — for a uniquely integrated experience.

FreemiumChatbotsProductivity

ElevenLabs

A leading AI voice platform for text to speech, voice cloning, speech to text, dubbing, and conversational agents

4.6

ElevenLabs combines premium text to speech, voice cloning, multilingual audio generation, speech to text, developer APIs, and voice agents in one AI audio platform.

FreemiumAudioMarketing

Fish Audio

Create expressive AI speech, cloned voices, dialogue, and transcription through a studio or API

4.4

Fish Audio is an AI voice platform for expressive text-to-speech, rapid voice cloning, multi-speaker dialogue, transcription, audio production, and developer integrations.

FreemiumAudioContent Creation

Read next

More on Video, Audio & Creative →