Google Launches Gemini 3.5 Transcribe for Real-Time Voice Apps
Google's dedicated speech-to-text model targets voice agents, captions, and call analytics, but language count and vendor accuracy claims are not substitutes for testing real audio.

Bottom line
Gemini 3.5 Transcribe is a dedicated real-time speech-to-text model for more than 85 languages. It gives developers a new path for voice agents and live captions, while production adoption still depends on domain accuracy, latency, privacy, and cost per accepted transcript.
Editorial accountability
Who checked this guide
- Evaluation type
- Research-based verification
- Last materially checked
- Evidence
- 4 listed sources
Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.
Editorial basis
What this guidance is based on
- Editorial basis
- Source-led analysis
- Primary references
- 4
- Products covered
- 1
- Last checked
- 2026-09-24
Important limits
- • DiscoverAI has not independently benchmarked Gemini 3.5 Transcribe.
- • Language coverage and vendor performance claims do not establish equal accuracy across accents, domains, audio conditions, or production workflows.
In this guide
Short answer
Google launched Gemini 3.5 Transcribe in the Gemini API and Google AI Studio as a dedicated real-time speech-to-text model for more than 85 languages. Google positions it for voice agents, live captioning, meeting notes, media transcription, and post-call analytics, with context-aware handling intended to improve noisy audio and specialized terminology.
The opportunity is a simpler audio layer for applications already using Gemini: developers can stream speech, add domain context, and turn transcripts into downstream summaries, routing, or analytics. The launch does not prove the model will be accurate enough for a particular language, accent, room, industry, or regulated workflow.
Free creative AI buyer checklist
Measure approved assets, rights, and revision time.
Get a checklist for quality, credits, consent, licensing, and cost per approved asset—plus weekly creative-AI changes.
What the model changes
Speech products often combine a transcription model with separate language models for cleanup, terminology, speaker interpretation, and downstream actions. A dedicated Gemini model can bring those stages closer to the same developer platform and make it easier to maintain context across a live interaction.
Real-time output matters because a voice agent cannot wait for an entire recording before responding. Captions also need stable partial results and sensible corrections. Teams should evaluate both final transcript quality and the behavior users experience while the words are still arriving.
Where the opportunity is strongest
Customer-service and appointment agents can use live transcription to understand callers before deciding what to retrieve or do. Accessibility products can generate captions across supported languages. Sales and research teams can create searchable call records, while media teams can accelerate rough transcripts and subtitle preparation.
The highest-value workflow is not necessarily the one with the lowest word-error rate. A model may be useful if it reliably captures names, numbers, intent, consent, and escalation signals while reducing review time. Conversely, a polished transcript can still be unsafe if it changes a dosage, account number, negation, or contractual commitment.
What Google's claims do not prove
Support for more than 85 languages does not mean equal accuracy, latency, or feature quality across all of them. Google-reported performance does not establish results for a buyer's microphones, codecs, overlapping speakers, background noise, dialects, proper nouns, or adversarial audio. Context-aware cleanup can also make text look more fluent while concealing a meaning-changing error.
Transcription availability does not establish that a complete voice agent is reliable. Retrieval, reasoning, tool calls, text-to-speech, interruptions, and human handoff introduce separate failure modes. Privacy and compliance also depend on the chosen API settings, region, account terms, logging, retention, consent, and downstream storage—not the model name alone.
A production evaluation set
Build a consented, de-identified set of representative audio covering supported languages, accents, devices, noise levels, jargon, interruptions, silence, numbers, names, and high-consequence statements. Create human-verified reference transcripts and label the fields that must never be guessed.
Measure word and entity error rates, meaning-changing errors, first-token and finalization latency, correction stability, speaker handling, interruption behavior, review minutes, failure and retry rates, and cost per accepted minute. Evaluate each important language separately rather than averaging away weak segments. Route low-confidence or high-risk cases to a human, and retain the original audio only when policy and consent permit it.
What developers should do now
Prototype one bounded workflow in Google AI Studio or the Gemini API, then validate it with real conditions before connecting it to customer records or irreversible actions. Confirm current model identifiers, quotas, regional availability, pricing, data-use terms, retention controls, and deprecation policy in official documentation.
For voice agents, keep transcript evidence visible to reviewers, require confirmation for critical values, and design a graceful fallback when the stream is unclear or disconnected. This is a research-based news analysis. DiscoverAI has not independently benchmarked Gemini 3.5 Transcribe, and model behavior, prices, quotas, and availability can change.
Sources and verification
Product details and claims were checked against the following primary sources.
Frequently asked questions
What is Gemini 3.5 Transcribe?
It is Google's dedicated real-time speech-to-text model in the Gemini API and Google AI Studio for voice agents, captions, transcription, and audio analytics.
How many languages does Gemini 3.5 Transcribe support?
Google says the model supports more than 85 languages, but teams should evaluate every important language and accent with representative audio.
Does it make a complete voice agent?
No. Transcription is one component. Retrieval, reasoning, tool execution, speech generation, interruption handling, and human escalation need separate design and testing.
How should teams test the model?
Use consented representative audio and measure entity and meaning-changing errors, latency, correction stability, reviewer time, failures, and cost per accepted minute for each important language.
Recommended tool
Use Google Gemini if this workflow fits your team
It has one of the clearest workflow fits in its category and is easier to recommend than tools that only look impressive in demos.
Tools mentioned in this article
Google Gemini
Google's deeply integrated AI assistant with unmatched access to Google's ecosystem
Gemini combines powerful AI with Google's vast data ecosystem — Search, Gmail, Docs, YouTube, and more — for a uniquely integrated experience.
Read next
Recommended for you

Google Beam Expands 3D Video Meetings to Six Countries
HP Dimension with Google Beam is moving beyond pilots, but buyers still need to test whether a specialized 3D room produces enough measurable value to justify its footprint and cost.
Google Beam is expanding its headset-free 3D meeting platform to six countries, 18 channel partners, and bookable Industrious locations. The rollout creates a practical trial path for distributed teams, while Google's internal results still need independent validation.
Read guide
Gemini Notebook Adds Live Study Conversations, Quizzes and Short Videos
Google Ads Adds AI Max Brief Controls and Unified Journey Reporting
Google Expands AI Plans With Voice, Pics, and Sheets Canvas