Gemini 3.8 Flash: Pricing, Agent Upgrades, and Who Should Switch
Google’s newest Flash model targets long-horizon coding and production agents at introductory prices that reset the speed-versus-capability tradeoff.

Bottom line
Gemini 3.8 Flash brings longer-horizon agent work to Google’s fast model tier. Compare its pricing, context, reasoning controls, limitations, and migration case.
Editorial accountability
Who checked this guide
- Evaluation type
- Research-based verification
- Last materially checked
- Evidence
- 4 listed sources
Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.
Editorial basis
What this guidance is based on
- Editorial basis
- Source-led analysis
- Primary references
- 4
- Products covered
- 1
- Last checked
- 2026-09-05
Important limits
- • Features, availability, and pricing can change after publication; confirm consequential details with the provider.
In this guide
*This research-based analysis covers Google’s September 2–3, 2026 Gemini 3.8 Flash materials. DiscoverAI has not independently benchmarked the model. Performance claims and comparisons are Google-reported unless otherwise stated.*
The short answer
Gemini 3.8 Flash is Google’s new general-availability workhorse for long-horizon software engineering, autonomous agents, and complex knowledge workflows. It keeps the Flash family’s emphasis on latency and scale while adding more deliberate multi-step reasoning, iterative tool use, and verification.
The release matters because “fast model” no longer means “simple task model.” Google is aiming 3.8 Flash at workflows that previously pushed buyers toward more expensive frontier tiers. Its introductory API price—$0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026—is also an aggressive invitation to test that premise in production.
What Gemini 3.8 Flash adds
Google describes stronger multi-file coding, deterministic tool execution, autonomous planning, and complex enterprise data work. The model supports text, image, audio, video, and PDF inputs, text output, a one-million-token context window, and 64,000 output tokens. It is available through the Gemini app, Gemini API, Google AI Studio, Gemini Enterprise Agent Platform, AI Mode, and Antigravity.
Three reasoning levels let developers trade latency and spend for more verification. Low effort fits fast chat, incident pipelines, drafts, and routine analysis. Medium is the default for agent and coding work. High is intended for the hardest multi-step tasks. Google explicitly warns that 3.8 Flash may use more tokens on complex jobs because it takes smaller reasoning steps, calls tools iteratively, and checks its work.
That warning is useful. A lower token price does not guarantee a lower task price. Long contexts, repeated tools, grounding requests, and reasoning tokens can dominate the bill. Teams should compare completed-work economics rather than multiplying a headline rate by an optimistic prompt size.
Gemini 3.8 Flash pricing
Through December 31, 2026, paid Standard API usage costs $0.75 per million input tokens and $3.75 per million output tokens, including thinking tokens. Context-cache reads cost $0.075 per million tokens, plus listed storage charges. On January 1, 2027, Google says Standard prices rise to $1.50 input and $7.50 output.
Batch and Flex rates are half of Standard. Priority processing costs more. Google Search and Maps grounding have separate allowances and per-request charges after those allowances. Free-tier data may be used to improve Google products, while the pricing page says paid-tier data is not. Procurement should verify the exact service, region, retention, grounding, and enterprise terms before treating the model rate as the whole cost.
The temporary discount deserves a migration footnote in every forecast. A workflow that only meets its ROI threshold at the promotional rate may stop meeting it in January. Run both present-price and standard-price scenarios now.
What the benchmarks do—and do not—show
Google reports gains over Gemini 3.7 Flash across long-horizon coding and professional evaluations. It lists 61.4% on Vals Finance Agent v2 and 54.9% on HLE-Verified, and showcases strong DeepSWE efficiency. The model card cautions that improved evaluation methods make some results unsuitable for direct comparison with older cards.
These numbers support a test, not a fleet-wide migration. Agent performance depends heavily on prompt scaffolding, tool schemas, retry policies, context management, permissions, and the definition of success. A two-point benchmark gain can be irrelevant if a model’s preferred tool-call format forces an expensive rewrite—or extremely valuable if it removes a human review loop.
Google also lists hallucinations, occasional slowness or timeouts, and higher token use at high effort among known limitations. Its documented knowledge cutoff is uneven by domain. Retrieval and source verification remain necessary for current facts.
Who should test it
Gemini 3.8 Flash is a strong candidate for teams already using Gemini, multimodal input pipelines, Google grounding, or high-volume agents that need more sustained execution. It may also pressure-test whether a premium model is truly necessary for code maintenance, document-heavy analysis, and operational orchestration.
Do not switch solely for introductory pricing if your current system depends on mature prompt behavior, stable structured output, a specific region, or a deeply tested safety envelope. Keep 3.7 Flash available during the evaluation; Google says it remains supported.
A fair migration test
Select 30–50 representative tasks, including long and short jobs, tool failures, ambiguous instructions, and adversarial inputs. Run 3.7 and 3.8 with equivalent effort settings and the same budgets. Measure acceptance rate, tool-call validity, loop failures, correction time, latency percentiles, token use, grounding fees, and cost per accepted result.
Then replay the cost model at January 2027 rates. Only migrate a workflow when quality or labor savings survive that price change and your logs can explain failures. For consequential actions, preserve human approval, least-privilege tools, and rollback.
The verdict
Gemini 3.8 Flash is notable less as another incremental model number than as evidence that cost-efficient models are absorbing serious agent work. Its promotional rates make experimentation easy; its variable reasoning effort and potentially larger token consumption make measurement essential.
The sensible move is a routed rollout: use low effort for latency-sensitive work, higher effort for tasks where verification pays, and retain smaller or older models when they already meet the bar. The best Flash workflow is the one whose completed-output economics still work after the promotion ends.
Sources and verification
Product details and claims were checked against the following primary sources.
Frequently asked questions
What is Gemini 3.8 Flash best for?
Google positions it for long-horizon software engineering, autonomous agents, complex enterprise workflows, multimodal understanding, and high-scale knowledge work.
How much does Gemini 3.8 Flash cost?
Paid Standard pricing is $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, rising to $1.50 and $7.50 on January 1, 2027.
Does Gemini 3.8 Flash have a free tier?
Google lists free Standard usage, but quotas and data-use terms differ from paid service. Check the current pricing page and your workspace terms before testing sensitive data.
Should developers replace Gemini 3.7 Flash immediately?
No. Run representative tasks side by side and compare accepted outputs, tool failures, latency, token use, and costs at both introductory and 2027 prices. Google says 3.7 remains supported.
Found this useful?
Get the next one in your inbox.
One five-minute briefing a week: a meaningful change, a practical workflow, and a clearer tool decision—already filtered for lean teams.
Free · one email a week · unsubscribe any time
Tools mentioned in this article
Google Gemini
Google's deeply integrated AI assistant with unmatched access to Google's ecosystem
Gemini combines powerful AI with Google's vast data ecosystem — Search, Gmail, Docs, YouTube, and more — for a uniquely integrated experience.
Read next
Recommended for you

TypingMind Review 2026: Multi-Model Chat, Pricing, and Privacy
A research-based TypingMind review covering capabilities, pricing, privacy, limitations, and a fair buyer test.
TypingMind unifies multiple AI APIs in a local-first interface, but API spend, browser storage, optional cloud services, and provider policies determine its real value and privacy.
Read guide
Dia Browser Review 2026: Is the AI Browser Worth Switching To?
Looker Studio vs Power BI for AI Analytics Workflows in 2026
Gamma AI Review 2026: Fast Presentations, but Is It Worth Paying For?