In this guide
Short answer
AQuA is relevant when your AI agent looks healthy operationally but users still report wrong outcomes. Treat it as an investigation aid. A machine-generated diagnosis still needs review, a reproducible example and a regression check before a fix reaches users.
The October 8 release
Google's announcement introduces AQuA, the Ambient Quality Agent, as open building blocks running alongside an agent in a Google Cloud project. It samples production sessions, reviews failures, groups related findings and verifies clusters against transcripts. Diagnosis can compare those failures with an immutable snapshot of deployed source.
Google says AQuA stays outside the request path and does not write back to the monitored agent. It can propose an edit but does not apply it or open a pull request on its own. Its demonstration is vendor evidence, not proof of the same improvement in another deployment.
Infrastructure success and user success are different
Our editorial interpretation is that teams should track two separate questions: did the system respond, and did it fulfill the authorized request? A fast, technically successful response can still omit a deadline, lose a constraint or describe an action that never happened.
For a support assistant, define success as a grounded answer or an appropriate escalation. For a booking workflow, define the authoritative evidence needed before anything is described as confirmed. Make these rules visible to both the operator and the evaluator.
A proposed evaluation pilot
Start with a small authorized sample after reviewing transcript access and retention. Add deliberately constructed cases with a changed instruction, missing tool output and a refusal to proceed. Prepare an answer key before looking at model-generated findings.
For every reported issue, ask a human reviewer to identify the exact turn, expected behavior and supporting evidence. Also inspect sessions that received no finding. Otherwise you can measure false alarms while overlooking missed failures.
Record disagreement between reviewers and the evaluator rather than silently treating the evaluator as correct. An ambiguous business rule should be clarified before it becomes a software defect. We have not run this pilot.
Turn a finding into a durable improvement
Save a minimized, non-sensitive reproduction with the relevant deployment revision. Change one cause at a time: an instruction, a tool contract or a handoff. Replay that case alongside cases that already worked. A fix that improves one conversation while breaking another is not ready to ship.
Keep an owner and a resolution note for each accepted issue. Revisit it after a model or tool change. Do not treat a period without another observation as proof that the underlying problem disappeared; sample coverage and traffic mix can change.
When this is worth the effort
A team needs usable transcripts, deployed-version records and someone responsible for corrections. If those foundations are missing, establish them first. More diagnostic output without ownership can become another ignored dashboard.
For the broader adoption context, read our [Gemini work-agent rollout guide](/articles/gemini-work-agent-governed-rollout-2026). For selecting a model before production, use the [AI model selector](/ai-model-selector).
Transparency
How this guide was checked
Editorial accountability
Who checked this guide
- Evaluation type
- Research-based verification
- Last materially checked
- Evidence
- 1 listed sources
Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.
Editorial basis
What this guidance is based on
- Editorial basis
- Source-led analysis
- Primary references
- 1
- Products covered
- 1
- Last checked
- 2026-10-09
Important limits
- • DiscoverAI has not independently tested the product or reproduced vendor benchmarks.
- • Availability, prices and policies may change. Evaluation exercises are proposed reader-run tests.
Sources and verification
Product details and claims were checked against the following primary sources.
Frequently asked questions
What is AQuA?
Google's Ambient Quality Agent is an open reference implementation for investigating production-agent failure patterns.
Does AQuA automatically fix production code?
Google says it proposes diagnoses and edits but does not apply edits or open pull requests itself.
Is a successful HTTP response enough?
No. A request can complete successfully while the answer ignores a constraint or invents a result.
Did DiscoverAI reproduce the demonstration?
No. The assessment is based on Google's announcement; the pilot described here is proposed.
Tools mentioned in this article
Google Gemini
Google's deeply integrated AI assistant with unmatched access to Google's ecosystem
Gemini combines powerful AI with Google's vast data ecosystem — Search, Gmail, Docs, YouTube, and more — for a uniquely integrated experience.
Read next

