GuideUpdated 2026-10-09

Google Releases AQuA: Find Agent Failures Behind Healthy Dashboards

By DiscoverAI Editorial TeamReviewed by DiscoverAI Editorial Review3 min readHow we evaluate

Bottom line

A healthy server can still deliver a wrong answer. Google's AQuA connects conversation failures to the deployed code for human investigation.

A magnifying glass connecting conversation cards to a faulty gear and a code snapshot
Original DiscoverAI editorial illustration. Original editorial illustration; not a product screenshot or measured result.
In this guide
  1. Short answer
  2. The October 8 release
  3. Infrastructure success and user success are different
  4. A proposed evaluation pilot
  5. Turn a finding into a durable improvement
  6. When this is worth the effort

Short answer

AQuA is relevant when your AI agent looks healthy operationally but users still report wrong outcomes. Treat it as an investigation aid. A machine-generated diagnosis still needs review, a reproducible example and a regression check before a fix reaches users.

The October 8 release

Google's announcement introduces AQuA, the Ambient Quality Agent, as open building blocks running alongside an agent in a Google Cloud project. It samples production sessions, reviews failures, groups related findings and verifies clusters against transcripts. Diagnosis can compare those failures with an immutable snapshot of deployed source.

Google says AQuA stays outside the request path and does not write back to the monitored agent. It can propose an edit but does not apply it or open a pull request on its own. Its demonstration is vendor evidence, not proof of the same improvement in another deployment.

Infrastructure success and user success are different

Our editorial interpretation is that teams should track two separate questions: did the system respond, and did it fulfill the authorized request? A fast, technically successful response can still omit a deadline, lose a constraint or describe an action that never happened.

For a support assistant, define success as a grounded answer or an appropriate escalation. For a booking workflow, define the authoritative evidence needed before anything is described as confirmed. Make these rules visible to both the operator and the evaluator.

A proposed evaluation pilot

Start with a small authorized sample after reviewing transcript access and retention. Add deliberately constructed cases with a changed instruction, missing tool output and a refusal to proceed. Prepare an answer key before looking at model-generated findings.

For every reported issue, ask a human reviewer to identify the exact turn, expected behavior and supporting evidence. Also inspect sessions that received no finding. Otherwise you can measure false alarms while overlooking missed failures.

Record disagreement between reviewers and the evaluator rather than silently treating the evaluator as correct. An ambiguous business rule should be clarified before it becomes a software defect. We have not run this pilot.

Turn a finding into a durable improvement

Save a minimized, non-sensitive reproduction with the relevant deployment revision. Change one cause at a time: an instruction, a tool contract or a handoff. Replay that case alongside cases that already worked. A fix that improves one conversation while breaking another is not ready to ship.

Keep an owner and a resolution note for each accepted issue. Revisit it after a model or tool change. Do not treat a period without another observation as proof that the underlying problem disappeared; sample coverage and traffic mix can change.

When this is worth the effort

A team needs usable transcripts, deployed-version records and someone responsible for corrections. If those foundations are missing, establish them first. More diagnostic output without ownership can become another ignored dashboard.

For the broader adoption context, read our [Gemini work-agent rollout guide](/articles/gemini-work-agent-governed-rollout-2026). For selecting a model before production, use the [AI model selector](/ai-model-selector).

Transparency

How this guide was checked

Editorial accountability

Who checked this guide

Meet the editorial team →
Evaluation type
Research-based verification
Last materially checked
Evidence
1 listed sources

Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.

Editorial basis

What this guidance is based on

Editorial basis
Source-led analysis
Primary references
1
Products covered
1
Last checked
2026-10-09

Important limits

  • • DiscoverAI has not independently tested the product or reproduced vendor benchmarks.
  • • Availability, prices and policies may change. Evaluation exercises are proposed reader-run tests.

Sources and verification

Product details and claims were checked against the following primary sources.

Frequently asked questions

What is AQuA?

Google's Ambient Quality Agent is an open reference implementation for investigating production-agent failure patterns.

Does AQuA automatically fix production code?

Google says it proposes diagnoses and edits but does not apply edits or open pull requests itself.

Is a successful HTTP response enough?

No. A request can complete successfully while the answer ignores a constraint or invents a result.

Did DiscoverAI reproduce the demonstration?

No. The assessment is based on Google's announcement; the pilot described here is proposed.

Free AI tool buyer checklist

Make the next AI subscription earn its place.

Get the printable buyer checklist now, plus one useful five-minute AI briefing each week.

Free · one email a week · unsubscribe any timeRead a sample email →Preview the checklist →

Tools mentioned in this article

Google Gemini

Google's deeply integrated AI assistant with unmatched access to Google's ecosystem

Not rated

Gemini combines powerful AI with Google's vast data ecosystem — Search, Gmail, Docs, YouTube, and more — for a uniquely integrated experience.

FreemiumChatbotsProductivity

Read next