What did Google release, and when?

The dated announcement is the 8 October 2026 Developers Blog post, “The Outer Loop, Insights First.” Authors listed are Dima Melnyk, a Cloud AI product manager, and Elia Secchi, a solutions specialist. We opened that page on 9 October. The object being released is AQuA — Ambient Quality Agent — described as composable building blocks, not a managed Google Cloud product with an SLA.

Google’s setup problem is the gap after the first 80%. Offline evals and a tight inner loop, the post says, can get an agent through anticipated test cases in days. After launch, usage drifts, models and tools change, and health checks stay green while conversations fail. The examples on the page are concrete: recording seat 3A as confirmed without asking the seat-selection sub-agent whether 3A is free, or recommending a Florentine steakhouse to a traveler whose profile says vegan.

That is a production-quality problem, not a new chat model. It sits next to last week’s Gemini agent for work, which is a user-facing enterprise agent, and next to our five-point agent evaluation guide. AQuA is the side-car that is supposed to notice when the agent already in production is quietly wrong.

How does the five-stage sweep work?

AQuA, the post says, runs unattended in the same Google Cloud project as the observed agent. It samples trajectories from Cloud Trace, Cloud Logging or BigQuery on a schedule, after each deployment, or on demand. Raw transcripts, source snapshots and BigQuery tables are described as staying inside the project boundary. AQuA “never sits in the request path or writes back to your agent.”

Each run is a five-stage pipeline. Sample: a random draw of up to 1,000 recent sessions, capped so run cost stays predictable. Review: a single-pass session_review judge grades each session against a nine-point checklist, with the agent’s instructions, tools and an optional goal.md in view, and writes an actual/expected finding when a session fails. Cluster: findings that share a failure mechanism become candidate issues. Verify: a separate model checks each candidate against up to three full transcripts and drops unsupported clusters. Track: surviving clusters are matched in BigQuery as NEW, RECURRING, or auto-RESOLVED after 14 days unseen.

Root-cause analysis is a separate, on-demand step from the dashboard chat or `agents-cli aqua run`. Google says AQuA then reads the failing traces against the immutable source snapshot stored at deploy time. If the defect is in the repo, it cites file-and-line ranges the server validates against that revision and proposes an edit. If the fault is an upstream dependency, a handoff or a retrieved payload, it attributes the step without a code diff. It never applies the edit and never opens a pull request. That “what actually changed” instinct is closer to ThinkingBox’s backend-state checks than to a preference leaderboard, but AQuA is a Google Cloud recipe, not a public benchmark.

What did the travel-concierge demo show?

The worked example is travel-concierge from google/adk-recipes, a multi-agent trip planner that routes across inspiration, place, POI, planning, flight-search, seat-selection, booking and trip-phase sub-agents. Google says attaching AQuA took three agents-cli commands, wrote a Revision 1 source snapshot to Cloud Storage, and used a goal.md that told the judge to ignore tone and abandoned browsing sessions.

The sweep Google reports is 32 replays of four scripted traveler journeys, producing 1,583 OpenTelemetry spans. Five sessions passed; 27 sessions produced 42 findings; clustering made 9 candidates; Gemini 3.7 Flash, as the verifier, rejected 3. The six surviving issues were led by a seat-availability bypass in planning_agent (15 sessions), a vegan-profile drop between inspiration_agent and place_agent (7 sessions), and a poi_agent prompt that required Maps fields while `tools=[]` (5 sessions).

Diagnosis against Revision 1’s 33-file snapshot, using Gemini 3.8 Flash, is said to have landed on line 93 of planning/prompt.py and line 23 of inspiration/prompt.py. After those two one-line prompt edits and a Revision 2 deploy, the same 32 sessions are reported as: seat bypass 15 to 2; vegan omissions 7 to 0; full-session passes 5/32 to 13/32; the untouched poi_agent `tools=[]` defect still in the queue. Those are Google’s scripted replays of Google’s demo agent. They are not a live-traffic study.

How can teams try it, and what does it depend on?

The post’s no-cloud path is `git clone https://github.com/google/adk-recipes.git`, then `core/python/ambient-quality-agent` and `make demo`, which serves the dashboard on synthetic data with no credentials and no model calls. We opened that tree on 9 October. The directory listing included a README, Dockerfile, Terraform, UI, tests, skills and an agents-cli extension manifest. The parent repository, google/adk-recipes, is Apache-2.0.

The README we opened is stricter than the blog’s three-command sketch. It wants Python 3.11–3.13, uv, npm, agents-cli 1.6 or later, Terraform, and an organization-backed Google Cloud project because the dashboard sits behind Identity-Aware Proxy. The documented pin is `agents-cli extension add google/adk-recipes#core/python/ambient-quality-agent --ref ambient-quality-agent/v0.1.0`. Without an organization, IAP can be disabled and the UI opened through `agents-cli aqua ui-proxy`.

Telemetry sources listed in the README are the ADK BigQuery Agent Analytics plugin, Cloud Telemetry linked BigQuery `_AllSpans` datasets, and agents-cli default Cloud Trace / Cloud Logging. That is a Google Cloud and ADK-shaped recipe. It is not a drop-in for an agent that only logs to a laptop.

What did Google say about cost, judges and trust?

The post is explicit that a diagnosing agent will sometimes be wrong. Clusters are claims until verified. The insight schema has no confidence field. Occurrences link to Cloud Trace session IDs. Root-cause citations that point at a missing file or line are rejected. Dismissed findings stay dismissed.

Cost figures are Google’s, at “standard Gemini platform pricing.” A 96-session single-agent sweep is quoted at $0.70 total, about $0.007 per session, using Gemini 3.1 Pro and Gemini 3.7 Flash. The 32-session travel-concierge sweep, with far deeper traces, is quoted at $3.76, about $0.12 per session. On-demand root-cause diagnosis with Gemini 3.8 Flash is quoted at $0.33 to $2.47 per insight. An 87-trace internal benchmark is said to have had the verifier reject 4 of 24 candidate clusters. We did not rerun those bills.

Google also lists the problems this MVP does not solve: random sampling of 1,000 sessions can miss a silent 1% break on a critical workflow; goal.md does not automatically match domain experts; verification is capped at 50 clusters × 3 transcripts; hundred-turn coding traces need compaction the current design does not have; `adk run --replay` cannot reconstruct external state that has moved on.

What is not established?

We did not deploy AQuA, attach it to an agent, or replay travel-concierge. We did not confirm that every agents-cli flag in the README matches a released binary on 9 October. We did not see AQuA run against live, non-scripted user traffic.

The 15-to-2 and 7-to-0 lines are Google’s after two prompt edits on a 32-session scripted set. They are not evidence that AQuA will find the same class of bug in a different harness, and they are not a claim that the remaining 19 failing sessions are now safe.

This is an evidence review of pages opened on 9 October. It is not a first-hand AQuA run and not a recommendation to replace human review of production agents.

Common questions

Does AQuA block or rewrite live agent replies?

No. Google says it stays outside the request path and does not write back to the observed agent. It samples after the fact and proposes edits for a human or a coding agent to apply.

Do I need Google Cloud to see the dashboard?

Not for the synthetic demo. `make demo` in the adk-recipes AQuA directory serves local fake data. Attaching AQuA to a real agent is documented as a Google Cloud, ADK and agents-cli path.

Are the travel-concierge gains a production result?

No. They are Google’s replay of 32 scripted sessions on its own demo agent after two one-line prompt edits. The poi_agent tools=[] defect was left unfixed in that write-up.

THE TAKEAWAY

What to remember

Read the 8 October post for the five-stage loop and the no-auto-apply rule. Clone adk-recipes for the demo. Treat the travel-concierge percentages as a scripted Google example, not a field trial.

Sources & further reading

  1. The Outer Loop, Insights First: An Ambient Quality Agent That Diagnoses Your Production Agent ↗
  2. Ambient Quality Agent (AQuA) README ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories