What did Asana and OpenAI publish?

Asana’s Inside Asana article dated for this cycle explains how StackAI’s browser agent was made cheaper and faster without cutting answer quality. OpenAI’s customer story, “Asana cuts model costs 76x in browser tests with GPT-6.1 Sol,” dated 9 October 2026, retells the same experiment for OpenAI’s audience.

StackAI is the workflow-automation platform Asana acquired; the browser agent navigates sites, fills forms, and gathers fields for customer automations. Frank Hidalgo, StackAI CTO at Asana, directed GPT-6 Astra in Codex to map the agent, propose fixes, and run comparisons he estimates would have taken one to two months by hand.

GPT-6.1 Sol is the cost-focused GPT-6 variant we covered earlier; see GPT-6.1 Sol’s cost-capability note for OpenAI’s own model framing, separate from this browser-agent study.

Two pages, one study: Asana’s engineering narrative and OpenAI’s customer story. Prefer Asana’s methodology details for cache-hit rates and history policies, and treat OpenAI’s page as a distribution channel for the same experiment. Neither page is an independent replication.

San Francisco downtown towers seen from a hillside park under a blue sky. No identifiable people appear.
San Francisco from Ina Coolbrith Park, 15 July 2021. CC BY-SA 4.0 archival photograph by Frank Schulenburg via Wikimedia Commons. Contextual Bay Area view; it does not show a browser-agent session or StackAI UI. Photo: Frank Schulenburg / Wikimedia Commons. CC BY-SA 4.0 · Cropped and resized.

How did caching cut the bill?

A browser agent resends tools, system prompt, page text, and screenshots on every model call. Prompt caching only reuses the longest unchanged prefix. Asana says the agent already cached tools and instructions but not the growing history, and that trimming screenshots and text on every step broke the prefix anyway.

Three changes were selected for testing: cache the browsing history with a marker on the latest tool result; prune screenshots in batches (keep up to 20, then cut back to 1) instead of every turn; raise the history budget from 120,000 to 480,000 characters so older text is not rewritten each step.

That is the same prefix-stability idea behind other agent context work; for a related compaction angle see context compaction for agents.

The core trick is structural: keep a stable prefix so prompt caching discounts apply, and stop appending full screenshot histories that bust the cache every turn. Batch pruning and history budgets are workload-specific; copy the principle, not the exact token counts, unless your agent’s observation format matches StackAI’s browser tool.

Rows of server racks with blue network cables in a brightly lit data hall. No people appear.
Wikimedia Foundation server room, photographed by Victorgrigas. CC BY-SA 3.0 via Wikimedia Commons (File:Wikimedia Foundation Servers-8055 13.jpg). Contextual compute hall; it is not Asana’s or OpenAI’s infrastructure and does not depict prompt-cache hardware. Photo: Victorgrigas / Wikimedia Commons. CC BY-SA 3.0 · Cropped and resized.

Which numbers are Asana’s?

The study ran six caching/history policies at two budgets across GPT-6.1 Sol and three anonymized frontier models (A, B, C), three runs each—144 runs—plus a short follow-up. Costs come from provider token counters; answers were scored against an independently prepared reference.

On Model B, the best condition cut cost about 29× and ran about 4× faster than the original production setup. On GPT-6.1 Sol, Asana reports 76× lower cost and 5× faster runs, about $0.47 and four minutes per run, with 89% of input read from cache. OpenAI’s post quotes the same $0.47 and four-minute averages.

The task collected six fields for each of 32 books from a public demo catalog—192 facts. Asana says every best-condition run encountered all 192 facts. Models A–C remain unnamed competitors in both write-ups.

The 76× cost and 5× latency lines compare optimized GPT-6.1 Sol runs to Asana’s original Model B production setup inside a 144-run matrix. They are not a promise that your agent will see 76× on GPT-6.1 Sol tomorrow. Cache discounts, screenshot resolution, and step counts dominate the arithmetic.

What shipped in StackAI, and what did Codex do?

Asana says the browser-navigation changes are released in StackAI and that similar experiments will feed platform evaluations so teams can compare cost, runtime, and answer quality when configuring agents.

OpenAI’s story stresses GPT-6 Astra in Codex as the experimenter that refactored the harness for parallel workflows and logged results into Asana Command. That is a customer narrative about Codex desktop workflows, not a claim we verified by opening Command.

Hidalgo is also quoted as using Astra to test product features before release. That forward-looking line is Asana’s; we did not observe those QA runs.

What limits should teams assume?

The 76× figure compares an optimized GPT-6.1 Sol workflow to Asana’s original Model B production setup, not to every browser agent on the market. Cache pricing, screenshot policy, and history budgets are workload-specific.

Models A–C are anonymized. You cannot map the competitor rows onto a named API SKU from these posts alone.

Prompt caching only helps when the prefix stops changing. Teams that rewrite history every turn will not see the same hit rates.

If you lack prompt caching on your provider, the same history rewrite still helps context size but will not reproduce the dollar curve. Pair cache hygiene with the isolation patterns in our AI agent sandboxing explainer so a cheaper browser agent is not also a wider blast radius.

What did we not test?

We did not run StackAI Browser Navigation, open Codex against Asana’s codebase, or replay the 144-run matrix. This article reports the Asana, StackAI, and OpenAI posts we opened on 10 October 2026.

Common questions

Is the 76× figure measured by AiLookout?

No. It is Asana’s reported comparison of its optimized GPT-6.1 Sol workflow to its original Model B production setup in a 144-run study.

Do you need GPT-6 Astra to apply the caching fixes?

The product changes are history caching and batch screenshot pruning. Astra in Codex is how Asana says it found and tested those changes; the posts do not require every reader to use Astra.

Are Models A, B, and C named?

No. Both Asana and OpenAI keep those frontier models anonymized.

THE TAKEAWAY

What to remember

Treat the 76× claim as Asana’s 9 October StackAI study on GPT-6.1 Sol after cache-friendly history edits—steal the prefix discipline, remeasure on your workload.

Sources & further reading

  1. How we cut a browser agent's cost 76x and made it 5x faster by keeping its cache intact ↗
  2. How StackAI by Asana used Astra, Codex and Command to Reduce Browser Agent Costs ↗
  3. Asana cuts model costs 76x in browser tests with GPT-6.1 Sol ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories