What did Postman and AWS publish on 9 October?

The dated object is the AWS Machine Learning Blog entry stamped 09 OCT 2026: “How Postman runs Agent Mode for 40 million developers on Amazon Bedrock,” by Srinivas Kini and Shubham Gupta. The opening contrast is explicit: building an agent for a demo and operating one for tens of millions of developers are different engineering problems.

Agent Mode is described as Postman’s portal for AI-native work across testing, documentation, discovery, and implementation after eleven years of interface-driven product habits. The authors say they expected model quality and prompt design to dominate, but the harder work was making a mature product legible to an agent—controlling tool sprawl, exposing schema-based reads, and treating context rather than raw capability as the primary bottleneck.

Bedrock continues to add agent-facing APIs this week. For reasoning summaries on OpenAI models in Bedrock’s Responses path, see Bedrock OpenAI reasoning summaries. For Anthropic’s hosted agent workflow packaging, see Managed Agents dynamic workflows.

Which product patterns does the post emphasize?

Tool sprawl: exposing every action as a separate tool scaled poorly. Postman reports that tool-selection errors rose as the visible toolset grew. The preferred pattern is dynamic tool selection per task, plus schema-aware reads through a query engine instead of dozens of single-purpose read tools. Actions should not require an open UI tab—that keeps the agent navigating screens instead of reasoning over data.

Context engineering: missing or incomplete context caused more failures than missing capabilities. The post distinguishes broad, shallow background context (gathered automatically and minified) from deep, user-selected context routed through entity-type handlers that distill what the agent needs. Serializing the rendering data model was not enough; handlers exist because UI objects were shaped for display, not reasoning. Truncation of open-ended user fields (descriptions, OpenAPI specs, payloads) is treated as a first-class budget problem, with a filesystem-backed approach mentioned as exploratory work.

Three runtime components aggregate before each foundation-model call: client-side tools in the Postman app (open requests, modify settings, run collections, inspect auth) plus some server-side tools; generic agent instructions for tone, uncertainty, and baseline product knowledge; and a RAG knowledge base seeded from Postman’s Learning Center so feature articles inject when relevant (for example, selecting a mock server pulls the related article).

Bedrock capabilities named for Agent Mode in the 9 October post
CapabilityWhat the post describesCaveat printed
Model flexibilityRoute workloads across supported Anthropic Claude models via Bedrock APIsSupported models vary by Region
Cross-Region inferenceGeographic or global inference profiles as modelId in Converse/InvokeModelIAM, SCPs, and quotas must allow every destination Region
Data retentionZero retention with data_retention_mode none for supported modelsAvailability and behavior are model-dependent
Prompt caching1h checkpoint for stable core; 5m for variable contextLonger-lived checkpoint must appear before shorter; model-dependent TTLs

How does Agent Mode use Amazon Bedrock?

Inference is framed as a routing-and-caching problem, not only model selection. Postman can point high-volume, latency-sensitive turns at faster Claude models and reserve larger models for complex reasoning. The authors say newer and larger models reduced tool hallucinations in Postman’s testing—still a vendor-reported observation, not our A/B test.

Cross-Region inference uses Bedrock inference profiles so bursty developer traffic can spread across destination Regions. Geographic profiles keep routing inside a defined geography (for example United States or European Union). Global profiles maximize throughput when the workload does not require a geographic boundary. Enterprise buyers are reminded that geographic profiles constrain Bedrock’s eligible Regions; they do not mean inference runs inside Postman’s own AWS account.

On retention, the post states AWS does not use prompts and completions to train AWS models or distribute them to third parties, and that Postman configured zero data retention for supported Agent Mode models. Each production model still needs a check against current Bedrock data-protection docs. For agent packaging that travels with sandbox policy, compare AWS Strands Box. For shipping API surface area as MCP skills, see Cloudflare’s API MCP skills changelog.

Rows of blue server cabinets in a large data-center aisle. No people appear.
Data-center aisle photographed by Connie Zhou for Google, CC BY-SA 4.0 via Wikimedia Commons (File:Google_data_center.jpg). No people appear. Archival data-center photo; not AWS Regions used by Postman and not a Bedrock control plane. Photo: Lambtron. CC BY-SA 4.0 · Cropped and resized.

What should builders take from the Bedrock caching section?

A production agent resends stable prefixes every turn: system prompt, agent instructions, core tools, selected knowledge, and conversation state. Bedrock prompt caching lets Agent Mode mark a one-hour checkpoint on near-immutable core blocks and a five-minute checkpoint on more variable context. The post notes Bedrock requires the longer-lived checkpoint to appear before the shorter-lived one, and that teams should verify `cacheReadInputTokens` / `cacheWriteInputTokens` plus time-to-first-token on their own models.

Best practices distilled in the article: budget tools as carefully as tokens; prefer schema-aware reads; decouple actions from UI state; engineer context handlers instead of dumping render models; manage the context window as a scarce resource; ship docs with features so RAG stays current; and match Claude model, inference profile, and cache TTLs per workload.

The conclusion states Postman’s production implementation is proprietary and is not available as a public sample repository. Readers looking for clone-and-run code will not find it in this post; the artifact is the architecture narrative plus pointers to Bedrock docs and Postman’s Agent Mode product documentation.

Network switches with patch cables in a rack, photographed close up. No people appear.
Network switches photographed by Håkan Johansson, CC BY-SA 3.0 via Wikimedia Commons (File:Network_switches.jpg). No people appear. Contextual networking photo; not Postman’s production fabric. Photo: ShakataGaNai. CC BY-SA 3.0 · Cropped and resized.

What is not established by this case study?

The “40 million developers” figure is Postman/AWS’s scale framing for the community Agent Mode serves; we did not audit active seats. Latency, cost, and hallucination improvements attributed to larger Claude models or caching are described qualitatively or as Postman’s internal testing—not as a public benchmark table with prompts we can rerun.

Zero data retention is model-dependent. Cross-Region geographic boundaries are profile-dependent. Neither replaces a customer’s own compliance review. Schema-based reads and dynamic tools are design patterns; they do not guarantee correct API mutations on every collection or auth type Postman supports.

This article reviews the AWS Machine Learning Blog HTML we opened on 10 October 2026. We did not authenticate to Postman, invoke Bedrock Converse, or measure cache hit rates.

Close-up of an Ethernet switch with many RJ45 ports and status LEDs. No people appear.
Ethernet switch photograph released as CC0 via Wikimedia Commons (File:EthernetSwitch.jpg). No people appear. Contextual networking photo; not AWS Bedrock hardware. Photo: Raysonho @ Open Grid Scheduler / Grid Engine. CC0 · Cropped and resized.

Common questions

Is Postman’s Agent Mode code open source?

No. The 9 October post states the production implementation is proprietary and not available as a public sample repository. The write-up shares patterns, not a cloneable stack.

Which models does Agent Mode use on Bedrock?

The post says Agent Mode can access supported Anthropic Claude models through Amazon Bedrock model inference APIs and route workloads by latency versus quality. Exact model IDs and Regions must be checked in current Bedrock documentation.

What caching TTLs does the post describe?

A one-hour cache checkpoint for the near-immutable core (system prompt, agent instructions, core tools) and a five-minute checkpoint for more variable context, with the longer-lived checkpoint ordered first.

THE TAKEAWAY

What to remember

The 9 October AWS ML Blog case study is useful as a Bedrock production checklist—tool budgets, context handlers, Claude routing, geographic inference, retention, and tiered prompt caching—while Postman’s Agent Mode itself remains closed source.

Sources & further reading

  1. How Postman runs Agent Mode for 40 million developers on Amazon Bedrock ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories