Announced 9 Oct 2026 · Sources checked
What did Microsoft publish on 9 October?
The dated announcement is Waldek Mastykarz’s post on Microsoft for Developers, stamped October 9th, 2026, titled “Introducing the Agent Experience (AX) Practitioner Playbook.” It defines AX as how well coding agents discover and correctly use a technology, and it argues that waiting for models to improve is not a strategy when teams can change docs, MCP tools, skills, plugins, instructions, CLIs, and APIs.
The downloadable objects resolve through Microsoft short links. On 11 October, aka.ms/ax-playbook returned HTTP 301 to https://github.com/microsoft/scope/blob/main/docs/ax-playbook/ax-playbook.pdf. aka.ms/ax-playbook/skill redirected to the ax-practitioner directory under the same docs tree. The PDF we fetched is 423,147 bytes with SHA-256 0e8bbce63a05425ab2fe711696b94275d05440f32254d8b86adc78f7e6514ce7. GitHub’s API lists microsoft/scope as MIT-licensed with description “Scope | Agentic Experience evaluation platform (Research preview).”
This playbook is about agent enablement surfaces, not Microsoft’s typed decision scorer. For Decision-1 in Foundry, see Microsoft Decision-1. For harnessed agentic RL tooling from the same vendor family, see Agent Lightning v1.0.

What problem is the playbook trying to solve?
The post’s opening example is familiar: ask a coding agent to build something with your SDK, watch it produce code that looks plausible, then find the wrong version, a deprecated auth pattern, or a setup no one on the product team would recommend. The agent followed its training data and whatever tools loaded. The claim is that owners of documentation and extensions can change those sources now, if their evaluations can prove which change helped.
Microsoft says DevRel has measured agents against Azure, Cosmos DB, SharePoint Framework (SPFx), and Microsoft 365 Copilot extensions since fall 2025 using prompts like those real developers write. Earlier pieces in an Agent Experience series are presented as parts; the playbook is the end-to-end method. Those measurement programs are Microsoft’s own; the post does not publish a public leaderboard of third-party SDKs.
Who it is for, per the post: people who build SDKs, APIs, services, CLIs, MCP servers, skills, plugins, instruction sets, or the docs behind them—and advocates who can run scenarios even when they do not own the surface. The method is described as evaluation-system agnostic as long as required capabilities are present.
How does the method diagnose agent mistakes?
The announcement highlights a failure split that other write-ups quoting the PDF also emphasize: an extension that never loaded (discovery), one that loaded but was never called (invocation), and one that was called but applied wrong (application). Each points at a different fix. The post also warns that evaluations can lie—perfect scores for code that never compiled, or a “used Platform X” check that passes whether or not Platform X was used—so criteria need meaning gates and run gates before a score is trusted.
The AX Practitioner skill’s SKILL.md, which we opened from microsoft/scope on 11 October, operationalizes that structure for interactive use. It instructs the skill to answer strictly from per-chapter playbook fragments, cite chapter and section, and stop with a labeled “not covered” block when the PDF does not settle the question. Diagnose mode is told to match misbehavior to Chapter 9’s discovery/invocation/application split and Chapter 10’s failure patterns, then point at the “From failure to surface” table and Chapter 11 fix guidance.
Numbers that appear in the skill’s own guidance—examples such as “at least 5 runs” or SPFx score transitions cited in the skill text—are playbook figures the skill is allowed to reproduce, not independent benchmarks. The blog’s “across 330 questions, its answers scored 95% on average against what the playbook says” measures skill fidelity to the PDF, not Cosmos DB or SPFx product quality.

What outcomes does Microsoft claim so far?
The post says evaluations already led to dozens of shipped fixes, “including 46 improvements to the Azure Cosmos DB Agent Kit,” and that an SPFx project-upgrade scenario runs end to end in the playbook from scenario through shipped fixes. Those counts and the worked example are Microsoft’s reporting about Microsoft surfaces. No outside team’s AX run is cited in the post we opened.
The skill is positioned as a way to learn the method while applying it: install it in a coding agent, ask how to word a criterion or why an extension was ignored, and optionally review scenarios against the playbook’s standards. SKILL.md repeats that outside content must stay in labeled blocks only after the user opts in. That design is about preventing the skill from inventing playbook text; it is not a guarantee that every agent host will load the skill correctly—an irony the discovery/invocation split would flag.
If your immediate problem is containing what a coding agent can touch on a laptop rather than fixing docs the agent reads, compare GitHub Copilot’s local sandboxing GA and AWS Strands Box as different layers.
| Claim | Source | What that is not |
|---|---|---|
| 46 Cosmos DB Agent Kit improvements | DevRel blog post | A third-party kit audit |
| 95% skill answers on 330 questions | DevRel blog post | An SDK quality score |
| PDF 423,147 bytes at sha256 0e8bbce6… | Fetched 11 October | A page count we OCR’d here |
| MIT license; research preview Scope repo | GitHub API for microsoft/scope | A GA Azure product SKU |
| SPFx upgrade worked example in PDF | Blog + skill Appendix B pointer | Your product’s eval result |
How do you get the PDF and the skill?
Start with the blog post for the framing, then fetch the PDF via aka.ms/ax-playbook or the raw GitHub blob under docs/ax-playbook/ax-playbook.pdf. For interactive guidance, follow aka.ms/ax-playbook/skill into docs/ax-playbook/ax-practitioner and install that skill in a coding agent that supports the format. The repository also hosts ax-practitioner-skill-report.html beside the PDF; we did not treat that HTML report as a separate announcement.
Scope’s GitHub description still says research preview. Placeholder links inside the playbook (the skill warns about TODO- targets) may not resolve yet. If you need a working link the playbook marks unfinished, the skill’s rule is to say so rather than invent a URL.
Practical first step for a platform team: pick one realistic developer prompt against your MCP server or SDK docs, decide in advance what “correct” means, and only then score trajectories with the discovery/invocation/application lens. The playbook’s value is that checklist discipline; the 46 and 95% figures are Microsoft’s own scoreboard.

Common questions
Is the AX Practitioner Playbook the same as Microsoft Scope?
The playbook and skill are files inside the microsoft/scope GitHub repository, which GitHub describes as an Agentic Experience evaluation platform in research preview. The 9 October post is the announcement of the method PDF and skill; it is not a claim that Scope is generally available as an Azure SKU.
Does the 95% figure mean agents use Cosmos DB correctly 95% of the time?
No. The blog says the skill’s answers scored 95% on average against what the playbook says across 330 questions. That is fidelity to the PDF, not a product success rate.
Did Ai Lookout run an AX evaluation?
No. We opened the blog post, followed the aka.ms redirects, fetched the PDF and SKILL.md, and recorded hashes. We did not score an agent against an SDK.
What to remember
Microsoft’s 9 October AX Practitioner Playbook is a downloadable eval method—plus a playbook-grounded skill—for fixing how coding agents discover and apply your docs and tools. Use the PDF’s discovery/invocation/application split; keep the 46-fix and 95% lines as Microsoft’s.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





