Announced 9 Oct 2026 · Sources checked
What did Microsoft publish on 9 October?
The dated object is the Command Line post. It is stamped 2026.10.09 and signed Achint Srivastava. The lede says decision models “are quickly emerging as an important new category” and that Microsoft-Decision-1 is “our new model for fast decision-scoring, available in Microsoft Foundry and coming soon through OpenRouter.” Named jobs are routing, classification, prioritization, verification and workflow control.
The catalog URL printed in the post is https://ai.azure.com/catalog/models/Microsoft-Decision-1. The page we opened labels the publisher Microsoft, type “Text classification, Zero shot classification,” lifecycle “Generally available (GA),” input type text, output type json, context window 32768. A “Direct from Azure models” block describes single-license Azure purchase and PTU portability. That is catalog packaging, not a bench.
This is a hosted decision API, not an open-weight drop. For Cloudflare’s same-day Apache-2.0 omni head, see Clef-omni. For the 8 October hosted diffusion scorer that already prices itself against Jev, see Celeris-1-decision. For OpenAI’s public-beta typed path, see the Decisions API.

How does Decision-1 work, and what does the catalog forbid?
Srivastava writes that Microsoft “post trained Qwen3.5-9B for fast, single-pass decision scoring and will soon rebase it on other models, including Microsoft AI (MAI) and OpenAI.” Given a fixed option set, the model “provides a calibrated probability score for each option.” Supported shapes are yes/no, multiple-choice, rating, and “rubric-based grading of AI responses and agent actions.”
The catalog repeats the Qwen3.5-9B base and adds a training sentence: “publicly available datasets subject to Microsoft's Open Data process, plus synthetic data created by the team.” Capabilities listed there include groundedness evaluation against evidence in the input, safety flagging at application-defined thresholds, and an explicit abstention option such as “cannot tell.” Single-pass scoring is “over inputs of up to 32K tokens.”
Out of scope is printed in full. The model “is not designed for text generation, open-ended question answering, conversation, translation, or summarization.” It “is not designed or evaluated for use as the sole automated decision-maker in consequential decisions about people” and “should not be used as the sole basis for decisions involving credit, employment, housing, insurance, education, healthcare, legal rights, or similarly consequential domains.” It is text-only and “does not generate explanations or rationales.” The integrating application owns options, thresholds, escalation and oversight.
A 9B post-train is not a new mixture-of-experts recipe. For what a MoE spends compute on, see mixture of experts explained. For why a typed answer is a different object from generated text, see structured AI output.

Which scores are Microsoft’s, and who is in the appendix?
The post’s headline comparison is a “36-benchmark comparison, spanning nearly 150,000 questions across benchmarks kept blind from training.” Microsoft-Decision-1 “achieved the highest accuracy” on that set. Latency: “4.5 times quicker than Quyet-1.0-Large, the runner-up, and 35 times quicker than GPT-6 Sol.” A later line restates P50 latency as “~35x faster than GPT-6 Sol.” Those are one vendor’s harness.
Robustness is a flip test. Microsoft “perturb[s] the same request in eight ways” and reports that Decision-1 “changes its decision on 1.3% of perturbations on average with zero flips when option descriptions are paraphrased or when options are reversed or shuffled.” Safety is “5,250 requests across 11 benchmarks” covering harmful content, jailbreaks and prompt injection; the model “successfully refused harmful behavior while retaining a high degree of utility.” No per-benchmark table is in the prose we opened.
The appendix names the models Microsoft says it benchmarked: Quyet-1.0-Large, Surogate Rune 26B-A4B, GPT-6 Luna Decisions, deck-31B, H2O-Lightning-4B and Strands-Decider 2B, each with a Hub, GitHub or docs link. Jev is not in that appendix list. The post separately says Microsoft “took several of the top public models on the popular open leaderboard JevBench” and then tested them on 36 additional public and private benchmarks. We did not open those extra sets.
| Field | What Microsoft prints | What that is not |
|---|---|---|
| Access | Foundry now; OpenRouter “coming soon” | A date for OpenRouter |
| Price | $0.042 / M input; output free | A PTU quote |
| Context | 32,768 tokens on the catalog | Jev’s 64k request window |
| Base | Post-trained Qwen3.5-9B; MAI/OpenAI rebase planned | Open weights for Decision-1 |
What do the internal pilots actually claim?
Xbox Research, the post says, used Decision-1 on “more than 10,000 open-ended pieces of feedback and reviews from surveys, STEAM, and Twitter/X” sorted into a fixed theme set. The finding: “competitive on quality with GPT-6 Sol while running over 14 times faster and 200 times less expensive.” The Copilot team’s quality-control note is “competitive with GPT5.6 Luna and 100 times faster.” Incident-response retrieval “performed better and faster than an LLM.” Microsoft Discovery’s adaptive replanning “scored as 46 times more consistent than the LLM-based score at three times the speed and resulted in nearly four times the speed on adaptive replanning.” Those are internal Microsoft notes. We did not see the labels, the bills, or the Discovery loop.
A use-case list then names agent controls, model routing, skill-based decisions, data labeling, AI judging, intent analysis, incident routing, validation, recommendations, search relevance, content filtering, code scanning, safety screening, computer use, robotics and scientific discovery. That list is a menu, not a customer appendix.
How does $0.042 sit next to Jev, Clef and Celeris?
TypeSafe’s Models page, opened the same evening as the TypeSafe $870 million Series A note, still lists Jev 1.13.0 at $0.042 per million input tokens with free output. Microsoft printed the same list price. That is a price match on two vendor pages, not a claim that the models agree. Jev’s documented request window is 64k; Foundry’s Decision-1 window is 32k. Jev’s docs remain text-only; so does Decision-1. Clef-omni, by contrast, takes embedded audio and video and lists $0.15 per million on Workers AI.
Celeris’s 8 October page priced `celeris-1-decision` at $0.04 per million and put Jev 1.13.0 in a vendor jev-bench table. Microsoft’s appendix does not reprint that table. A typed gate can still fail open; see the option-channel preprint.
What should a team try before swapping a router?
The useful trial is a set of questions you already score: a ticket, a policy clause, an agent tool-call proposal, and one case that should abstain. Compare Decision-1’s probabilities with Jev 1.13, Clef, or your current LLM JSON prompt at the same schema. The 35× GPT-6 Sol line is not a production SLA, and the catalog’s people-decision ban is on the same product page as the GA badge.
Pin the Foundry deployment. OpenRouter is still “coming soon” with no date. The planned MAI and OpenAI rebases are future work. We did not create an Azure resource, send 32K tokens, or rerun JevBench.
This is an evidence review of the 9 October Command Line post and the Foundry catalog page. We did not watch the backpack or classification demos as timed measurements.

Common questions
Are Microsoft-Decision-1 weights on Hugging Face?
Not on the pages we opened. The catalog describes a hosted Foundry model post-trained from Qwen3.5-9B. The Command Line post does not announce an open-weight Decision-1 repo.
Is OpenRouter already serving Decision-1?
Microsoft’s post says “coming soon through OpenRouter” and gives no date. The live path it names is the Foundry catalog URL.
Did Microsoft say Decision-1 beat Jev?
Not in those words. Jev is not in the appendix list. The post says Microsoft took top JevBench public models and then scored them on 36 other benchmarks, where Decision-1 “performed the best.” That is Microsoft’s comparison, not JevBench’s live board.
What to remember
Use the 9 October Command Line post for Foundry availability, $0.042, and the vendor latency rows. Use the catalog for the 32K window, the Qwen3.5-9B base, and the ban on sole use in consequential people decisions. Keep OpenRouter and the MAI/OpenAI rebases in the planned column.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





