The Model That Answers With a Number
A trillion-dollar American lab just shipped a flagship capability built on an open model from Alibaba. It refuses to talk.
The lab is Microsoft, the base is Alibaba’s open Qwen3.5-9B, and the thing itself is Microsoft-Decision-1, out on Thursday. Its one rule: it will not write you a sentence. Give it a situation and a short list of options and it returns a probability for each, a calibrated score your software can act on at once.
The real news is the speed. The category is three weeks old. Jev, a startup, opened it in mid-September. OpenAI and Cloudflare answered within days. Now the largest company in the field has a version at four cents per million tokens, and a clever idea is a commodity before anyone has checked it works.
The primary record
Achint SrivastavaVP of Software Engineering, Office of the CTO, MicrosoftCommand Line · 9 Oct 2026
“To build Microsoft-Decision-1, we post trained Qwen3.5-9B for fast, single-pass decision scoring and will soon rebase it on other models, including Microsoft AI (MAI) and OpenAI. When given a fixed set of answer options, Microsoft-Decision-1 provides a calibrated probability score for each option.”
Reproduced from Microsoft’s launch post, not a live embed. Listed in Microsoft Foundry and OpenRouter.
Decision-1 was post-trained for what Microsoft calls single-pass decision scoring. It handles yes or no, multiple choice and rating scales, and grades an agent’s next move against a rubric. The word to slow down on is calibrated: a probability that means something is different from a number that only ranks.
Interactive · cast a decision
A ticket: “I have been charged twice.” Four moves. Pick one.
Microsoft-Decision-1 · probability per option
Pick one and it shows its working.
You chose A. The model leans to B.
You chose B, the model’s pick at 0.71.
You chose C. Only 0.20 on asking.
You chose D. Just 0.03 on escalating.
Illustrative, not live output.

Fast, and vendor-measured
Every decision costs time: a tenth of a second on twenty decisions is two seconds. In Microsoft’s tests Decision-1 was quickest, about two and a half times faster than the runner-up and thirty-five times faster than GPT-6 Sol at the median.
Read it as a claim
Every number here is Microsoft’s, on benchmarks Microsoft chose. Clef was absent from the table. Read the chart as a claim, not a measurement.
The independent number
Outside evidence exists, but not for this model. Two vendor-free benchmarks cover the category. DMB, a reproducible harness, ran 256 decisions per contender: GPT-6 Sol scored 85.7%, and Jev, the category’s originator, 80.5% on banking and 98% on spam. sysone-bench put Jev at 0.9065 and the best open model at 0.8556 across 1,240 decisions. DMB adds the line that matters most: Jev’s confidence is a “provider-defined score, not a calibrated probability of correctness”.
So the category’s promise is unproven for its flagship, and Microsoft’s is untested. I used both benchmarks; I did not run Decision-1, so this stays a record, not a test.
Who actually loses
Jev invented this category and watched a trillion-dollar lab repackage it for pennies three weeks later. Anyone who raised money to sell a decision model just learned their moat was one good idea wide. And the buyer wiring a vendor’s 83.5% into production is trusting a number only the seller checked. Cheap is not the same as verified.
What to do on Monday
Decision-1 is cheap and fast enough to wrap into a routing or agent-control step today. The move is not to trust the 83.5% and not to dismiss it: run your own calibration on your own cases, then let the confidence score decide what it may decide alone. Want a second pair of eyes? Email brandon@kreostudio.co.uk.
How this was made, and what is confirmed
- Confirmed, opened directly: Microsoft's launch post, 9 October 2026. Decision-1 is post-trained from Qwen3.5-9B and returns a calibrated probability per option, at $0.042 per million input tokens.
- Reported, not confirmed: the 83.5% and 85 ms (The Decoder, reading Microsoft's chart); Microsoft chose the benchmarks. Independent, for the category: DMB (GPT-6 Sol 85.7%; Jev 80.5% and 98%, its confidence 'provider-defined') and sysone-bench (Jev 0.9065; best open 0.8556).
- Ran: nothing; no account, so the scorecard is illustrative and the hero hand-drawn. What would change my mind: an independent reproduction of the 83.5% and the 85 ms.
Sources & references
- Achint Srivastava, Introducing Microsoft-Decision-1, our model for fast decision-making, Microsoft Command Line, 9 October 2026 (primary).
- microsoft-decision-1, Microsoft Foundry model catalog (primary); OpenRouter (primary).
- Microsoft’s Decision-1 model enters the fast-growing AI decision model race, The Decoder, 10 October 2026 (the 83.5% and 85 ms figures, Jev, and the Clef omission).
- DMB, decision-model benchmark (independent, no vendor sponsorship), github.com/nibzard/decision-model-benchmark (pilot of 26 September 2026; the “provider-defined score” line).
- sysone-bench (independent head-to-head on byte-identical inputs), github.com/instax-dutta/sysone-bench (Jev 0.9065; best open model 0.8556; 1,240 decisions).
Reader signal
Was this useful?
Work with KREO Studio
AI engineering, data science and design architecture, from Plymouth to the wider UK.
Next Article
