Microsoft introduced Microsoft-Decision-1 — a narrow decision model built as a post-train on top of Qwen3.5-9B. It does not generate text: in a single pass, the model outputs calibrated probabilities for a fixed set of options, and the software only needs to make a decision. On 36 benchmarks hidden from training, covering almost 150,000 questions, it showed the best accuracy among LLMs and other decision models, and according to the stated latency, was 35 times faster than the broad generative model GPT-6 Sol. The model is already available in Microsoft Foundry, and the page with the price is open in the OpenRouter catalog.

image

What happened

The announcement was published in the Microsoft Command Line blog along with a model card in the catalog and an open API. Microsoft-Decision-1 solves four classes of tasks: answering according to a yes/no scheme, choosing one item from a list, rating, and rubric-based evaluation of responses from language models and agent actions; instead of text, the software receives probabilities for each option. By P50 latency, the model is 4.5 times faster than the nearest competitor Quyet-1.0-Large. The stability of the decision was checked by perturbations of the same request: with eight types of distortions, the choice changes on average only in 1.3% of cases, and when paraphrasing the descriptions of options, their rearrangement, and shuffling, the decision did not change even once. According to internal Microsoft measurements, the Xbox Research team labeled more than 10,000 reviews from Steam and X surveys with the model at a quality level of GPT-6 Sol about 200 times cheaper, and Copilot accelerates such decisions 100 times at a quality level of GPT-5.6 Luna. The cost of a call is $0.042 per million input tokens, output tokens are free.

Context

The idea of evaluating and classifying others' responses with language models — LLM-as-judge — has been known for a long time: broad models have been used for years for labeling, scoring, and moderation. The novelty of Microsoft-Decision-1 is not scientific, but product: the output space is strictly fixed, so the model gives not text that needs to be parsed, but probabilities suitable for direct use in code. Microsoft itself marks decision models as a category by direct comparisons with competing solutions and broad generative models, building a separate direction within Foundry. The key promise of the product is calibration: if the stated probability corresponds to the actual frequency of successes, the service can set a confidence threshold and automatically process disputed cases, not spending expensive judgment of a large model where a cheap choice from ready-made options is enough.

Why this is important for the industry

For the industry, the event fixes decision models as a separate class of products: the API immediately gives usable probabilities, not text for subsequent analysis. The economics of the change are specific: agent pipelines perform dozens of sequential decisions — routing, classification, prioritization, moderation, control of agent actions, and each call to a broad LLM adds seconds of latency and fractions of the budget. A narrow decision model solves the same tasks orders of magnitude faster and cheaper, which collapses the cost of such pipelines and puts pressure on the margin of products whose profit relies on expensive LLM calls for simple micro-decisions. Microsoft promises to rebase the model on other backbones, including MAI and OpenAI, and connect it to GitHub Copilot for choosing between local and cloud processing — this will make the pattern visible to the mass developer. If the statements are confirmed, it is worth expecting the appearance of analogs from other providers and the consolidation of the decision layer as a standard node of agent frameworks and SDKs.

Why this is important for users

If you are assembling bots, labeling pipelines, or agent systems, the task of "choose one of N" can now be given to a narrow model for a fraction of a cent instead of calling a large LLM. From the calibrated probability in the response, a "judge" of agent responses is built: it shows when a decision can be trusted without checking, and when a person is needed — this is convenient for moderation, evaluation of AI responses, and choice of agent actions, where the flow of decisions is too large for manual control. Practical limitation: the win works only where the output can be reduced to a fixed set of options; the model does not replace free generation. A reasonable first step is to run the model on your own sample of 100–1000 examples against the current baseline, the same broad LLM, and compare accuracy, cost, and latency on your data. A pilot with an open API is set up in days, not quarters.

What is still unknown / limitations

All key quantitative indicators — best accuracy on hidden benchmarks, wins in latency and cost, labeling quality in production — are measurements by Microsoft itself, there is no independent check yet, so the results should be considered a vendor statement. The comparison of a narrow 9B classifier with a broad generative model is methodologically non-homogeneous: these systems have different computation budgets and different types of outputs, so the correct figure in its category is a comparison with the direct competitor Quyet-1.0-Large. Calibration is the central promise of the product, but in the presented materials there are no metrics like ECE, Brier, or coverage and error curves, so the main selling argument is not yet confirmed by data. The stated robustness is also under-described: the number of trials for each type of perturbation and the criterion by which the decision is considered changed are not disclosed. Finally, in the text of Microsoft itself it says "soon on OpenRouter", while the page with the price is already open in the catalog — this discrepancy should be clarified with the vendor.

Sources

Author

Look at AI, editorial team