The open-source inference engine SGLang has introduced /v1/decisions and /v1/systemone endpoints: a standard chat model responds in them not with text, but with probabilities — a choice from options, a rating on a scale, or a “yes/no” answer. The changes were accepted into the repository on September 24 and 25, 2026; until the next stable release, the feature is available in nightly builds. In practice, any server already deployed on SGLang can be used as a probabilistic classifier without a separate classification model, although the values in the responses are not yet calibrated, so decision thresholds will have to be tuned on your own labeled sample.

What happened
The PR #40992 change with the /v1/decisions endpoint was created in the SGLang repository on September 24, 2026. The next day, on September 25, PR #41208 (commit 16d1c93b) with the /v1/systemone route, compatible with the TypeSafe API, was merged into the main branch; until the next release, both features are available only in nightly builds. The mechanics are as follows: the engine renders the question through the model's chat template with the “thinking” block disabled, assigns single-token labels, and computes a softmax based on the log-probabilities of the next token, taking temperature into account. The set of labels depends on the task format: letters A–Z for choosing from 2–26 options, two-letter AA/AB and so on up to 255 options in System One, digits 0–9 for a scale of 2–10 levels, and the yes/no pair for a “yes/no” question. Along with the distribution, the response stores label_mass — the share of probability mass that fell on the labels. The performance of the combination was demonstrated on a live example: in a video published on September 30, 2026, the Qwen3.8-27B model plays through Pokémon FireRed, making decisions via /v1/decisions in under 100 milliseconds.
Context
The technique itself has been known for a long time: classification by the softmax of the log-probabilities of the next token with single-token labels is a standard method of verbalizer prompting from large language model research. Previously, it had to be assembled manually on top of standard generation: generate a string, parse the answer, ensure the label fits in one token, and suppress the reasoning block. The novelty here is engineering: the probabilistic classification pipeline has for the first time been moved to the inference server level, alongside the existing /v1/score and /v1/rerank. External compatibility is also notable: the /v1/systemone route replicates the API of the System One service by TypeSafe AI with the proprietary Jev model, meaning SGLang positions itself as an open replacement for a closed interface. Finally, the approach is built as stateless classification: a single input is processed without dialogue history, so any state, such as game progress in the Pokémon FireRed demo, must be fully serialized into a single prompt.
Why this matters for the industry
For the industry, this is a change in the very optics: inference engines no longer compete only on generation speed. SGLang takes on the server level a non-trivial part of the work — rendering the prompt through the chat template, guaranteeing the single-token nature of the label, and suppressing the reasoning block — which previously each team assembled on its own. A new class of endpoints emerges, where the model returns not a string, but a probability distribution: classification, routing, scoring, and guardrails based on application state are performed by a synchronous request in milliseconds, without separate classification checkpoints and without parsing free text. Compatibility with the TypeSafe AI System One API turns any SGLang server into an open replacement for the proprietary Jev API, reducing the cost of migrating already-written clients. Since the technique with single-token labels is trivially reproducible, it is logical to expect similar “decision” endpoints from other engines, including vLLM; then competition will shift to standardizing the API format and calibration tools, and LLM classification as a separate paid feature will begin to commoditize.
Why this matters for users
For a reader who already serves models through SGLang, there is no need to deploy anything additional: a nightly build and a request to /v1/decisions with a set of options or a binary question are enough. There is no need to train a separate reward or classification checkpoint for such tasks — the same model that answers in chat works. Ready-made TypeSafe SDK clients (the typesafe-sdk package, installed with the pip install typesafe-sdk command) connect to your own server by replacing base_url, so existing integrations with the TypeSafe API can be migrated without rewriting code. Typical scenarios — ticket triage, routing requests between models, and simple guardrails in your own projects. The recommended implementation order in the documentation: label a sample, fix the temperature, and calibrate thresholds for your tasks.
What is still unknown / limitations
The main methodological nuance is that the probabilities are not calibrated: the values fluctuate by up to 0.07 between a cold request and a request from the prefix cache, and the softmax is passed through temperature, so decision thresholds are valid only on your own labeled sample and with fixed settings. Operability also depends on the tokenizer of a specific model: the label must fit in one token in the answer position, which is why the above limitations on the number of options and scale levels follow. The combination has so far been tested on a pair of models from the Qwen family, and there is no stable release yet, so it is too early for production, while for pilots the feature can already be used. The figure “under 100 milliseconds” is taken from one demonstration video with Qwen3.8-27B and Pokémon FireRed: latency distributions, data on the accuracy of choices, and a comparison with the “generate a string and parse” baseline are not provided. Finally, the PR #41922 change with image input has not yet been merged, so multimodal input is not yet confirmed by code in the main branch.
Sources
- SGLang Docs — Decision models section (/v1/decisions and /v1/systemone)
- SGLang PR #40992 — /v1/decisions endpoint code (created 24.09.2026)
- [SGLang PR #41208 “[feat] add a system one compatible /v1/systemone route” (merged 25.09.2026)](https://github.com/sgl-project/sglang/pull/41208)
- SGLang (@sgl_project) on X — Qwen3.8-27B demo in Pokémon FireRed (30.09.2026)
Author
Look at AI, editorial team
