Liquid AI has released a new class of models — decision models, created not for text generation but for structured decisions. The first model in the family is called d1 and is available for free via API as d1:free: it takes the task state as plain text or JSON and returns probabilities for predefined answer options in a single call, without generating a single output token. The developers position d1 as a replacement for generative LLMs where the answer is known in advance to be chosen from a known set — in classification, routing, and scoring.

image
image
image

What happened

The d1 model is open at the endpoint https://api.liquid.ai/decisions/v1/systemone. According to the API contract, the task state (state) is fed as input, and the response contains only the declared calibrated probabilities for predefined options; at the same time, the usage.output_tokens field is always equal to zero. Three question primitives are supported: Noul — yes or no with a probability from 0 to 1, Choice — selection of one option with a full probability distribution and a confidence field, Score — rating on an ordered scale, calculated as a probability-weighted position. Multiple questions can be mixed and executed in a single request, so three sequential classification calls are merged into one. The response always matches the question type — this guarantee is also fixed in the contract, so there are no schema parsing errors.

Context

Today, classification, request routing, scoring, and moderation are usually performed by generative LLMs: the model spends paid token generation to pronounce a label from a predefined set of options, after which the output needs to be parsed, and in case of a format error — asked again. d1 solves the same task in a discriminative setting: the guarantee of zero output tokens indicates a probabilistic head on top of the model instead of autoregressive generation, although the specific architecture and training method are not disclosed in open sources. In parallel, the market for structured decisions is being formed as a separate segment alongside text LLMs: TypeSafe Jev 1.13, available via OpenRouter, Convai Laya with open weights under the Apache 2.0 license, and AutoTrust JEV-27B work with the same primitives.

Why this matters for the industry

For the industry, d1 is a new API primitive, not just another model: classification, ticket routing, scoring, and moderation are moving from token generation with JSON parsing to a single deterministic call. The cost of a request is reduced to input tokens, latency becomes predictably low, and an entire class of failures in response parsing disappears along with retries, because the response by contract always matches the requested schema. The confidence field provides a ready-made escalation template: disputed cases are passed to a large model or a human, and the rest are resolved by a cheap model. Thus, d1 becomes a direct competitor to LLM-router and LLM-as-judge niches and the pattern of multiple sequential classification calls, and if independent checks confirm the calibration, a separate decision layer with a decision model — LLM — human cascade for intent routing and triage is likely to appear in agent stacks.

Why this matters for users

You can try d1 for free and without preparation: the d1:free model is available with a key from console.liquid.ai, there are SDKs typesafe-sdk for Python and @typesafe-ai/sdk for TypeScript, official documentation, a migration guide, and a road-decider demo on Node.js 18+, where d1 chooses a driving lane two to five times per second. The migration rule from the official materials is simple: if the answer is one of N predefined options, take a decision model; if you need to compose new text, keep the LLM. As a working template, probability-based escalation is proposed: if the value is above 0.8, the action is blocked, below 0.2 — it is passed, and the middle of the spectrum goes to a human for review. A separate lesson from the road-decider demo: the state format is more important than expectations — short summaries for each option give the model more confidence.

What is still unknown / limitations

There are no published calibration metrics in open sources — for example, ECE or Brier score on an external dataset, so trusting the 0.8 and 0.2 thresholds without your own validation on historical examples means transferring the vendor's claim to production without verification. There are also no comparisons with LLMs using structured output or constrained decoding on the same tasks, so the advantage in price, latency, and schema errors has not yet been measured by a third party. The architecture and training method of d1 are not disclosed, making it impossible to assess the real scientific novelty. The vendor's claim about calibration and leadership on Hugging Face's Decision Index has not yet been confirmed by independent measurements; the model is only available via API, and there are no signs of open weights from Liquid AI in the sources.

Sources

Author

Look at AI, editorial team