Cloudflare has released two open decision models under the Apache 2.0 license — Clef (27B) and Clef-flash (9B) — with weights available on Hugging Face and already accessible on the Workers AI cloud platform. Unlike generative LLMs, these models do not write text: in a single pass, they return the probability of each option for typed questions of type choice, bool, and score, based on which an agent selects its next action. The release is notable because a category previously dominated by closed APIs now has an option that can be self-hosted and immediately used in the cloud.



What Happened
Cloudflare has published the weights of two decision models under the Apache 2.0 license on Hugging Face and included both models in the Workers AI platform. The larger Clef model is fine-tuned on Qwen3.8-27B, and the smaller Clef-flash on Qwen3.5-9B. Instead of generating an answer in words, each model evaluates all options from a given schema of choice, bool, and score type questions in a single pass and outputs the probability of each option, which the agent uses to choose its next step. Both models accept text, JSON, images, and video, and maintain a context of 65,536 tokens. Fine-tuning was conducted using the Brier loss function and the RLCD method, aimed at calibrating these probabilities. According to Cloudflare, on the median of 43 tests, Clef responds in 209 ms compared to 524 ms for the closed Jev model, and Clef-flash in 39 ms; in the BANKING77 classification, Clef achieves a macro-F1 of 94.2 compared to 79.7 for Jev, and in ToolRet shows 69.2 compared to 65.3 nDCG@10. Along with the models, Cloudflare announced an RL fine-tuning platform that guides a service through a chain of Worker, AI Gateway, Containers, and Trainer, and then redeploys the updated model.
Context
Decision models address a long-standing problem in applying LLMs where a choice, not text, is needed. A generative model formulates an answer in words, and confidence must be extracted from the token distribution, which poorly aligns with the actual frequency of correct decisions, whereas for agent branching, it is precisely the calibration of probability that matters, not the familiar accuracy. Clef is structured differently: based on the input context, a single pass is performed over the base Qwen model, after which all options from the schema are evaluated in parallel by logits, and no decoding cycle occurs at all. It is precisely the refusal to generate text that explains the stated latency, which is incomparable to the time of a generative answer. Training with the Brier loss function and RLCD specifically targets the alignment of the stated probability with the actual frequency of correct answers. The market context is as follows: similar capabilities were previously provided only by closed APIs, primarily the Jev model from TypeSafe AI, and Clef's compatibility with this API is designed as a market move — competition is entering a category where developers previously had no choice.
Why This Matters for the Industry
The main consequence for the industry is portability: full compatibility with the Jev/SystemOne API means that under an existing integration, it is enough to switch the endpoint without rewriting code, and changing the decision model provider becomes a cheap operation. In a category that operated on a closed API, this changes the very logic: value shifts from access to the model to its calibration and fine-tuning for a specific domain, and this is exactly what the announced RL platform targets, rolling up the service update into a single cycle. The stated latencies allow placing classification, routing, and tool selection in the hot path of an agent and product logic, not just in background pipelines, meaning treating the decision model as a regular fast call within a transaction. Open weights also give teams a tool for independent verification of vendor figures on their own data, which is important in itself for a category where previously one could only take things on faith.
Why This Matters for Users
Practically, the models are available through three paths. The weights of Clef and Clef-flash can be downloaded for free from Hugging Face and deployed on your own GPU — Cloudflare has verified local launch on a single H200 with a stack of torch 2.11 and transformers 5.10.2. You can use the cloud endpoints @cf/cloudflare/clef and @cf/cloudflare/clef-flash in Workers AI, where one million input tokens costs $0.24 for Clef and $0.09 for Clef-flash, with a guarantee that Cloudflare does not read or store requests and does not train on them. A working scenario is simple: a schema of typed choice, bool, or score questions is described in JSON, text, JSON, images, or video are fed as input, and the probabilities of each option are obtained as output without parsing free text. A reasonable first step for a reader is to take one classification step of their product, such as triage of requests or routing, run their own set of questions through the model, and compare the result with what the existing Jev integration currently does.
What Is Still Unknown / Limitations
All key figures, including latency, BANKING77, and ToolRet, are reported by Cloudflare itself and have not yet been confirmed by independent measurements. The published sources do not include calibration metrics on held-out data such as ECE or reliability diagrams, so training with the Brier loss function and RLCD promises calibration but does not prove it, especially out of distribution of the fine-tuning, and it is premature to consider the returned probabilities a ready-made basis for confidence-gating in interfaces. The stated gap of about fourteen macro-F1 points in BANKING77 is unusually large for this task and is more plausibly explained by fine-tuning for the format and close domains than by universal model superiority, and the difference in ToolRet is small and may lie within the range of variation between pipelines. In addition, the independent fine-tuning service on the announced RL platform has not yet been launched: for now, domain fine-tuning is performed only by the Cloudflare FDE team, which limits the mass domain cycle.
Sources
- Introducing Clef: our open-source decision models, and new RL fine-tuning platform — Cloudflare Blog
- Clef — Cloudflare Workers AI Models Documentation
- Cloudflare/clef — Hugging Face (open weights, Apache-2.0)
Author
Look at AI, editorial team
