Cloudflare is offering developers a way to replace some general-purpose LLM calls with bounded decisions: choose a support team, assess urgency, or classify an input, then return typed results and probabilities that application code can act on.
In its October 1, 2026 announcement, the company released two internally trained models, Clef and Clef-flash, on Workers AI. Their weights are also downloadable under Apache 2.0. The models have a narrower role than a conversational assistant: make the decision inside an agent workflow without generating a free-form answer.
The hosted models and downloadable weights are available now. Cloudflare is also offering hands-on fine-tuning through its forward-deployed engineering team. The self-serve reinforcement-learning platform described in the announcement remains a future offering, with no published launch date.
For developers, the immediate question is which existing LLM calls need language generation and which only need a reliable choice.
Replace the Decision Step, Not the Whole Agent
Cloudflare describes Clef as a decision model. Developers provide context and a schema defining the questions and permitted answers; the model returns schema-constrained results with probabilities instead of composing an explanation.
Its support-ticket example asks three questions about a checkout failure: whether the request is urgent, which team should handle it, and how severe the customer impact is. The team field has defined choices such as billing, technical, and sales. Severity uses an ordered scale.
Application code can use the selected team directly, without extracting a department name from a paragraph. It can escalate an urgent ticket or send an uncertain classification to a human.
This bounded interface is especially useful when one agent combines several jobs. A decision model could route an incoming request, while a general-purpose LLM investigates the problem, calls tools, or drafts a response. Replacing the routing call does not require replacing the rest of the agent.
Other plausible applications include choosing between retrieval sources, selecting a tool category, or assessing whether submitted content belongs in a defined policy category. These are workflow designs, not evidence that Clef has been validated for every such task.
Policy classification also needs a boundary. A model’s assessment can inform an enforcement workflow, but hard permissions and mandatory restrictions should remain explicit application rules. A valid schema output guarantees an allowed answer format, not a correct judgment. Confidence thresholds, human review, and workload-specific evaluation still matter.
Clef Ships Hosted and Downloadable
Both Clef and Clef-flash are available through Workers AI and as Apache 2.0-licensed weights. Cloudflare positions Clef as the higher-precision option and Clef-flash as the model for latency-sensitive decisions.
The announcement explicitly gives Clef vision support and a 64K context window, making it relevant to classifications that require images or substantial input context as well as short text. Those capabilities should not automatically be attributed to Clef-flash without checking its model-specific documentation.
The models are also compatible with the Jev API, according to Cloudflare. Developers already using Jev can experiment without redesigning their entire request schema, though API compatibility does not establish equivalent behavior. Teams still need to compare classifications, confidence scores, and thresholds on their own data.
Downloadable weights provide another deployment option, but local inference is not necessarily lightweight. In The Register’s reporting, Cloudflare said Clef-flash requires 41 GB of GPU VRAM and Clef requires 85 GB, assuming single concurrency and a 64K context window. Those figures describe a stated deployment configuration, not every possible optimized setup.
The Register also reports hosted Clef pricing of $0.24 per million tokens, compared with $0.042 for Jev. Faster does not necessarily mean cheaper. Developers need to consider request sizes, throughput, infrastructure costs, and the consequences of incorrect decisions.
“Open-weight” is the useful description of this release. Although Cloudflare calls the models open source, The Register reports that their training datasets are not public. Apache 2.0 weights enable local experimentation and deployment; they do not expose the complete training process.
Clef Scores Choices Instead of Writing an Answer
The architectural change explains why Cloudflare expects these models to reduce decision latency.
According to its technical description, Clef uses a Qwen backbone for a prefill-only pass, then scores valid schema choices in parallel. Prefill is the stage in which a language model processes its input. Clef’s decision step does not subsequently generate an intermediate answer token by token.
Cloudflare says a specialized routing mechanism extracts evidence relevant to each option and allows fields to attend to other fields and the original input before scoring. Those internal representations determine the output without generated prose.
This trades an open-ended answer space for a defined set of decisions. If the application needs a department and a severity level, generating a written explanation first can be unnecessary work.
Tasks requiring investigation, explanation, or information outside the supplied context still need other components. Developers must also define meaningful answer choices. A fast classifier cannot compensate for a schema that leaves out the right action.
The Speed Claims Need Workload-Specific Testing
Cloudflare reports the following median latencies across its test set:
| Model | Cloudflare-reported median latency |
|---|---|
| Clef | 209.3 ms |
| Clef-flash | 38.8 ms |
| Jev | 524.1 ms |
These are Cloudflare’s benchmark results, not independently validated measurements. The company says it ran 43 evaluation benchmarks, but the aggregate latency figures should not be treated as service-level guarantees for a production application.

Frequently Asked Questions
3 questions
1Can Clef replace a general-purpose LLM?
Clef can replace some LLM calls that only require a bounded classification or choice, such as support routing, urgency triage, and selecting a defined tool category. It does not replace an LLM’s broader role in investigation, explanation, or response generation. Developers should evaluate decision accuracy and escalation behavior before automating consequential actions.
2
Sources
- October 1, 2026 announcementblog.cloudflare.com
- The Register’s reportingtheregister.com





