Clef-omni evaluates text, images, audio, and video together, returning probabilities for predefined answers supplied by an application. It won’t write a response explaining what it found. Cloudflare’s new model serves as a decision layer for agents, not a conversational LLM.
Announced on October 9, 2026, the model extends Cloudflare’s open-weight Clef family to workflows that previously needed separate transcription or media-processing steps. An application can submit an installation photo, a recording of machinery, and a video of its operation, then ask several schema-bound questions in one request.
The launch also changes the economics of the existing lineup. Cloudflare cut Clef-flash’s hosted input price from $0.090 to $0.038 per million tokens, but reduced its hosted context window from 64K to 24K. Separately, the company reports faster hosted Clef inference without changing that model’s weights.
Each change needs its own evaluation: broader media support for Clef-omni, a price and context trade-off for Clef-flash, and a serving upgrade for Clef.
A Multimodal Decision Layer, Not a Chatbot
A conventional generative workflow might ask an LLM to inspect evidence, write an answer, and format it for an application. Clef-omni scores the allowed options directly.
The application supplies a state, which can contain text or structured data, and a questions schema. Questions can ask for a probability, a choice among named options, or a score across defined levels. Cloudflare’s documentation includes support-ticket examples that evaluate urgency, select a responsible team, and score customer impact.
A routing system typically needs a team identifier or an escalation decision, not a paragraph. Clef-omni outputs those constrained decisions; the application remains responsible for what happens next.
Cloudflare says the model uses the comprehension backbone of Qwen3-Omni-30B-A3B-Instruct, a mixture-of-experts architecture with 30 billion total parameters and 3 billion active parameters. It discards the speech-output components and trains low-rank adapters while freezing the backbone.
According to the company’s architectural description, candidate answers gather evidence across the input before a joint scoring mechanism computes confidence values. The model processes the payload without generating an output-token sequence.
Cloudflare distinguishes Clef models from LLMs because they do not generate text. The useful technical distinction is their inference behavior: Clef-omni inherits a multimodal foundation-model backbone but exposes a schema-bound scoring interface.
Probabilities still require validation. A constrained answer space prevents unexpected prose, but it does not establish that the selected action is correct or that the confidence value is reliable on a particular deployment’s data.
One Request Can Combine Media, Within Specific Limits
The hosted model is available as @cf/cloudflare/clef-omni. Requests also require the model selector clef-omni.
The Clef-omni model documentation specifies these supported media formats and limits:
- Images: PNG, JPEG, or WebP, with up to four images per request. Each can be at most 4 MiB and 16 megapixels, with an 8 MiB combined decoded-image limit.
- Audio: WAV or MP3, with up to four clips. Each clip can be at most 8 MiB and 300 seconds long.
- Video: MP4 or WebM, with up to two videos. Each can be at most 16 MiB and 60 seconds long, sampled at two frames per second.
Audio and video together may total no more than 16 MiB decoded. Media is supplied as embedded base64 data URLs in the images, audio, and videos fields; remote media URLs are not accepted.
Cloudflare says a video’s soundtrack is processed alongside its frames. The documentation qualifies this: the soundtrack is heard with the frames when every video in the request has one.
Developers evaluating brief visual events should test whether the two-frame-per-second sampling rate preserves the evidence their decision requires. Accepting video does not imply inspecting every original frame.
Clef-omni has a 64,000-token hosted context window, shared by the questions, text state, and media. Images are resized and converted into patch tokens, with a maximum of 1,024 tokens per image. Audio uses approximately 780 tokens per minute. Video can consume approximately 15,400 frame tokens per minute at maximum resolution, plus about 780 tokens per minute for sound.
At the $0.15-per-million-input-token rate, one minute of audio contributes roughly $0.000117 in media-token charges. A minute of maximum-resolution video with sound contributes approximately $0.00243, before text and questions. These are calculations from Cloudflare’s documented tokenization rates, not measurements of a deployed workload.
If media and questions exceed the context window, the request fails. Otherwise, the text state is truncated to fit the remaining space. Developers should budget context explicitly: a long text record may not survive unchanged beside media.
Clef-flash’s Price Cut Comes With a Smaller Hosted Window
Cloudflare’s Workers AI changelog lists the updated hosted lineup:
| Workers AI model ID | Price per million input tokens | Hosted context window |
|---|---|---|
@cf/cloudflare/clef-flash | $0.038 | 24K tokens |
@cf/cloudflare/clef | $0.240 | 64K tokens |
@cf/cloudflare/clef-omni | $0.150 | 64K tokens |
Clef models do not charge for output tokens. Media is converted into input tokens and billed at the applicable model’s input rate.
Clef-flash’s reduction from $0.090 to $0.038 is approximately 57.8%. At one billion input tokens, its listed inference charge falls from $90 to $38.
For an illustrative workload of one million decisions consuming 1,000 input tokens each, that is a $52 reduction. The calculation assumes those 1,000 tokens include everything billed for each request; it excludes other application and infrastructure expenses.
Existing integrations also need to account for the smaller context window. A workflow designed around the previous 64K hosted allowance cannot assume that the cheaper service accepts the same inputs.
Cloudflare says only 0.24% of its observed requests exceeded 24K input tokens, which informed the change. That aggregate usage figure does not tell an individual customer whether its long documents, accumulated agent history, or large schemas fit the new limit.
The unchanged Clef-flash weights support up to 256K context when self-hosted, according to Cloudflare. The price cut and smaller window describe the Workers AI deployment, not a newly trained, shorter-context model.
The company says it has also released Clef-omni’s weights. Self-hosting provides an alternative deployment route, but the Workers AI token prices are not self-hosting cost estimates. Hardware, serving software, utilization, and operating effort determine those economics.
Clef’s Speedup Is a Serving Change
The speed improvements apply to hosted , not automatically to every member of the family.
Sources
- Clef-omniblog.cloudflare.com
- Clef-omni model documentationdevelopers.cloudflare.com
- Workers AI changelogdevelopers.cloudflare.com





