Announced 9 Oct 2026 · Sources checked
What did Cloudflare publish on 9 October?
The blog post is dated 9 October 2026 and timestamped 18:27:34 UTC. It says Cloudflare is releasing Clef-omni, cutting the price of Clef-flash, and making Clef faster. It points back to the 1 October Clef and Clef-flash launch. The Workers AI changelog “Clef-omni adds audio and video input, Clef-flash is now cheaper, and Clef is faster” is dated 9 October and names the hosted id @cf/cloudflare/clef-omni.
The Hub API object we opened lists id Cloudflare/clef-omni, createdAt 2026-10-09T04:10:22.000Z, lastModified 2026-10-09T17:49:32.000Z, sha 0db1cd2607d76a7bdb2a382f659e7b313079f84b, license apache-2.0, and base_model Qwen/Qwen3-Omni-30B-A3B-Instruct. The safetensors block reports 35,259,818,545 BF16 parameters. The card’s own line is “30B-A3B.” usedStorage is 70,753,087,464 bytes. Fifteen model-*.safetensors shards sit next to joint_head.safetensors and joint_schema_model.py.
This is a typed decision model, not a chat checkpoint. For the same noul / choice / score shape on a hosted API, see Celeris-1-decision. For an open-weight edge pair that also refuses to emit tokens, see Liquid AI’s open d1 models.
How does Clef-omni work, and what does it not do?
The card’s opening sentence is that Clef-Omni “turns a state and a schema of typed questions into decisions.” It reads text, JSON, images, audio or video and “returns a probability for every allowed option of every question in a single forward pass.” “There is no free-form text generation and no output parsing.” The API is “fully compatible with Jev and SystemOne.”
The backbone is “the Qwen/Qwen3-Omni-30B-A3B-Instruct thinker with its vision and audio encoders.” The card says the base model’s speech-output weights (talker and code2wav) “are included unchanged but are not used; load_release_model does not load them.” A joint schema head “reads the backbone’s final hidden states, routes evidence from the state to each question, and scores all options of all questions jointly.” The blog adds that Cloudflare froze the Qwen3 backbone, trained LoRA adapters, and combined label-smoothed cross-entropy with a Brier score.
That is a classifier head on an omni backbone, not a new mixture-of-experts recipe. For what a MoE spends compute on, see mixture of experts explained. For the hosted Decisions shape that also returns probabilities instead of prose, see OpenAI’s Decisions API. For why a typed answer is a different object from generated text, see structured AI output.
The Hub page still shows default chat widgets (“Tell me an interesting fact about the universe”). Those strings are not a supported generation mode on the card we opened. The card’s local recipe wants torch 2.11 and transformers 5.10.2 on a single H200 and “about 64 GB of GPU memory in bfloat16.” Image inputs need pillow; audio and video need av. SGLang is marked “coming soon.”
What do the hosted limits and prices actually say?
The Workers AI model page prices @cf/cloudflare/clef-omni at $0.15 per million input tokens with a 64,000-token context. The changelog table matches that figure and lists Clef at $0.240 with 64K and Clef-flash at $0.038 with 24K. A “How media inputs are billed” block says Clef models convert media to input tokens and “do not charge for output tokens.”
Hosted media is embedded only. The model page rejects remote URLs. Images: at most four, 4 MiB and 16 megapixels each, 8 MiB total decoded, 64 to 1,024 tokens each. Audio: at most four clips, 8 MiB and 300 seconds each. Video: at most two clips, 16 MiB and 60 seconds each, sampled at 2 frames per second. Audio and video together may total 16 MiB decoded. A video’s soundtrack is heard with its frames “when every video in the request has one.” Each second of video “costs up to 256 input tokens.” Audio is “about 780 tokens per minute.” A 480p clip is “about 8,600 frame tokens per minute.”
The changelog’s latency lines for omni are “about 20 ms” for text, “under 100 ms” for image or audio, and “about 300 ms” for a 21-second video with sound. The blog’s latency lines are different: “about 130 ms” median for text, “about 150 ms” for images, “a few hundred milliseconds” for audio, and “about 1.5 seconds” for that 21-second clip. Those are two Cloudflare pages. We timed neither.
| Model | Hosted price | Hosted context | What the 9 October notes add |
|---|---|---|---|
| @cf/cloudflare/clef-omni | $0.15 / M input tokens | 64K | Audio and video fields; Apache-2.0 weights |
| @cf/cloudflare/clef-flash | $0.038 / M input tokens, was $0.090 | 24K, was 64K | Self-host card still cites 256K |
| @cf/cloudflare/clef | $0.24 / M input tokens | 64K | Serving speed-up only; weights unchanged |
Which scores are Cloudflare’s, and where does omni lose?
The card labels its large table “Per-benchmark results from our internal run of the Decision Index 0.2.1 suite.” The blog reprints a shorter slice. BANKING77 macro-F1 is 94.8 for Clef-omni against 94.20 for Clef, 90.93 for Clef-flash and 79.74 for Jev. CLINC150+OOS is 97.7 against 97.43, 66.77 and 89.27. Amazon ESCI is 57.8, a thin lead over Clef’s 57.48. BFCL case-exact is 98.2, behind Clef-flash at 98.76.
Several rows go the other way. Home-appliance case-exact is 69.3 for omni against 82.95 for Clef and 97.73 for Clef-flash. PhishNChips is 73.2 against 79.60 for Clef. When2Call is 63.3 against 80.97 for Jev. GPQA Diamond is 47.4 against 78.3 for Jev. RAGTruth hallucination F1 is 42.0 against 79.4 for Clef. POP909-CL is 2.6. Those are Cloudflare’s numbers.
A second table, “Workflow evals” on the card and “TypeSafe evals” on the blog, scores invoice, customer-service, security-incident and agent-trace workflows. Clef-omni does not lead any of those five printed rows. Invoice exact actions are 60.2 against 64.7 for Clef. We did not rerun Decision Index or Typesafe.
What changed for Clef-flash and Clef?
Clef-flash’s hosted price is now $0.038 per million input tokens, down from $0.090. The changelog says that makes it cheaper than Jev. The trade is context: hosted Clef-flash is 24K tokens, down from 64K. Cloudflare writes that 0.24% of requests exceeded 24K. “The Clef-flash weights on Hugging Face are unchanged and support up to a 256K context window if you self-host.” Callers who need the old hosted window are told to use Clef.
Clef’s speed-up is a serving change. The changelog table for ~800 tokens is 262 / 438 ms median / p95 before and 152 / 351 after (1.7×). At ~3,400 tokens the median halves, 616 to 305 (2.0×). At ~16,000 tokens it is 2,721 to 1,635 (1.7×). “The model weights are unchanged.” Part of the speed-up is a move to SGLang. Clef support is “coming to SGLang in version 0.5.22 (PR #42721).” We did not open that pull request or time a server.
What should a team test before swapping a classifier?
The useful trial is a set of questions you already score: a support ticket, a short WAV, a 10-second MP4 with sound, and one image that must be embedded as a data URL. Compare Clef-omni’s probabilities with your current router at the same schema. The card’s BANKING77 lead is not a production SLA, and the home-appliance miss is on the same table.
Pin the revision. joint_schema_model.py is custom code. The hosted path is @cf/cloudflare/clef-omni with model set to clef-omni; AI Gateway is listed as compatible. Self-host wants an H200-class card on the recipe we read. We did not download 70 GB of shards, start Workers AI, or score a clip.
This is an evidence review of the 9 October blog, the changelog, the Workers AI model page, the Hub card, the Hub API object, and the Apache-2.0 LICENSE file.
Common questions
Are the Clef-omni weights actually open?
Yes on the object we opened. Cloudflare/clef-omni is public, license apache-2.0, with 15 safetensors shards and a joint head. The LICENSE file is Apache License 2.0, January 2004. The Hub widget’s chat prompts are not a generation API.
Did Cloudflare publish one latency number for a 21-second video?
No. The changelog says about 300 ms. The blog says about 1.5 seconds. Both are Cloudflare pages. We timed neither.
Is Clef-omni the best Clef on Cloudflare’s own board?
Not on the tables we opened. It leads some intent rows and trails on home appliances, phishing, When2Call and the Typesafe workflow set. Those are vendor rows.
What to remember
Use the 9 October blog and changelog for the hosted id, the $0.15 / $0.038 / $0.24 price list, and the 24K Clef-flash cut. Use the Hub card for Apache-2.0 weights and the no-text-output rule. Keep every bench and both video-latency lines in Cloudflare’s column.
Sources & further reading
- Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash ↗
- Clef-omni adds audio and video input, Clef-flash is now cheaper, and Clef is faster ↗
- Cloudflare/clef-omni model card ↗
- Cloudflare/clef-omni Hub API object ↗
- clef-omni on Workers AI ↗
- Cloudflare/clef-omni LICENSE ↗
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





