What Clef Actually Does
Clef and its faster sibling Clef-flash are decision models—a category distinct from large language models. Instead of generating free-form text, they take structured inputs and return typed outputs with probabilities. Think: a support ticket goes in, and out comes { team: "technical", urgent: true, severity: 0.94 }.
No hallucinations. No token-by-token generation. Just fast, bounded answers your code can act on directly.
The concept isn’t entirely new—classifier models have existed for years—but the framing here borrows from Typesafe AI’s Jev model, which popularized the idea of decision models as a distinct primitive in agentic workflows. Clef is Cloudflare’s answer: more capable, faster, and fully open-source under Apache 2.0.
The Numbers That Matter
Latency is where Clef makes its clearest case. Based on Cloudflare’s own benchmarks:
- Clef median latency: 209ms
- Clef-flash median latency: 38.8ms
- Jev median latency: 524ms
Cloudflare tested this internally on their Threat Intelligence team, using Clef to classify website domains. The result: 2.2 seconds to fetch, render, and classify a domain versus 4.7 seconds with their fastest general LLM—and Clef returned more classification categories in that time.
That’s not a marginal improvement. For anything in a hot path, it’s a meaningful architectural shift.
On quality benchmarks across the Jev Decision Index, Clef scores competitively—leading on several evals including BFCL case exact (98.47), BANKING77 macro-F1 (94.20), and API-Bank accuracy (91.93). Clef-flash punches above its weight given how fast it runs.
What Makes Clef Different
A few things separate Clef from other decision models currently available:
- Vision encoder — Clef can classify images, not just text. Jev currently handles text only.
- 64k context window — double Jev’s 32k, useful when you need to pass in richer state.
- Jev-API compatible — swapping in Clef from Jev is straightforward by design.
- Edge-hosted on Workers AI — Cloudflare’s GPU infrastructure means low network latency on top of the model’s already fast inference.
The architecture is worth a brief note: Clef uses a prefill-only pass through a Qwen backbone, then scores valid schema choices in parallel. There’s no autoregressive token generation in the decision step, which is a large part of why it’s fast.
The RL Fine-Tuning Layer
Beyond the model itself, Cloudflare is launching a reinforcement learning fine-tuning service for enterprise customers who need Clef adapted to specific domains.
The pitch is sensible: a generic decision model is good, but a model trained on 15 years of your network data is better. Cloudflare is already using this internally—for Trust & Safety submissions, support triage, and bot classification.
The fine-tuning pipeline connects existing Cloudflare primitives:
- AI Gateway captures your traffic as a training dataset
- Workers AI generates rollouts against the base Clef model
- Containers serve as RL sandboxes
- A new Trainer component updates the fine-tuned model weights
- Workers AI + BYO Model redeploys the result
Initially this runs as a hands-on service with Cloudflare’s forward-deployed engineering team. A self-serve platform is planned to follow.
Who This Is For
Clef is most useful if you’re building agentic workflows where decisions happen frequently, latency matters, and you don’t need a full LLM to make a call. Customer support routing, content moderation, security triage, domain classification—anywhere you’re currently using an LLM just to get a structured yes/no or a category label.
It’s also worth noting the open-source angle. The weights are on Hugging Face under Apache 2.0, so you can run Clef locally without touching Cloudflare’s infrastructure at all. That’s a meaningful option for teams with data residency requirements or who just want to experiment without API costs.
The Practical Takeaway
Clef isn’t trying to replace LLMs—it’s trying to stop you from using them where you don’t need to. If your agent is calling a large model just to route a ticket or flag a domain, you’re paying in latency and cost for capability you’re not using.
A fast, typed, open-source decision model that slots cleanly into an existing workflow is a useful tool to have. Whether Clef becomes a default primitive in agentic stacks depends on how the ecosystem around decision models matures—but the benchmarks and the open-source release make it worth evaluating now rather than later.
Comments (0) No comments yet
Want to join this discussion? Login or Register.
No comments yet. Be the first to share your thoughts!