Cloudflare Clef: When the Edge Network Learned to Make Decisions
Cloudflare's new open-source decision models run at the edge with 40ms latency, beating Jev on accuracy while adding vision support and a 64k context window. Here's why decision models are the missing piece in agentic AI.
For years, the AI community has been obsessed with making language models bigger, smarter, and more capable of open-ended reasoning. But a funny thing happened along the way: we discovered that most real-world AI applications don't need a model that can write poetry or debate philosophy. They need a model that can make a decision — fast, reliably, and with calibrated confidence.
That's exactly what Cloudflare just bet on with Clef, a pair of open-source decision models that launched this week and immediately shot to the top of Hacker News with over 500 upvotes. And unlike the typical AI announcement that promises the moon and delivers a demo, Clef is already running in production on Cloudflare's edge network with numbers that make you pay attention.
What Is a Decision Model, Exactly?
The concept of a decision model is deceptively simple. Instead of asking a large language model to generate free-form text responses, you ask a specialized model to classify inputs and return structured, bounded outputs with probability scores. Think of it as the difference between asking a human expert to write an essay about a customer support ticket versus asking them to simply categorize it: urgent or not, which team handles it, severity level one through four.
The category gained momentum recently with Typesafe AI's Jev System One model, which introduced the idea of a dedicated decision model that produces typed answers with probabilities rather than open-ended text. Cloudflare took one look at that concept and decided they could do it better — and faster — by leveraging their edge infrastructure.
The result is Clef and Clef-flash, two models that are fully Jev-API compatible but bring some serious advantages to the table.
What Makes Clef Different
Cloudflare didn't just clone Jev. They built something that pushes the category forward in three meaningful ways:
- Vision support — Clef includes a vision encoder, so it can classify images and visual content. Jev only handles text today. This is a massive advantage for use cases like content moderation, document classification, or visual threat intelligence.
- 64k context window — Double Jev's 32k context. More context means more state to reason over, which matters when you're feeding an agent conversation history or a large document for classification.
- Edge deployment — Running on Cloudflare's Workers AI means decisions happen at the edge with single-digit millisecond network latency. The model itself adds only 39ms median latency for Clef-flash, compared to 524ms for Jev.
The benchmark numbers are striking. On the Jev Decision Index, Clef leads or ties on 7 out of 10 benchmarks. On BANKING77 (intent classification), Clef scores 94.2 macro-F1 versus Jev's 79.7. On CLINC150+OOS (out-of-scope detection), it's 97.4 versus 89.3. Even Clef-flash, the smaller and faster variant, beats Jev on several benchmarks while running 13x faster.
The Technical Approach: Non-Autoregressive Decisions
Here's where it gets technically interesting. Most LLMs generate text autoregressively — one token at a time, each depending on the previous. That's slow by definition. Clef takes a fundamentally different approach.
It uses Qwen as the backbone model (Qwen3.8-27B for Clef, Qwen3.5-9B for Clef-flash) but only for a prefill-only pass. The actual decision step is non-autoregressive — it scores valid schema choices in parallel directly from internal representations, rather than generating intermediate text. A specialized two-stage attention routing process lets each valid choice extract relevant context from the prompt, cross-attend with other fields, and reference the original payload before scoring.
In plain English: instead of writing out an answer word by word, Clef reads the input once and simultaneously scores all possible answers. That's why it's fast enough to sit in the hot path of an agent workflow without adding noticeable latency.
Real-World Impact: Cloudflare's Own Dogfooding
Cloudflare is already using Clef internally on their Threat Intelligence team to classify website domains. Give Clef a domain and it fetches, renders, and classifies it in 2.2 seconds — returning probability scores across multiple categories (95% fashion, 85% ecommerce, less than 1% phishing). Their fastest general LLM, gpt-oss-120b, took 4.7 seconds on the same task and returned fewer classifications.
That 2x latency improvement matters when you're classifying millions of domains. But the bigger story is the pattern: Clef sits between an agent's reasoning steps, making quick routing decisions that determine what happens next. It's the connective tissue that makes agentic workflows practical.
The RL Fine-Tuning Platform
Perhaps the most strategically interesting part of this announcement is the reinforcement learning fine-tuning platform. Cloudflare is offering a service where customers can fine-tune Clef for their specific workloads, using their own labeled data. The pipeline leverages existing Cloudflare primitives:
- AI Gateway captures your request/response traffic to build training datasets automatically
- Workers AI generates rollouts against the base Clef model
- Containers serve as RL sandboxes for scoring and replaying agent actions
- A new Trainer component updates weights of the fine-tuned model
- Workers AI + Bring Your Own Model redeploys the fine-tuned model at the edge
This is Cloudflare building the full stack: capture data, train, deploy, and run — all within their platform. It's an ambitious play that positions them not just as a CDN or edge compute provider, but as an end-to-end AI infrastructure company.
Why This Matters for the Agentic AI Stack
The current agentic AI stack has a gap. LLMs are great at reasoning and generating actions, but they're slow, expensive, and non-deterministic. Traditional classifiers are fast and cheap but brittle — you have to retrain them every time your categories change. Decision models sit in the sweet spot: fast enough for real-time use, flexible enough to handle new categories without retraining, and structured enough to programmatically drive agent behavior.
Cloudflare's bet is that decision models become a standard primitive in every agentic workflow — the routing layer that determines whether an agent should escalate to a human, which tool to call next, or whether a response is safe to send. By open-sourcing the weights under Apache 2.0 and hosting them on their edge network, they're making it trivially easy to adopt.
The name itself is a nice metaphor. In music, a clef sits at the beginning of a staff and defines the context for everything that follows. Cloudflare's Clef does the same for agents — it makes the first decision that shapes all subsequent actions. Whether that bet pays off depends on whether developers actually build decision-model-driven agents, or whether LLMs get fast and cheap enough that a dedicated routing layer becomes unnecessary.
Given that Clef-flash returns decisions in 39 milliseconds at the edge while a fast LLM takes 500+ milliseconds, the performance gap would need to close by an order of magnitude before decision models become redundant. For now, at least, the clef is setting the tempo.
Related Posts
Strata: When a 125B AI Model Ran on a Gaming PC at 100 Tokens Per Second
A new open-source tool called Strata lets you run Qwen 3.8 Flash Next — a 125-billion-parameter model — on an ordinary gaming PC with an RTX 4090. Nothing leaves your machine, and it's faster than you can read.
When AI Agents Spend Your Money While You Sleep: Why Hard Budget Caps Are Becoming Non-Negotiable
AWS and Google Cloud finally launched hard spending limits in the same month. It's not a coincidence — it's a response to AI agents that can rack up thousands of dollars before you wake up.
When Utah Banned VPNs: How a Court Stopped a Law That Demanded the Technically Impossible
A federal judge just blocked Utah's unprecedented anti-VPN law, ruling that lawmakers cannot mandate perfect geolocation — a technical impossibility. The case reveals a deeper problem: when legislation outruns engineering.