On October 1, 2026, Cloudflare released Clef — two open-source decision models that don’t generate text, don’t write prose, and don’t care about prompt engineering. They take a question with a list of options and return calibrated probabilities for each one. If you’re building AI agents and paying frontier LLM rates for routing and classification work, Clef is a direct challenge to that cost structure.
What Cloudflare Clef Decision Models Actually Do
Most developers treat LLM calls as the default answer to any input-to-output problem. That works, but it’s expensive when your agent’s “reasoning” is just picking between three options. Decision models take a state blob — some context — and a structured question: choose A, B, or C; is this message urgent (yes/no); which tool should fire next. They return probabilities per option, not a paragraph. No hallucinations, no latency from token generation, no $15-per-million-token bill.
The pattern emerging in production agent stacks has three layers: frontier LLMs for planning and reasoning, decision models for routing and gating, and deterministic code for anything with a known answer. If your agent makes 100 calls per task and 70 are routing decisions — “which tool?”, “safe to run?”, “should I escalate?” — those 70 should not be hitting GPT-6. Jev, the incumbent decision model from TypeSafe AI, prices at $0.042 per million input tokens. Frontier LLMs run $2–15. The cost difference is 50–300x for calls that don’t require reasoning.
Related: OpenAI Decisions API: Skip the Parser, Get Typed Answers
Clef vs Clef-Flash: Specs and Benchmarks
Cloudflare released two variants. Clef is the 27B flagship, built on a Qwen 3.8-27B backbone with a non-autoregressive scoring architecture, 64k context window, and a vision encoder — it can classify screenshots and images alongside text, something Jev can’t do. Cloudflare reports a 209ms median latency and a benchmark index score of 61.2, compared to Jev’s 57.9. On BANKING77 classification, Clef scores 94.20 macro-F1 against Jev’s 79.7.
Clef-flash is the 9B model for latency-critical paths. Cloudflare’s benchmarks show 38.8ms median and 122.4ms at p95 — roughly 13x faster than Jev’s 524ms. The weights for both are on Hugging Face under Apache 2.0, and they’re Jev-API compatible, so switching is a configuration change, not a rewrite.
| Model | Latency (CF claim) | Independent Test | Index Score | Price/M tokens |
|---|---|---|---|---|
| Clef (27B) | 209ms | N/A | 61.2 | TBD |
| Clef-flash (9B) | 38.8ms | ~661ms (beta) | 57.1 | ~$0.09 |
| Jev | 524ms | ~230ms | 57.9 | $0.042 |
The Benchmark Gap You Need to Know About
Cloudflare’s benchmarks show Clef-flash at 38.8ms. However, independent developers testing on Workers AI during the beta reported medians closer to 661ms, with one noting it “takes 3 seconds, which is useless.” Cloudflare’s numbers reflect their own tuned edge infrastructure. The public beta runs differently. That gap should close as the platform matures, but it’s worth knowing before you build a latency-critical path around the benchmark sheet.
An independent developer tested 42 real-world agent decisions — computer-use actions, supervision classification, message triage — against both Clef-flash and Jev. Overall accuracy: Jev 71.4%, Clef-flash 66.7%. Agreement between the two: 83%. The useful finding wasn’t a winner. It was this: on high-consequence decisions, the free local model and the paid API agreed 83% of the time. The conclusion isn’t “use Jev” or “use Clef.” It’s that you’re probably paying frontier LLM rates for work that either model handles correctly.
On the “open-source” label: the weights are Apache 2.0, which means you can run, modify, and deploy them. Nevertheless, the training data and pipeline are not public. That’s open-weight, not open-source in the full sense. For self-hosting, the distinction doesn’t matter. For teams that need reproducibility or want to audit training, it does.
The RL Fine-Tuning Platform Is the Real Play
Cloudflare shipped something more interesting alongside the models: a reinforcement learning fine-tuning platform. Phase 1 is hands-on — Cloudflare’s Forward-Deployed Engineers help you adapt Clef to your domain. Phase 2, coming later, will be self-serve: AI Gateway captures your production request/response data, Workers AI runs rollouts, a Trainer component updates the weights, and you redeploy.
Consequently, this is the actual long-term differentiation. A Clef model fine-tuned on your domain’s classification data — your specific ticket categories, your agent’s action vocabulary, your safety rules — will outperform a general-purpose frontier LLM at your specific task at a fraction of the cost. The open-weight models are the entry point. The fine-tuning flywheel is the moat.
Related: IBM Bob 2.0: The Multi-Model IDE That Routes to the Right LLM
How to Start Using Clef Today
Clef and Clef-flash are live on Workers AI — any developer with a Workers Paid plan can use them now. The weights are also available on Hugging Face for self-hosting. The official changelog has integration details. API compatibility with Jev means if you’re already using Jev, migration is minimal.
The practical approach: audit your agent’s call log. Find calls that are pure classification — routing decisions, priority scoring, tool selection, safety gates. Move those to Clef-flash first. Keep frontier LLMs for planning, multi-step reasoning, and output generation. The split-stack approach isn’t theoretical; it’s what production agents are running now.
Key Takeaways
- Decision models return typed probabilities, not text — they’re the right tool for routing, classification, and gating inside agent stacks, not frontier LLMs
- Clef-flash’s official latency (38.8ms) and independent beta measurements (~661ms) diverge significantly — the infrastructure is still maturing
- At 83% agreement between Clef-flash and Jev on agent decisions, the cost argument for switching matters more than the accuracy delta
- The RL fine-tuning platform is Cloudflare’s real differentiator — domain-specific fine-tuning will outperform general models at your specific classification task
- Weights are Apache 2.0 on Hugging Face; “open-source” is technically open-weight — training data and pipeline remain closed













