THE CRUNCH
Cloudflare has released two open-source decision models, Clef and Clef-flash, alongside a reinforcement learning platform for fine-tuning them. These models are designed to make bounded, structured decisions cheaply and quickly, contrasting with the non-deterministic nature of large language models. The company claims Clef leads the market on the Jev Decision Index and is fully Jev-API compatible. Both models are hosted on Workers AI and available under an Apache 2.0 license on Hugging Face. Cloudflare also debuted a new RL product to let customers fine-tune Clef for specific use cases.
A decision model makes classifications to help agents decide how to act, based on probabilities. For example, a model can take a customer support message and return a typed answer indicating whether it is urgent and which team should handle it. This allows agents to programmatically gather context, make decisions, and take actions, or defer to a human when needed. Cloudflare tested Clef on its Threat Intelligence team to classify website domains. The model can identify categories, such as classifying a domain with a 95% chance it is a fashion website. In this test, Clef took 2.2 seconds to fetch, render, and classify a website, compared to 4.7 seconds for the general LLM gpt-oss-120b, which returned only two classifications.
Clef differs from other decision models in three key ways. First, it has a vision encoder that can classify visual content, unlike Jev, which only handles text. Second, it has a 64k context window, compared to Jev’s 32k, allowing more input state. Third, it is accurate and powerful, scoring competitively on various benchmarks. Cloudflare ran 43 eval benchmarks and found its Clef models beat other decision models on latency, except for Laya, which is very fast but trades off quality. The models are hosted on Workers AI, taking advantage of Cloudflare’s GPUs at the edge to reduce network latency.
Cloudflare has open-sourced the models on Hugging Face under an Apache 2.0 license, allowing users to run them locally. The company also debuted a new reinforcement learning product that allows customers to fine-tune Clef to suit their use cases. This RL platform is part of Cloudflare’s broader strategy to provide tools for building agentic workflows that can autonomously decide, reason, and execute.
The benchmarks show Clef and Clef-flash performing well across various tasks. For instance, on the BRIGHT benchmark, Clef achieved an nDCG@10 score of 45.91, while Clef-flash scored 39.26. In the BANKING77 benchmark, Clef scored a macro-F1 of 94.20, and Clef-flash scored 90.93. The models also performed well on Cloudflare’s own eval suite, beating Jev in three out of four areas. The median latency for Clef was 209.3ms, and for Clef-flash, it was 38.8ms, significantly faster than Jev’s 524.1ms.


