Cloudflare says its new Clef model means humans no longer need to be in the loop for AI agents

In this articleClef-flash returns a decision in 39 millisecondsCloudflare is based on Qwen modelsCustomers will be able to fine-tune Clef for their…

By Vane October 2, 2026 4 min read
Cloudflare says its new Clef model means humans no longer need to be in the loop for AI agents


Cloudflare releases Clef and Clef-flash to let AI agents decide without human oversight

Cloudflare has launched two new decision models, Clef and Clef-flash, designed to allow autonomous AI agents to make choices without waiting for human approval.

Instead of generating free-form text, these models output probabilities for a set of predefined answers. Downstream systems can then act immediately on those results. For example, if a customer support message arrives, the model assesses its urgency and identifies the correct team. Code can use this data to route a ticket or trigger an escalation automatically.

Cloudflare claims this removes the need for a human to be in the loop for agentic decisions. Agents can now programmatically gather context, make decisions, and take action on tasks, or defer to a human only when necessary.

The name comes from music, where a clef assigns pitches to lines on a staff. Similarly, a decision model sets the framework for actions that follow. The phonetic similarity to TypeSafe AI’s competing model, Jev, is likely intentional.

These “decision models” sit between large language models and traditional classifiers. Language models can reason and call tools but often produce variable outputs and can be slow. Traditional classifiers are fast but require retraining for every new category. Cloudflare views Clef as a direct response to TypeSafe AI’s Jev model and has kept the API fully compatible to allow customers to switch easily.

Clef-flash returns a decision in 39 milliseconds

Across 43 benchmarks, Clef and the smaller Clef-flash are faster than all relevant competing decision models, according to Cloudflare. The company reports median latency of about 39 milliseconds for Clef-flash and about 209 milliseconds for Clef, compared with just over 524 milliseconds for Jev. Both models run directly on Cloudflare’s infrastructure, which the company says also lets them benefit from proximity to edge data centers.

Cloudflare’s threat intelligence team is already testing Clef to classify websites. In one example, it assigns a domain a 95 percent probability of being a fashion website and 85 percent of being an online store. The probability of it being a phishing site is under one percent. Fetching, rendering, and classifying the site took 2.2 seconds. The company’s fastest general-purpose language model took 4.7 seconds for the same process and returned only two categories.

Clef can also process images, according to Cloudflare, while Jev is limited to text so far. Its 64,000-token context window holds twice as much input as Jev’s. Clef also leads on the Jev Decision Index in Cloudflare’s own benchmarks.

Benchmark · accuracyClefClef-flashJevDiffusionGemma JevKev 9BLaya
API Bank91.9393.1188.1983.6656.3011.41
When2Call72.3765.5880.9775.4449.6211.94
PhishNChips79.6075.0562.5585.3550.7550.15

Cloudflare is based on Qwen models

Clef is based on Qwen3.8-27B and Clef-flash on the smaller Qwen3.5-9B, according to Cloudflare. Cloudflare leaves the base models unchanged during training and uses its own synthetic data to train extra components. These components derive answer options and probabilities from the models’ internal computations.

Cloudflare also uses its own variant of Reinforcement Learning for Calibrated Decisions (RLCD), the training method TypeSafe used to train Jev. RLCD trains models to answer multiple questions about an input in a single call, aiming to assign probabilities that match how often the answers are actually correct. The company previously experimented with DiffusionGemma to derive fixed decision values from a language model’s internal computations.

Customers will be able to fine-tune Clef for their own tasks

Alongside the launch, Cloudflare is rolling out a reinforcement learning service that lets customers tailor Clef to their own tasks. A team of forward deployed engineers will initially handle fine-tuning with customers, with a self-service platform planned for later.

Customers will be able to build a dataset by logging requests through AI Gateway, then evaluate those requests in containers that serve as an RL sandbox. A new trainer component will let them deploy the fine-tuned model on Workers AI. To run custom models, Cloudflare uses technology from Replicate, which it acquired in late 2025.

Cloudflare says it plans to use Clef internally to review abuse reports, sort support requests, and distinguish useful bots from harmful ones. Both models run on Cloudflare’s Workers AI platform and are available on Hugging Face under the Apache-2.0 license.

The approach was popularized by TypeSafe AI, the startup founded by former OpenAI researcher Diogo Almeida, which introduced Jev in mid-September. TypeSafe markets Jev as a model “without hallucinations,” though that only guarantees it stays within predefined answer options and doesn’t prevent it from choosing incorrectly. In late September, OpenAI followed with a Decisions API built on GPT-6 Luna that also accepts context as text or images.

Cloudflare primarily runs a global network for content delivery, DNS, and security services and has attracted little attention for its own AI models. Its recent AI headlines have focused on giving website owners control over access. In July, it let site operators block or allow AI bots based on their purpose.

What it means

For the people building these systems, the change is practical. Developers no longer need to write code that waits for a human to click a button or approve a step. The system can read a request, pick a category, and send the task to the right queue instantly.

This reduces the friction in automating workflows. A support bot can route a complaint to the billing team without asking a supervisor to confirm the action first. The models run fast enough to handle this in real time, making the automation feel immediate rather than delayed.


Scroll to Top