Cloudflare has released Clef and Clef-flash, the first models built by its Workers AI team. These are decision models, not chatbots. Each takes an input state and a schema of typed questions. It returns a probability for every allowed answer, with no free-form text. Both are open-weight under Apache 2.0 and compatible with TypeSafe AI’s Jev API.
In this article
Deployment status
Both models run today on Workers AI, and the weights are on Hugging Face for self-hosting.
What a decision model does
An LLM generates tokens one at a time, and its output still needs parsing. A decision model only answers a fixed set of questions about an input. Clef supports three question types:
noul: yes/no, returns the probability of yes.choice: picks 1 named option, with per-option probabilities and a confidence value.score: rates against an ordered rubric, returning a probability-weighted score.
On Workers AI, 1 request carries up to 64 questions and up to 4 images.
TypeSafe AI launched Jev, its first System One model, on September 15, 2026. Open alternatives like Kev-9B and Laya followed. Clef uses the same System One API. Switching from Jev means changing the endpoint and model name.
How Clef works
Clef is post-trained from Qwen3.8-27B, and Clef-flash from Qwen3.5-9B. Both keep the backbone’s vision encoder.
Inference has two stages. The backbone first runs a single prefill-only pass over the state and questions. A small transformer, the joint schema head, then reads the final hidden states. It routes evidence to each question, lets fields cross-attend, and scores all options jointly. A per-question softmax turns logits into probabilities.
Training froze both backbones and jointly optimized the routing head with rank-256 low-rank adapters. The loss pairs label-smoothed cross-entropy with a Brier loss for calibration. A secondary objective, Reinforcement Learning for Calibrated Decisions (RLCD), gives partial credit to adjacent ordinal choices.
Benchmarks: Where Clef wins and where it does not
On Cloudflare’s 10-benchmark shortlist from the Decision Index 0.2.1 suite, a Clef model scored highest on 7.
- BANKING77 (macro-F1): Clef 94.20 vs Jev 79.74.
- CLINC150+OOS (macro-F1): Clef 97.43 vs Jev 89.27.
- Home appliances (case exact): Clef-flash 97.73 vs Jev 52.27.
Jev keeps clear leads elsewhere. The full model card shows Jev ahead on GPQA Diamond (78.3 vs 48.0). It also leads MMLU-Pro (82.7 vs 65.9) and BBH (92.9 vs 73.7).
On TypeSafe’s own workflow evals, Clef beat Jev in 3 of 4 areas, by small margins. Invoice processing was 64.7 vs 61.8, customer service 76.3 vs 76.0, and security incidents 62.9 vs 61.7. Jev leads agent trace observability, 71.6 vs 68.5.
In Cloudflare’s threat intelligence workflow, Clef classified a domain in 2.2 seconds. gpt-oss-120b took 4.7 seconds.
All numbers are vendor-reported, with no independent replication yet.
Feature Comparison
| Feature | Clef | Clef-flash | Jev | Kev-9B | Laya |
|---|---|---|---|---|---|
| Developer | Cloudflare | Cloudflare | TypeSafe AI | Jared Palmer | Convai Innovations |
| Size | 27B | 9B | Not disclosed | 9B + 45.4M LoRA | 421M |
| Backbone | Qwen3.8-27B | Qwen3.5-9B | Not disclosed | Qwen3.5-9B-Base | ModernBERT-large |
| Weights | Apache 2.0 | Apache 2.0 | Hosted API | Apache 2.0 | Apache 2.0 |
| Image input | Yes | Yes | No | No | No |
| Context | 65,536 | 65,536 | 32K (per Cloudflare) | 65,536 (8,192 validated) | 512 (English) |
| Median latency* | 209.3 ms | 38.8 ms | 524.1 ms | 51.4 ms | 5.8 ms |
| Hosted price (input) | $0.24/M | $0.09/M | $0.042/M | Self-host | Self-host |
*Cloudflare’s internal Decision Index run. All 5 implement the System One API. Sources: Clef docs, Clef-flash docs, Clef card, Jev post, Kev-9B card, Laya card.
Deployment and fine-tuning
Both models are callable through the Workers AI binding (env.AI.run()), the REST API, or AI Gateway. For self-hosting, the model cards list testing on a single H200 with BF16 weights.
Cloudflare also announced a reinforcement learning service for tuning Clef on private data. It starts with Cloudflare’s forward-deployed engineers, with a self-serve platform later. The pipeline combines AI Gateway, Workers AI, Containers and a new Trainer component. Teams can apply via the design partner form.
What it means
Developers using decision models now have an open-weight alternative to TypeSafe AI’s Jev. The shift is technical rather than philosophical. Users must change the endpoint and model name to switch from Jev to Clef. The trade-off is speed and accuracy on specific tasks. Clef-flash runs in 38.8 ms, while Jev takes 524.1 ms. Clef also handles images and video within a 64K-token context window. For teams needing custom training, Cloudflare is rolling out a reinforcement learning service to tune these models on private data.




