<h2>Decision AI models return choices, not paragraphs</h2>
<p>You send text with typed questions and the model returns scores, probabilities, or yes/no answers your code can branch on directly.</p>
<p>The category went mainstream when TypeSafe AI launched Jev after two years in stealth. TypeSafe calls it a 'System One model,' after Daniel Kahneman's fast, intuitive System 1 thinking.</p>
<p>Within three weeks, Fastino Labs shipped two rival models, and open-source developers published several Jev-style reproductions. This article covers how the category works, where it fits, and how the options compare.</p><h2>How Jev works</h2>
<p>Jev accepts a 'state' (a string, array or set of name-value pairs) and one or more questions. According to the TypeSafe docs, it supports three primitives:</p>
<ul>
<li><strong>Choice:</strong> pick one option from a list, with probabilities and confidence. TypeSafe says Jev supports up to 255 options.</li>
<li><strong>Score:</strong> rate the state on a rubric of ordered levels, with probabilities and confidence.</li>
<li><strong>Noul:</strong> a 0 to 1 probability that a statement is true. The name is short for Bernoulli.</li>
</ul>
<p>Every question is evaluated in parallel and in isolation against the same state. Adding questions barely changes response time. Because Jev never generates strings, TypeSafe says it cannot return a type error.</p>
<p>Under the hood, TypeSafe describes a new architecture, a parallel sampler and a training method called Reinforcement Learning for Calibrated Decisions (RLCD). RLHF optimises for human preference. RLCD optimises for calibrated probabilities, where higher confidence should mean higher accuracy.</p>
<p>Pricing is the key thing to watch. Jev costs $0.042 per million input tokens, and output is free. OpenRouter lists a 32K context window. TypeSafe reports end-to-end responses in 70 to 500 milliseconds.</p>
<h2>Benchmarks: What TypeSafe's workflow evals show</h2>
<p>TypeSafe built workflow evals across four tasks: security incidents, agent trace observability, invoice processing and customer service. Reference labels come from averaging GPT-6 Astra and Claude Fable 5.1 at high thinking.</p>
<ul>
<li><strong>Jev:</strong> 67.8% mean accuracy, $0.0004 per case, 0.4 seconds.</li>
<li><strong>Claude Sonnet 5 (same workflow):</strong> 67.8%, $0.1174 per case, 78.1 seconds.</li>
<li><strong>Best comparison model (OpenAI "sol"):</strong> 74.1%, $0.0836 per case, 23.3 seconds.</li>
</ul>
<p>Jev matched Sonnet 5 on accuracy at a fraction of the cost and latency. It still trails the top frontier configuration by 6.3 points. Per task, Jev scored 76.0% on customer service but only 61.8% on invoice processing.</p>
<h2>Use cases: Where decision models fit</h2>
<p>The rule of thumb is simple. If your code needs a bounded answer it will branch on, a decision model is a candidate. If a human needs to read the output, use an LLM.</p>
<h3>Agent control flow</h3>
<ul>
<li><strong>Choosing the next tool or subagent:</strong> Vercel lists this as a primary use.</li>
<li><strong>Continue, retry, ask the user, or stop:</strong> A single Choice question replaces a fragile JSON-parsing step.</li>
<li><strong>Model routing:</strong> Send easy requests to a cheap model and hard ones to a frontier model. Fastino lists routing by destination, complexity or escalation level.</li>
</ul>
<h3>Classification and triage</h3>
<ul>
<li><strong>Support ticket routing, email triage and intent detection:</strong> These make up much of Fastino's 17-dataset benchmark.</li>
<li><strong>Spam detection, label suggestions and prioritization:</strong> Simon Willison calls these natural classification fits.</li>
<li><strong>Security alert triage:</strong> TypeSafe's simplest published workflow decides whether to close an alert, pass it to an analyst, or contain it.</li>
</ul>
<h3>Verification and safety</h3>
<ul>
<li><strong>Guardrails and jailbreak detection:</strong> TypeSafe pitches Jev for scoring prompts, reasoning traces and outputs.</li>
<li><strong>Joint safety checks:</strong> GLiNER2.5-Decide can decode "safety" and "harm type" together so the answers never contradict.</li>
<li><strong>Verified cascades:</strong> OpenRouter's Jev guide describes drafting with a cheap model, checking with Jev, and escalating only on failure.</li>
</ul>
<h3>Evaluation and observability</h3>
<ul>
<li><strong>LLM-as-a-judge replacement:</strong> Arize and Langfuse now run Jev evaluators on traces.</li>
<li><strong>CI/CD gates:</strong> Buddy lets pipelines score, classify or gate runs with a Jev action.</li>
</ul>
<h3>Search and data processing</h3>
<ul>
<li><strong>Reranking:</strong> Score 100 BM25 candidates for relevance in one parallel call.</li>
<li><strong>Map-reduce over large datasets:</strong> TypeSafe pitches turning bulk data into features at low cost.</li>
<li><strong>Context pruning:</strong> Fastino lists choosing which context to keep before an LLM call.</li>
</ul>
<h3>Real-time applications</h3>
<ul>
<li><strong>Games and simulations:</strong> TypeSafe demoed Jev playing Doom and Wikiracing.</li>
<li><strong>Latency-critical UX:</strong> Sub-second decisions make AI usable inside interactive flows.</li>
</ul>
<h3>When not to use a decision model</h3>
<ul>
<li>You need generated text, summaries or explanations.</li>
<li>The task needs exact arithmetic, counting or date math. TypeSafe's Jev 1.13 jaggedness guide flags all three.</li>
<li>The decision affects people's livelihoods, such as hiring. Willison warns that hidden bias is hard to inspect.</li>
</ul>
<h2>Case studies: Where decision models are already working</h2>
<h3>1. Vercel AI Gateway adoption</h3>
<p>Vercel reports that Jev became the fastest-adopted model in AI Gateway history. Within 24 hours, nearly 13% of paid teams were using it. That was 2x the GPT-5.6 family's share and more than 6x Fable 5.1's.</p>
<h3>2. Search reranking</h3>
<p>Simon Willison fetched 100 candidates with BM25, then had Jev score each for relevance. This pattern replaces an expensive LLM reranker with cheap, parallel Score questions.</p>
<h3>3. LLM evaluation pipelines</h3>
<p>Arize and Langfuse both shipped Jev-as-a-judge evaluators. Langfuse labels the feature "decision-model evaluators" so other models can be added later. Its Jev support is still marked experimental.</p>
<h3>4. 6G edge network orchestration</h3>
<p>A new arXiv paper used Jev to interpret service contracts at the network edge. At matched correctness, Jev cut median decision latency by 22.4% versus DeepSeek and 61.9% versus Gemini.</p>
<h3>5. Real-time game agents</h3>
<p>TypeSafe's Doom demo ran Jev at 10 queries per second. The team estimated the cost at roughly $7 per hour.</p>
<h2>Competitor landscape: 7 decision models compared</h2>
<p>The table below covers Jev, its closest commercial rival, and the leading open-weight options. All specs come from each project's own release page or model card.</p>
<figure class="wp-block-table">
<table class="has-fixed-layout">
<thead>
<tr>
<th>Model</th>
<th>Developer</th>
<th>License / access</th>
<th>Size and architecture</th>
<th>Question types</th
Source Read original →Decision AI Models Explained: TypeSafe Jev vs Fastino GLiDE, GLiNER2.5-Decide and Open-Source Competitors
<h2>Decision AI models return choices, not paragraphs</h2> <p>You send text with typed questions and the model returns scores, probabilities, or yes/no answers…




