AWS Strands Labs Releases Strands Decider 2B: An Open Source Decision Model That Picks Options in About 115 ms

AWS Strands Labs has released Strands Decider 2B, an open source model designed to select options rather than generate text. The system…

By Vane October 2, 2026 4 min read
AWS Strands Labs Releases Strands Decider 2B: An Open Source Decision Model That Picks Options in About 115 ms

AWS Strands Labs has released Strands Decider 2B, an open source model designed to select options rather than generate text. The system takes a state and a typed question, then outputs a choice, a yes/no probability, or a score with a calibrated confidence level. It contains 1.9 billion parameters and runs locally on a CPU, a consumer GPU, or an Apple silicon Mac.

Deployment is available for local and self-hosted environments. Weights are hosted on Hugging Face under an Apache-2.0 license. Installing the package via pip install strands-decider provides a command-line interface and an HTTP server. The bundled server binds to 127.0.0.1 without authentication, meaning production setups must add their own security layer. No hosted inference provider currently offers this model.

What a decision model does

Decision models, sometimes called System One models, emerged as a distinct category following the launch of Jev by TypeSafe AI last month. While a standard large language model can produce arbitrary output, a decision model restricts itself to picking between options or rating items on a scale.

Strands Decider supports three specific question types:

  • choice: selecting one option from a set of N.
  • noul: a yes/no probability between 0 and 1.
  • score: a level on an ordered rubric.

Every response is drawn from the allowed options and includes a confidence metric. The team notes the model performs worse than reasoning models on complex problems. It is not suitable for coding, chat, or summarisation tasks.

Architecture: an LLM with its mouth removed

The team begins with Qwen3.5-2B-Base and removes the language modelling head. A small pointer head of roughly 1 million parameters replaces it. This head compares the hidden state at the <answer> position against the hidden state at the final token of each option. One forward pass yields the result, with no decoding loop.

The torso uses a rank-16 LoRA, while the head runs in fp32. Label sets come from the request, so nothing caps the option count. The released checkpoint is version 19.

Asking several questions about a single text is cheap. The state is read once, and each additional question adds only its own tokens.

Benchmarks and latency

The team measures accuracy and calibration on JevBench, a public benchmark for Jev-class models.

Published v19 figures:

  • JevBench v1 public accuracy: 0.723 (167 of 231 tasks).
  • Brier score 0.342, expected calibration error 0.052.
  • Tier accuracy: easy 1.000, standard 0.875, hard 0.505.
  • Latency on an RTX 3090: 115 ms median, 299 ms p95.
  • Latency on an M3 Pro: 153 ms warm median under 300 tokens.

On the v1.4.2 board of September 25, v19 ranked 3rd of 33 in the 2B class. Excluding three models just over 2B, it ranked 1st of 30. The repository also flags a caveat. Mapika’s newer decider-2b v11 scores 175 of 231 on the Strands harness, eight tasks ahead. Strands Decider was not on the newer v1.5.4 composite board at the time of writing.

Calibration is the practical win. On unseen short classification tasks, answers at 0.9 confidence or higher were correct about 95% of the time. The team advises confirming or escalating below that threshold.

How it compares

FeatureStrands Decider 2B (v19)Jev 1.13.0decider-2bDecision 2B
DeveloperStrands Agents (AWS)TypeSafe AIMapikaFlyMy.AI
AccessOpen weights, Apache-2.0Closed hosted APIOpen code or weightsOpen code or weights
Base modelQwen3.5-2B-Base + LoRA + pointer headUndisclosedQwen3.5-2B-Base + trained readoutMiniCPM5-2B + LoRA + pointer head
Size1.9BUndisclosed1.9B2.5B dense
Self-hostingYesNoYesYes
Full training recipe and data list publishedYesNoNot verifiedNot verified
JevBench public accuracy (v1.4.2 board)0.723Not in source table0.7100.753
Reported latency115 ms median (RTX 3090)70 to 500 ms (vendor)Not comparedNot compared

Sources: Strands JevBench comparison, Benchmark Heaven JevBench, Jev product page. Latencies come from different hardware and harnesses, so they are not directly comparable.

Use cases and a guardrail example

The team reports early success in model routing, tool selection, argument checking, triage, guardrails, evals, and hybrid agents. In a hybrid agent, an LLM makes the hard calls and the decider handles rote ones.

The repository example gates a weather tool call inside a Strands agent. A before_tool_call intervention asks two yes/no questions. Are the arguments grounded in what the user said? Is calling now premature? If the agent guessed a city, it asks the user instead.

From the CLI, routing “Help! My payouts have been failing for 3 days!” across billing, sales and retail returns billing with confidence 0.768.

What it means

For developers building agents, this tool offers a way to validate actions before they happen. Instead of relying on a model to guess correctly, you can force it to check specific conditions first. This reduces the risk of hallucinated tool calls and makes the system’s behaviour more predictable. The low latency allows this check to happen in real time without slowing down the user experience.

Scroll to Top