In this article
A customer’s kitchen appliance sits stuck in a courier exception at a Nashville distribution centre, fifteen days past its delivery date. The AI agent processes nine tool calls, checks the refund policy, and closes the ticket as resolved. The database disagrees. The required end state is on hold, and the customer receives no real answer.
That gap is what ThinkingBox measures. Across 507 stateful business workflows, each run 20 times against various LLM models, the benchmark grades agents on terminal backend state and side effects. This post covers what we found, what consistency costs, and how to run the benchmark yourself through OpenEnv.
The example above comes from a benchmark task sandbox. The executable check that fails is a single field: the ticket’s status is solved where the required end state is hold. The full trace is in Appendix D.4, Case 3 of our paper.
Contents
- A tool call is not an outcome
- One success is not reliability
- Can you depend on the model behind your agent?
- What consistency costs
- Failure signatures
- How it works
- Run it yourself
- Where this goes next
Want to try it before reading the results? Skip to section Run it yourself.
A tool call is not an outcome
Final responses and valid tool calls are only proxies. An agent can sound correct while leaving the wrong value, changing the wrong record, or creating an extra side effect. Only the records it leaves behind settle the question.
The gap is substantial. In a common-set ablation covering 121,680 valid trials across 12 LLM models, 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error. Executable checks nevertheless found wrong field values in 77.61% of them, unintended extra effects in 43.30%, and missing required effects in 25.36%. Those state-check findings overlap.
A trajectory is a claim. Database state is the evidence. Repetition is the trust test.
One success is not reliability
An agent that processes a refund correctly once and mishandles it the next four times is not a working refund agent. So every task runs 20 independent times, each from an identical clean backend, and we report three different things:
Table 1: The three numbers we report, and the question each one answers
| Metric | What it measures | What it answers |
|---|---|---|
| pass@1 | Share of all attempts that succeeded | How does it usually do? |
| pass@20 | Share of tasks solved at least once in 20 tries | Can it ever do this? Breadth. |
| Observed 20/20 | Tasks that actually passed all 20 recorded attempts | Can it always be correct? |
We use observed 20/20 in this blog post as the literal count of how many of the 507 tasks passed 20 out of 20. No estimator, no smoothing.
The table below reports pass@1, the single-attempt score estimate, broken out by domain. This is the number most leaderboards publish, and on its own it reads like an ordinary capability ranking.
Table 2: ThinkingBox-Bench pass@1 (%) by domain
| Model | Retail (98) | Auto insurance (100) | Travel (104) | Neobank (104) | Consulting (101) | Overall, task-weighted (507) |
|---|---|---|---|---|---|---|
| Proprietary models | ||||||
| Claude Opus 5.5 | 80.97 | 68.40 | 54.28 | 71.25 | 61.58 | 67.16 |
| Claude Opus 5 | 80.71 | 65.80 | 49.95 | 70.62 | 66.19 | 66.50 |
| GPT-5.4 | 76.33 | 62.65 | 68.12 | 65.34 | 54.60 | 65.36 |
| GPT-5.6 Sol | 67.65 | 65.30 | 60.34 | 59.09 | 57.52 | 61.91 |
| Claude Sonnet 4.6 | 72.35 | 54.40 | 58.94 | 56.39 | 54.31 | 59.19 |
| GPT-6 Astra | 71.73 | 46.55 | 55.87 | 60.87 | 56.83 | 58.31 |
| GPT-5.2 | 70.20 | 22.40 | 53.70 | 51.15 | 34.06 | 46.28 |
| Claude Opus 4.6 | 68.62 | 8.30 | 21.11 | 35.67 | 27.82 | 32.09 |
| o3-pro | 37.70 | 2.95 | 17.31 | 24.28 | 14.60 | 19.31 |
| Grok-4.3 | 43.93 | 2.60 | 15.14 | 1.78 | 9.55 | 14.38 |
| Open-weight models | ||||||
| Kimi-K3 | 82.24 | 50.80 | 61.83 | 41.35 | 51.63 | 57.37 |
| Qwen3.8-27B | 64.03 | 47.85 | 53.41 | 47.88 | 45.69 | 51.70 |
| DeepSeek-V4-Pro | 68.21 | 29.65 | 43.13 | 44.86 | 31.04 | 43.26 |
| Kimi-K2.6 | 53.72 | 24.50 | 39.52 | 33.65 | 37.33 | 37.66 |
| GLM-5.1 | 58.67 | 25.70 | 35.43 | 13.27 | 34.06 | 33.19 |
| Qwen3.6-27B | 43.11 | 29.00 | 46.39 | 27.84 | 18.37 | 32.94 |
| Qwen3.5-9B | 19.90 | 0.70 | 4.71 | 1.15 | 2.33 | 5.65 |
| Mistral-Large-3 | 11.28 | 1.30 | 8.99 | 1.15 | 0.74 | 4.66 |
Claude Opus 5.5 leads overall at 67.16%, two-thirds of a point above Claude Opus 5. Kimi-K3 is the strongest open-weights model, within a point of GPT-6-Astra. Domain matters just as much: Claude Opus 4.6 scores 68.62% on retail




