The Agent Said It Was Done. The Database Disagreed.

In this articleContentsA tool call is not an outcomeOne success is not reliability A customer’s kitchen appliance sits stuck in a courier…

By Vane October 3, 2026 3 min read
The Agent Said It Was Done. The Database Disagreed.


A customer’s kitchen appliance sits stuck in a courier exception at a Nashville distribution centre, fifteen days past its delivery date. The AI agent processes nine tool calls, checks the refund policy, and closes the ticket as resolved. The database disagrees. The required end state is on hold, and the customer receives no real answer.

That gap is what ThinkingBox measures. Across 507 stateful business workflows, each run 20 times against various LLM models, the benchmark grades agents on terminal backend state and side effects. This post covers what we found, what consistency costs, and how to run the benchmark yourself through OpenEnv.

The example above comes from a benchmark task sandbox. The executable check that fails is a single field: the ticket’s status is solved where the required end state is hold. The full trace is in Appendix D.4, Case 3 of our paper.

Contents

  • A tool call is not an outcome
  • One success is not reliability
  • Can you depend on the model behind your agent?
  • What consistency costs
  • Failure signatures
  • How it works
  • Run it yourself
  • Where this goes next

Want to try it before reading the results? Skip to section Run it yourself.

A tool call is not an outcome

Final responses and valid tool calls are only proxies. An agent can sound correct while leaving the wrong value, changing the wrong record, or creating an extra side effect. Only the records it leaves behind settle the question.

The gap is substantial. In a common-set ablation covering 121,680 valid trials across 12 LLM models, 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error. Executable checks nevertheless found wrong field values in 77.61% of them, unintended extra effects in 43.30%, and missing required effects in 25.36%. Those state-check findings overlap.

A trajectory is a claim. Database state is the evidence. Repetition is the trust test.

One success is not reliability

An agent that processes a refund correctly once and mishandles it the next four times is not a working refund agent. So every task runs 20 independent times, each from an identical clean backend, and we report three different things:

Table 1: The three numbers we report, and the question each one answers

MetricWhat it measuresWhat it answers
pass@1Share of all attempts that succeededHow does it usually do?
pass@20Share of tasks solved at least once in 20 triesCan it ever do this? Breadth.
Observed 20/20Tasks that actually passed all 20 recorded attemptsCan it always be correct?

We use observed 20/20 in this blog post as the literal count of how many of the 507 tasks passed 20 out of 20. No estimator, no smoothing.

The table below reports pass@1, the single-attempt score estimate, broken out by domain. This is the number most leaderboards publish, and on its own it reads like an ordinary capability ranking.

Table 2: ThinkingBox-Bench pass@1 (%) by domain

ModelRetail (98)Auto insurance (100)Travel (104)Neobank (104)Consulting (101)Overall, task-weighted (507)
Proprietary models
Claude Opus 5.580.9768.4054.2871.2561.5867.16
Claude Opus 580.7165.8049.9570.6266.1966.50
GPT-5.476.3362.6568.1265.3454.6065.36
GPT-5.6 Sol67.6565.3060.3459.0957.5261.91
Claude Sonnet 4.672.3554.4058.9456.3954.3159.19
GPT-6 Astra71.7346.5555.8760.8756.8358.31
GPT-5.270.2022.4053.7051.1534.0646.28
Claude Opus 4.668.628.3021.1135.6727.8232.09
o3-pro37.702.9517.3124.2814.6019.31
Grok-4.343.932.6015.141.789.5514.38
Open-weight models
Kimi-K382.2450.8061.8341.3551.6357.37
Qwen3.8-27B64.0347.8553.4147.8845.6951.70
DeepSeek-V4-Pro68.2129.6543.1344.8631.0443.26
Kimi-K2.653.7224.5039.5233.6537.3337.66
GLM-5.158.6725.7035.4313.2734.0633.19
Qwen3.6-27B43.1129.0046.3927.8418.3732.94
Qwen3.5-9B19.900.704.711.152.335.65
Mistral-Large-311.281.308.991.150.744.66

Claude Opus 5.5 leads overall at 67.16%, two-thirds of a point above Claude Opus 5. Kimi-K3 is the strongest open-weights model, within a point of GPT-6-Astra. Domain matters just as much: Claude Opus 4.6 scores 68.62% on retail

Scroll to Top