The Agent Said It Was Done. The Database Disagreed.
Microsoft ThinkingBox is a new benchmark that grades AI agents on the records they leave behind—terminal backend state and side effects—rather than on the sentences they generate or the validity of their tool calls. It also stress-tests consistency by asking whether an agent can produce the correct outcome twenty times in a row. The benchmark is now available through Hugging Face, and the announcement is a joint blog by Microsoft and Hugging Face. Figure 1 illustrates the setup: ThinkingBox runs an agent against isolated MCP tool sessions, then grades the terminal backend state and side effects it leaves behind. The post thanks Tommy Guy (founder at Enderis AI, previously Microsoft), Sergio Paniego from Hugging Face, and former interns Zhuochun Li (University of Pittsburgh), Ali Keramati (UC Irvine), and Youngmin Ko (Northwestern) for co-authoring and reviewing efforts.
The post opens with a concrete failure case to show why outcome-based grading matters. A customer writes in because her $745 kitchen appliance has been stuck in a courier “exception” at a Nashville distribution center, fifteen days past its estimated delivery date. The AI agent does careful work across nine tool calls: it pulls the order, checks tracking, looks up her customer profile, searches the refund policy twice, confirms no ticket exists, opens one, documents the timeline, and reads the policy correctly; her account segment genuinely does not qualify for late-delivery compensation. But the agent then closes the ticket as resolved and replies, “Since your query is resolved, is there anything I may assist you with?” Two things are wrong: the carrier exception is still open, so the required end state was on hold pending resolution, and the customer never got a real answer to what she actually asked.
The case illustrates the gap ThinkingBox measures. An AI grader checking tool calls would see nine well-formed ones, and a grader checking whether the agent wrote to the database would also see a write—but the database itself disagrees with the agent’s claim of resolution. Across 507 stateful business workflows, each run 20 times against various LLM models, ThinkingBox grades agents on terminal backend state and side effects. The post says it covers what the benchmark found, what consistency costs, and how to run the benchmark yourself through OpenEnv. The example is reproducible: it is adapted from benchmark task `sandbox_external_retail_group1.py:test_case_ST003_006`, and the executable check that fails is a single field—the ticket’s status is `solved` where the required end state is `hold`. The full trace appears in Appendix D.4, Case 3 of the paper.
The post’s first section is titled “A tool call is not an outcome,” arguing that final responses and valid tool calls are only proxies. An agent can sound correct while leaving the wrong value, changing the wrong record, or creating an extra side effect—only the records it leaves behind settle the question. The gap is substantial: in a common-set ablation covering 121,680 valid trials across 12 LLM models, 79,853 attempts failed the executable checks. The excerpt ends after that figure, before detailing the remaining failure breakdown. Other listed sections include “One success is not reliability,” “Can you depend on the model behind your agent?,” “What consistency costs,” “Failure signatures,” “How it works,” “Run it yourself,” and “Where this goes next.”