Microsoft has released ThinkingBox, a benchmark that evaluates AI agents by the actual state they leave behind in business systems rather than simply judging their responses or tool calls.
Microsoft has released ThinkingBox, a benchmark that evaluates AI agents by the actual state they leave behind in business systems rather than simply judging their responses or tool calls.
The benchmark runs 507 stateful workflows 20 times per model, checking whether records were changed correctly and whether agents introduced missing or unintended side effects. Results show a wide gap between one-off capability and repeatable reliability, with GPT-6 Astra retaining 78% of its single-run success rate across repeated trials, the highest retention reported.