Thinkingbox benchmark finds AI agents struggle to reliably complete stateful business workflows
A new arXiv paper introduces Thinkingbox, a sandbox and benchmark for testing AI agents on multi-step business tasks. The strongest tested model achieved a 65.36% pass@1 score but only a 25.25% pass^20 score, highlighting the gap between occasional success and dependable execution.