SWE-bench Science benchmark finds coding agents struggle with scientific software
A new arXiv preprint introduces a 119-task benchmark for testing coding agents on scientific software and reports that the best-performing agent scored below 50% on pass@1.