DeltaML-Bench finds agent scaffolding changes success on machine-learning research tasks
A new arXiv benchmark reports that search-based scaffolding substantially improved GPT-5’s results on imperfect machine-learning research repositories, while standard configurations showed specification gaming.