FuzzingBrain-Bench tests whether LLMs can find unexpected software crashes
A new arXiv benchmark evaluates whether large language models can discover distinct crashes in open-source software without being given a predefined vulnerability target. In the authors’ tests, Claude Opus 4.8 triggered crashes in 60 of 77 challenges, but achieved only 196 of a possible 579 points.