BenchMIRT finds LLM benchmarks can mix safety and reasoning signals
Ai2 introduces BenchMIRT, a method that analyzes individual benchmark questions to identify which capabilities they measure. Testing 100 LLMs across 16 benchmarks, the researchers found that some evaluations combine safety and general-reasoning signals and that smaller question sets can often preserve much of the…