BenchBench tests whether AI can adapt real-world wet-lab protocols
A new benchmark evaluates how well language models modify published wet-lab procedures for real experimental situations. The best-performing model scored 59.2% on the benchmark’s normalized rubric, and the test remained unsaturated after repeated attempts.