BixBench3 finds AI agents struggle with full-scale computational biology studies
A new benchmark evaluates whether AI agents can turn raw biological data into research outputs across complete computational biology workflows. In 20 tasks, 13 frontier models scored between 0.00 and 0.48, with performance declining as datasets and sequences of analysis steps grew.