返回新闻
创新AI Understanding 简报

BenchBench 测试人工智能是否可以适应现实世界的湿实验室协议

一个新的基准评估语言模型如何针对真实实验情况修改已发布的湿实验室程序。表现最好的模型在基准标准化评分标准上得分为 59.2%,并且在多次尝试后测试仍然不饱和。

5 min readRead the primary source
Primary-source image accompanying BenchBench tests whether AI can adapt real-world wet-lab protocols
主要来源文件来源记录
出版商
arxiv.org
来源链接
arxiv.orghttps://arxiv.org/abs/2608.23898
来源类型
主要文件——我们直接阅读的官方公告、文件、文件或第一方页面。
背景60 秒内了解这一点

从这里开始

关键术语

基准测试
用于测量和比较模型性能的标准化测试或数据集。
检索
从知识源中查找相关文档或记录以进行查询。
迅速的
提供给生成模型的输入指令和上下文。
测试一下自己AI 模型解释测验

发生了什么

Researchers introduced BenchBench-Protocol, a built from 149 real-world wet-lab protocol modifications. Across nine models, Claude Opus 5 achieved the highest reported normalized rubric score, but the benchmark exposed substantial room for improvement.

The arXiv paper introduces BenchBench-Protocol as a for large language models handling wet-lab protocol reasoning and modification. It contains 149 tasks reconstructed from differences between published protocols and versions that scientists modified during real experimental work. The authors say the benchmark draws on 96 source protocols covering nine wet-lab biology domains, and that the included tasks were rated highly after review by domain experts.

This design makes the ’s central object a concrete change to an existing procedure rather than a question generated solely by asking experts to imagine a challenge. The paper frames protocol adaptation as a routine but consequential laboratory task. A researcher modifying a published procedure must account for earlier choices and for steps that follow the change. BenchBench-Protocol uses those real modifications to form queries and to create weighted rubric elements for judging responses. The source presents this as a way to ground evaluation in the structure of real experiments, while retaining an open-ended format rather than reducing every answer to a single exact string. The benchmark therefore tests whether a model can preserve relevant dependencies across a procedure, not merely produce scientifically plausible language.

The authors evaluate nine closed and open models. Claude Opus 5 receives the highest reported result, a 59.2% normalized rubric score; the other models score between 34.1% and 47.1%, according to the abstract. The also remains unsaturated when the best of ten attempts is used, meaning repeated attempts do not appear to exhaust the available performance under the paper’s evaluation setup. The source does not identify every evaluated model in the abstract, explain the precise interpretation of the normalized score, or report that any model’s answer was used to carry out a successful experiment.

来源详情: arxiv.org ↗

为什么这很重要

AI systems are increasingly used to support life-sciences research, where seemingly small procedural changes can affect later steps and experimental outcomes. Testing models against modifications made during actual scientific work offers a more practical measure of reliability than evaluating only invented questions.

The research addresses a practical weakness in many AI evaluations for science: a model may appear capable on broad scientific questions while struggling with the detailed, dependent choices that make laboratory procedures work. By basing tasks on changes scientists actually made to published protocols, the gives evaluators a way to examine whether an AI system can reason through local modifications without losing track of downstream consequences. That is directly relevant to researchers using language models for planning, documentation, or procedural support.

The weighted rubric is also important because protocol changes can be partly correct while still omitting a detail that matters later. A response might identify the obvious substitution but fail to adjust a subsequent step, condition, quantity, or control. The source does not provide a breakdown of these error types, but its design is intended to make such omissions visible rather than rewarding answers solely for sounding fluent. In practical terms, this could help laboratories compare systems on the work they actually need assistance with, rather than relying on general-purpose model scores. The reported results point to capability gaps even among the strongest systems tested.

A top normalized score of 59.2% should not be read as a 59.2% success rate in laboratory experiments, because the abstract describes a rubric score rather than completed experimental outcomes. The spread between the highest and lower reported scores suggests that model choice can affect performance on this task, but the source does not establish which model characteristics explain the difference. It also does not show that performance predicts safety, reproducibility, or scientific validity outside the tested protocols.

Interactive Mechanism

互动机制:它实际上是如何运作的

以交互方式探索这一发展背后的基础技术。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
交互式概念检查+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

接下来看什么

The next questions are whether the is independently reproduced, how models perform on individual protocol changes, and whether stronger scores translate into safer and more useful assistance in real laboratories. The source does not establish that any evaluated model can independently conduct experiments or produce validated procedures.

A key next step is detailed examination of the full and its scoring process. Readers should look for the distribution of tasks across the nine domains, examples of high- and low-scoring modifications, the identities and versions of all nine models, and the individual rubric elements that models most often miss. Those details would show whether the result reflects a broad limitation in protocol adaptation or a smaller set of recurring difficulties. The source’s abstract alone does not answer those questions.

Independent replication would help establish how stable the findings are. Useful checks would include rerunning the models, testing different formats, measuring variation across repeated attempts, and evaluating whether the same ranking holds on protocols withheld from construction. Because the benchmark remains unsaturated after ten attempts, future work should also clarify whether additional sampling, tools, , or human review materially changes performance. None of those outcomes is reported in the source.

The most consequential question is whether better scores correspond to better laboratory assistance. Future evaluations could compare model suggestions with expert-reviewed modifications and, where appropriate, controlled experimental validation. Such work would need to distinguish textual correctness from actual experimental success and would need safeguards around errors. For now, BenchBench-Protocol provides evidence about model responses to a defined set of real-world-derived tasks, not evidence that these systems can safely modify or execute wet-lab procedures without qualified human oversight.

相关指南和测验

人工智能模型解释人工智能培训ChatGPT 与大语言模型AI 伦理测试你所知道的——尝试免费的人工智能测验在我们的词汇表中查找人工智能术语关注 AI 模型发布跟踪器
觉得这有用吗?