Haberlere Geri Dön
YenilikAI Understanding brifing

BenchBench, yapay zekanın gerçek dünyadaki ıslak laboratuvar protokollerini uyarlayıp uyarlayamayacağını test ediyor

Yeni bir kıyaslama, dil modellerinin gerçek deneysel durumlar için yayınlanmış ıslak laboratuvar prosedürlerini ne kadar iyi değiştirdiğini değerlendiriyor. En iyi performans gösteren model, karşılaştırma ölçütünün normalleştirilmiş değerlendirme tablosunda %59,2 puan aldı ve tekrarlanan denemelerden sonra test doymamış kaldı.

5 min readRead the primary source
Primary-source image accompanying BenchBench tests whether AI can adapt real-world wet-lab protocols
Birincil kaynak belgeKaynak kaydedildi
Yayıncı
arxiv.org
Kaynak bağlantısı
arxiv.orghttps://arxiv.org/abs/2608.23898
Kaynak türü
Birincil belge – doğrudan okuduğumuz resmi bir duyuru, belge, dosyalama veya birinci taraf sayfası.
Bağlam60 saniyede bunu anlayın

Buradan başlayın

Anahtar terimler

Karşılaştırma
Model performansını ölçmek ve karşılaştırmak için kullanılan standartlaştırılmış bir test veya veri kümesi.
Geri alma
Bir sorgu için bir bilgi kaynağından ilgili belgeleri veya kayıtları bulma.
İstemi
Üretken bir modele sağlanan girdi talimatları ve bağlam.
Kendinizi test edinYapay Zeka Modelleri Açıklaması Testi

Ne oldu?

Researchers introduced BenchBench-Protocol, a built from 149 real-world wet-lab protocol modifications. Across nine models, Claude Opus 5 achieved the highest reported normalized rubric score, but the benchmark exposed substantial room for improvement.

The arXiv paper introduces BenchBench-Protocol as a for large language models handling wet-lab protocol reasoning and modification. It contains 149 tasks reconstructed from differences between published protocols and versions that scientists modified during real experimental work. The authors say the benchmark draws on 96 source protocols covering nine wet-lab biology domains, and that the included tasks were rated highly after review by domain experts.

This design makes the ’s central object a concrete change to an existing procedure rather than a question generated solely by asking experts to imagine a challenge. The paper frames protocol adaptation as a routine but consequential laboratory task. A researcher modifying a published procedure must account for earlier choices and for steps that follow the change. BenchBench-Protocol uses those real modifications to form queries and to create weighted rubric elements for judging responses. The source presents this as a way to ground evaluation in the structure of real experiments, while retaining an open-ended format rather than reducing every answer to a single exact string. The benchmark therefore tests whether a model can preserve relevant dependencies across a procedure, not merely produce scientifically plausible language.

The authors evaluate nine closed and open models. Claude Opus 5 receives the highest reported result, a 59.2% normalized rubric score; the other models score between 34.1% and 47.1%, according to the abstract. The also remains unsaturated when the best of ten attempts is used, meaning repeated attempts do not appear to exhaust the available performance under the paper’s evaluation setup. The source does not identify every evaluated model in the abstract, explain the precise interpretation of the normalized score, or report that any model’s answer was used to carry out a successful experiment.

Kaynak ayrıntıları: arxiv.org ↗

Neden önemli?

AI systems are increasingly used to support life-sciences research, where seemingly small procedural changes can affect later steps and experimental outcomes. Testing models against modifications made during actual scientific work offers a more practical measure of reliability than evaluating only invented questions.

The research addresses a practical weakness in many AI evaluations for science: a model may appear capable on broad scientific questions while struggling with the detailed, dependent choices that make laboratory procedures work. By basing tasks on changes scientists actually made to published protocols, the gives evaluators a way to examine whether an AI system can reason through local modifications without losing track of downstream consequences. That is directly relevant to researchers using language models for planning, documentation, or procedural support.

The weighted rubric is also important because protocol changes can be partly correct while still omitting a detail that matters later. A response might identify the obvious substitution but fail to adjust a subsequent step, condition, quantity, or control. The source does not provide a breakdown of these error types, but its design is intended to make such omissions visible rather than rewarding answers solely for sounding fluent. In practical terms, this could help laboratories compare systems on the work they actually need assistance with, rather than relying on general-purpose model scores. The reported results point to capability gaps even among the strongest systems tested.

A top normalized score of 59.2% should not be read as a 59.2% success rate in laboratory experiments, because the abstract describes a rubric score rather than completed experimental outcomes. The spread between the highest and lower reported scores suggests that model choice can affect performance on this task, but the source does not establish which model characteristics explain the difference. It also does not show that performance predicts safety, reproducibility, or scientific validity outside the tested protocols.

Interactive Mechanism

İnteraktif Mekanizma: Aslında Nasıl Çalışıyor?

Bu gelişmenin arkasında yatan teknolojiyi etkileşimli olarak keşfedin.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
İnteraktif Konsept Kontrolü+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Bundan sonra ne izlenecek?

The next questions are whether the is independently reproduced, how models perform on individual protocol changes, and whether stronger scores translate into safer and more useful assistance in real laboratories. The source does not establish that any evaluated model can independently conduct experiments or produce validated procedures.

A key next step is detailed examination of the full and its scoring process. Readers should look for the distribution of tasks across the nine domains, examples of high- and low-scoring modifications, the identities and versions of all nine models, and the individual rubric elements that models most often miss. Those details would show whether the result reflects a broad limitation in protocol adaptation or a smaller set of recurring difficulties. The source’s abstract alone does not answer those questions.

Independent replication would help establish how stable the findings are. Useful checks would include rerunning the models, testing different formats, measuring variation across repeated attempts, and evaluating whether the same ranking holds on protocols withheld from construction. Because the benchmark remains unsaturated after ten attempts, future work should also clarify whether additional sampling, tools, , or human review materially changes performance. None of those outcomes is reported in the source.

The most consequential question is whether better scores correspond to better laboratory assistance. Future evaluations could compare model suggestions with expert-reviewed modifications and, where appropriate, controlled experimental validation. Such work would need to distinguish textual correctness from actual experimental success and would need safeguards around errors. For now, BenchBench-Protocol provides evidence about model responses to a defined set of real-world-derived tasks, not evidence that these systems can safely modify or execute wet-lab procedures without qualified human oversight.

İlgili kılavuzlar ve testler

Yapay Zeka Modellerinin AçıklamasıYapay Zeka EğitimiChatGPT ve LLM'lerYapay Zeka EtiğiBildiklerinizi test edin; ücretsiz bir yapay zeka testini deneyinSözlüğümüzde bir yapay zeka terimine bakınAI modeli sürüm izleyicisini takip edin
Bunu yararlı buldunuz mu?