Skip to main content
Aggregate arXiv cs.AI 人工智能 26 Aug 2026 - 14:30

BenchBench-Protocol: Evaluating Real-World Wet-Lab Protocol Reasoning and Modification

RSS 官方收录 · 可信分层展示

关键摘要

arXiv:2608.…

  • 23898v1 Announce Type: new Abstract: We introduce BenchBench-Protocol,…
  • Adapting a published protocol to a new experiment is a routine task fo…
  • Recent life-science benchmarks have moved toward open-ended, rubric-gr…

摘要引擎:抽取

正文提要

arXiv:2608.23898v1 Announce Type: new Abstract: We introduce BenchBench-Protocol, a benchmark for large language models of 149 protocol-modification tasks recovered from modifications that scientists made to published protocols during real experimental work. Adapting a published protocol to a new experiment is a routine task for a wet-lab scientist, and a correct modification requires accounting for prior choices and downstream steps. Recent life-science benchmarks have moved toward open-ended, rubric-graded tasks, but tasks are typically elicited from experts rather than reconstructed from real-world modifications. BenchBench-Protocol tasks are derived from differences between a published protocol and a version a scientist modified, which provides the basis for the query and the weighted rubric elements for a correct response. The benchmark draws from 96 source protocols across nine domains of wet-lab biology and only includes tasks rated highly after review by domain experts. We evaluate nine closed and open models; Claude Opus 5 scores highest at 59.2% normalized rubric score, with other models between 34.1% and 47.1%, and the benchmark remains unsaturated when taking the best of ten attempts. As models are increasingly helpful in life-sciences research, evaluating them on routine wet-lab tasks becomes correspondingly important. We present BenchBench-Protocol as both a grounded assessment of wet-lab reasoning and evidence for the utility of real-world experiments to construct benchmark tasks.

来源:https://arxiv.org/abs/2608.23898

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表