protocol-reasoning

Tag

Cards List
#protocol-reasoning

BenchBench-Protocol: Evaluating Real-World Wet-Lab Protocol Reasoning and Modification

arXiv cs.AI · 2026-08-26 Cached

The article introduces BenchBench-Protocol, a benchmark of 149 real-world wet-lab protocol-modification tasks to evaluate large language models' reasoning, with Claude Opus 5 scoring highest at 59.2%.

0 favorites 0 likes
← Back to home

Submit Feedback