Tag
The article introduces BenchBench-Protocol, a benchmark of 149 real-world wet-lab protocol-modification tasks to evaluate large language models' reasoning, with Claude Opus 5 scoring highest at 59.2%.