K-Bench: measuring model performance on real scientific agent requests

arXiv cs.AI Papers

Summary

The paper introduces K-Bench 01, a benchmark for evaluating AI agents on real scientific requests, revealing that no model consistently meets the threshold for acceptable performance, with overclaiming as a common failure.

arXiv:2608.21601v1 Announce Type: new Abstract: Benchmarks for scientific artificial intelligence are mostly written to be scored: multiple-choice questions, curated agent tasks with reference solutions, or simulators with a known generative structure. Real scientific requests arrive differently. They are underspecified, they carry attachments, and lack ground truth. We report K-Bench 01, an evaluation built from first-turn requests sampled from live user traffic on K-Dense Web and run end to end by nine frontier models in identical sandboxes, yielding 1,602 completed agent runs. Three blinded language-model judges scored every run against an eight-dimension rubric. On a rubric whose 8-anchor instructs judges that a domain scientist would accept the work with minor edits, no model clears the line under all three judges. gpt-5.6-sol has the highest pooled mean, 8.04, but its 95% interval [7.80, 8.23] spans the threshold, and two of the three judges rank claude-opus-5 first instead. We therefore report the ordering of systems as the reproducible quantity, the absolute level as an attribute of the instrument, and the top of the table as unresolved. Across all 39,934 scored judgments -- the eight dimension scores plus a holistic overall for each assessment, excluding not-applicable cells -- 47.6% fall below the 8-point threshold. Difficulty is not uniform across the rubric: scientific accuracy averages 6.22 against 7.33 for communication, on identical denominators and in the same direction within every one of the nine models. The single leading failure tag is overclaiming, on 31.4% of assessments. We argue that the informative quantity for scientific agents is not a leaderboard position but the joint distribution of what was delivered, what was claimed, and what artifacts were produced.
Original Article
View Cached Full Text

Cached at: 08/25/26, 04:20 AM

# measuring model performance on real scientific agent requests
Source: [https://arxiv.org/html/2608.21601](https://arxiv.org/html/2608.21601)
Darshil PatelAffiliation:K\-Dense, Inc\.Yuhuan HeAffiliation:K\-Dense, Inc\.Timothy KassisAffiliation:K\-Dense, Inc\.Affiliation:Corresponding author: timothy\.kassis@k\-dense\.ai

###### Abstract

Benchmarks for scientific artificial intelligence are mostly written to be scored: multiple\-choice questions, curated agent tasks with reference solutions, or simulators with a known generative structure\. Real scientific requests arrive differently\. They are underspecified, they carry attachments, and lack ground truth\. We report K\-Bench 01, an evaluation built from first\-turn requests sampled from live user traffic on K\-Dense Web and run end to end by nine frontier models in identical sandboxes, yielding 1,602 completed agent runs\. Three blinded language\-model judges scored every run against an eight\-dimension rubric\. On a rubric whose 8\-anchor instructs judges that a domain scientist would accept the work with minor edits, no model clears the line under all three judges\.gpt\-5\.6\-solhas the highest pooled mean, 8\.04, but its 95% interval \[7\.80, 8\.23\] spans the threshold, and two of the three judges rankclaude\-opus\-5first instead\. We therefore report the ordering of systems as the reproducible quantity, the absolute level as an attribute of the instrument, and the top of the table as unresolved\. Across all 39,934 scored judgments — the eight dimension scores plus a holistic overall for each assessment, excluding not\-applicable cells — 47\.6% fall below the 8\-point threshold\. Difficulty is not uniform across the rubric: scientific accuracy averages 6\.22 against 7\.33 for communication, on identical denominators and in the same direction within every one of the nine models\. The single leading failure tag is overclaiming, on 31\.4% of assessments\. We argue that the informative quantity for scientific agents is not a leaderboard position but the joint distribution of what was delivered, what was claimed, and what artifacts were produced\.

![Refer to caption](https://arxiv.org/html/2608.21601v1/figures/graphical_abstract.png)Figure 1:Graphical abstract\.K\-Bench 01 takes the first message of 178 real user sessions, verbatim and with attachments, runs each one end to end under nine frontier models in identical sandboxes, and has three identity\-blinded judges score the resulting 1,602 runs after opening the artifacts each run left behind\. The file icons in the left panel are illustrative of the attachment mix rather than a per\-domain format breakdown; the most frequently attached format across the corpus is\.docx\(Table[30](https://arxiv.org/html/2608.21601#A1.T30)\)\. Tile values are the headline results: the best model mean \(8\.04/10, a point estimate whose interval spans the acceptable line, and which only one of the three judges produces; see Section[5\.4](https://arxiv.org/html/2608.21601#S5.SS4), and Section[5\.2](https://arxiv.org/html/2608.21601#S5.SS2)for why the count of models reaching 8 is zero, one or two depending on the judge\), the share of the 39,934 scored judgments below the acceptable line \(47\.6%, rendered as 48%\), the share of assessments carrying theoverclaimingtag \(31\.4%\), and the share of tasks no model solved \(12\.4%, rendered as 12%\)\. The strip beneath contrasts the mean of the execution dimensions \(6\.93\) with the mean of the substance dimensions \(6\.34\) and Section[4\.2](https://arxiv.org/html/2608.21601#S4.SS2)gives the sharper and denominator\-matched version, namely that scientific accuracy trails communication by 1\.11 points within every one of the nine models\.## 1 Introduction

A scientist sends an agent a count matrix from an RNA\-seq experiment and asks which genes are differentially expressed, and whether the batch effect is real\. Another attaches three papers and asks whether an effect replicates\. These requests vary in shape; they are the first message of a real working session and may arrive with files, often specify the goal partly, and typically lack a validated reference answer\.

Today, few benchmarks used to track progress in scientific artificial intelligence have this shape, and for a defensible reason: a benchmark has to be scorable\. This has yielded exam\-style suites\. Each design choice leaves a gap where the typical real\-world request lives\. The gap is a construct\-validity problem\([Bean et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib3)\)\.

To date, the K\-Dense Web platform has processed over 75,000 interactions across approximately 18,000 user sessions\. We utilized 178 first\-turn requests from live K\-Dense Web traffic\([K\-Dense Inc\. 2026](https://arxiv.org/html/2608.21601#bib.bib22)\)and ran each one under nine models in an identical stock harness \(Section[3](https://arxiv.org/html/2608.21601#S3)\)\. Figure[1](https://arxiv.org/html/2608.21601#S0.F1)summarizes the pipeline and the headline results\. The design question is whether a traffic\-derived set still separates frontier systems, and on which axes\.

Every score in this paper is awarded by a panel of three language\-model judges applying a written rubric to a run’s transcript and to the files it left on disk\. The measured quantity is therefore panel\-assessed rubric compliance, and Section[5\.2](https://arxiv.org/html/2608.21601#S5.SS2)sets out what that does and does not establish\. In short: the panel’s ordering of systems reproduces across judges, and the level at which it places the scale does not\.

Our results show that the axes are not the ones a capability\-first reading would predict\. These findings inform the two framing commitments\. The first is that eloquence is not a deliverable; an evaluation that cannot see the difference between prose and an empty directory will systematically overstate progress\([Si et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib53);[Si et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib54)\)\. The second is that a single leaderboard number is the wrong summary of an agentic scientific benchmark\.

This paper contributes the following\.

- •We construct a benchmark from unmodified deployment traffic: 178 first\-turn scientific requests taken verbatim from live users, with their attachments, without reference answers \(Section[3](https://arxiv.org/html/2608.21601#S3)\)\.
- •We grade the files a run left behind rather than its prose alone\. Judges receive read\-only access to the run’s output tree \(Section[3](https://arxiv.org/html/2608.21601#S3)\)\.
- •We run a balanced campaign at scale: nine frontier models on identical tasks in identical sandboxes, 1,602 runs, three blinded judges, 4,806 assessments and 39,934 scored judgments \(Section[4](https://arxiv.org/html/2608.21601#S4)\)\.
- •We treat the judging panel as an object of study rather than as an instrument, separating what it reproduces from what it asserts, and showing that the count of models clearing the rubric’s threshold is panel\-dependent \(Section[5](https://arxiv.org/html/2608.21601#S5)\)\.
- •We report where the deficit sits: scientific accuracy trails communication within every model in the field, overclaiming is the leading failure tag at 31\.4% of assessments, and 47\.9% of runs finish with no file on disk \(Sections[4\.2](https://arxiv.org/html/2608.21601#S4.SS2)–[4\.7](https://arxiv.org/html/2608.21601#S4.SS7)\)\.

## 2 Related work: the 2020–2026 evaluation landscape

Evaluation of scientific artificial intelligence has moved through three overlapping generations in six years\. A useful way to read the field is by what each generation treats as the object of measurement \(Figure[2](https://arxiv.org/html/2608.21601#S2.F2)\)\. The first generation measures recalled and reasoned\-about knowledge\. The second measures written procedure\. The third and current generation measures executed work, and it is only in the third that the question of artifact quality becomes askable at all\.

![Refer to caption](https://arxiv.org/html/2608.21601v1/figures/landscape_timeline.png)Figure 2:Representative benchmarks for scientific and agentic capability, 2020–2026, grouped by what they measure\.Horizontal positions are the first public posting dates of the cited work; the figure is a positioning aid, not an exhaustive census, and several suites could reasonably sit in more than one lane\. Several suites discussed below are omitted here for legibility; Table[1](https://arxiv.org/html/2608.21601#S2.T1)gives the fuller comparison\. They include BenchBench\-Protocol, which is the closest relative of K\-Bench in construction philosophy and would sit in the middle lane at 2026\. K\-Bench 01 is placed at the right\. Building a benchmark out of deployment traffic is not itself new: WildBench and Arena\-Hard curate items from chat logs and RealClawBench reconstructs developer\-agent sessions\([Lin et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib30);[Li et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib28);[Lv et al\. 2026](https://arxiv.org/html/2608.21601#bib.bib38)\), and within science AstaBench is the nearest precedent, with problems inspired by requests to its deployed agents\. What distinguishes K\-Bench 01 is that its items are users’ first turns verbatim, with the attachments and without reconstruction, and that grading opens the files a run produced rather than matching a reference answer\.### 2\.1 Knowledge and exam suites

The first generation established the discipline of reproducible scoring\. MMLU set the template of large multiple\-choice coverage\([Hendrycks et al\. 2020](https://arxiv.org/html/2608.21601#bib.bib16)\), BIG\-bench pushed breadth to hundreds of tasks\([Srivastava et al\. 2022](https://arxiv.org/html/2608.21601#bib.bib60)\), and HELM formalized multi\-metric reporting across scenarios\([Liang et al\. 2022](https://arxiv.org/html/2608.21601#bib.bib29)\)\. In science specifically, SciBench targets college\-level problem solving with worked numerical answers\([Wang et al\. 2023](https://arxiv.org/html/2608.21601#bib.bib65)\), SciEval builds a multi\-level scientific evaluation spanning knowledge and reasoning\([Sun et al\. 2023](https://arxiv.org/html/2608.21601#bib.bib62)\), and GPQA raises the ceiling with graduate\-level questions written to resist search\([Rein et al\. 2023](https://arxiv.org/html/2608.21601#bib.bib49)\)\. Humanity’s Last Exam extends the same logic to the frontier of closed\-form difficulty\([Phan et al\. 2026](https://arxiv.org/html/2608.21601#bib.bib47)\)\. These suites remain the right instrument for the property they measure, with their weaknesses well documented from inside the field: benchmark choice itself shapes conclusions\([Dehghani et al\. 2021](https://arxiv.org/html/2608.21601#bib.bib9)\), contamination is measurably present in public benchmark items\([Golchin and Surdeanu 2023](https://arxiv.org/html/2608.21601#bib.bib13)\), and leaderboard dynamics can reward selective reporting rather than capability\([Singh et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib56)\)\. The general form of the objection is construct validity\. A systematic review of 445 benchmarks finds the measured phenomenon, the task and the scoring metric routinely coming apart\([Bean et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib3)\), which is what the older warning against treating a single suite as a general measure of progress looks like when it is applied to language models\([Raji et al\. 2021](https://arxiv.org/html/2608.21601#bib.bib48)\)\.

### 2\.2 Expert\-authored research skills and laboratory procedure

The second generation moves the content toward how practicing scientists work\. LAB\-Bench assembles biology\-research tasks covering literature reasoning, protocol comprehension, sequence manipulation and figure interpretation\([Laurent et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib24)\), and LABBench2 revises the suite with free\-response tasks set in more realistic contexts\([Laurent et al\. 2026](https://arxiv.org/html/2608.21601#bib.bib25)\)\. LifeSciBench evaluates language models on expert\-level life\-science tasks with free\-response rubrics rather than option keys\([Liu et al\. 2026a](https://arxiv.org/html/2608.21601#bib.bib31)\)\. Procedure\-focused work is a distinct and clean sub\-field: BioPlanner automates evaluation of protocol planning\([O’Donoghue et al\. 2023](https://arxiv.org/html/2608.21601#bib.bib45)\), BioLP\-bench measures understanding of laboratory protocols by injecting and detecting errors\([Ivanov 2024](https://arxiv.org/html/2608.21601#bib.bib19)\), BioProBench scales protocol reasoning to a corpus\-and\-benchmark pairing\([Liu et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib35)\), ChemReason\-Bench does the analogous work for experimental chemistry\([Zhang et al\. 2026](https://arxiv.org/html/2608.21601#bib.bib69)\), and PhysDox audits whether proposed physiological\-sensing protocols are physically feasible\([Liu et al\. 2026b](https://arxiv.org/html/2608.21601#bib.bib32)\)\.

Benchling’s BenchBench\-Protocol is the closest relative of K\-Bench in construction philosophy to date\([Sivakumar et al\. 2026](https://arxiv.org/html/2608.21601#bib.bib57)\)\. Rather than authoring items, it recovers them from real edits that scientists made to published wet\-lab protocols, so the task distribution is inherited from practice instead of invented for measurement\. The key difference is scope and container: BenchBench\-Protocol stays inside protocol reasoning and modification with weighted rubrics over free\-response answers, while K\-Bench takes the whole heterogeneous distribution of what users send an agent\. Both designs accept the same tradeoff of a clean reference key for fidelity to real scientific work\.

### 2\.3 Agents, code and discovery

The third generation grades execution\([M\. Bran et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib39);[Boiko et al\. 2023](https://arxiv.org/html/2608.21601#bib.bib4);[Lu et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib36)\)\. General agentic suites established the harness conventions: AgentBench for multi\-environment agent evaluation\([Liu et al\. 2023a](https://arxiv.org/html/2608.21601#bib.bib33)\), GAIA for assistant tasks that require tool use\([Mialon et al\. 2023](https://arxiv.org/html/2608.21601#bib.bib42)\), SWE\-bench for repository\-scale software engineering\([Jimenez et al\. 2023](https://arxiv.org/html/2608.21601#bib.bib20)\), andτ\\tau\-bench for tool–agent–user interaction\([Yao et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib67)\)\. Terminal\-Bench moved the container into the specification, scoring agents on hard tasks inside command\-line environments of the kind this campaign uses\([Merrill et al\. 2026](https://arxiv.org/html/2608.21601#bib.bib41)\), and the Holistic Agent Leaderboard makes the harness itself a first\-class variable, running 21,730 rollouts across models, scaffolds and benchmarks to show how much of a reported result belongs to the scaffold rather than to the model\([Kapoor et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib23)\)\. Data\-analysis suites narrowed this to work that resembles the analytical core of science, and they differ in what they grade: InfiAgent\-DABench converts open\-ended analysis questions into a closed\-form format so answers can be checked automatically\([Hu et al\. 2024a](https://arxiv.org/html/2608.21601#bib.bib17)\), DSBench scores end\-to\-end analysis and modeling deliverables drawn from data\-science competitions\([Jing et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib21)\), and BLADE grades the analysis decisions themselves — which variables, transformations and models an agent chose — against ground truth collected from independent expert analyses\([Gu et al\. 2024b](https://arxiv.org/html/2608.21601#bib.bib15)\)\. Research\-engineering suites raised the horizon further: MLE\-bench\([Chan et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib6)\), MLGym\([Nathani et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib44)\), RE\-Bench, which compares agents against human experts on frontier research\-engineering tasks\([Wijk et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib66)\), and PaperBench, which asks agents to replicate published results\([Starace et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib61)\)\.

In science proper, ScienceAgentBench grades data\-driven discovery tasks distilled from published work with rubric and output checks\([Chen et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib7)\), DiscoveryBench formalizes data\-driven hypothesis search with verifiable targets\([Majumder et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib40)\), BixBench evaluates open\-ended bioinformatics analysis over real notebooks\([Mitchener et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib43)\), CORE\-Bench tests computational reproducibility of published papers\([Siegel et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib55)\), and SciGym turns systems biology into a dry lab where an agent designs experiments against an SBML simulator with known ground truth\([Duan et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib11)\)\. AstaBench packages a scientific research suite with a strong emphasis on harness control and reproducible agent comparison, and is the prior scientific suite that comes closest to deployed\-agent traffic, with a portion of its problems inspired by real requests to its Asta agents\([Bragg et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib5)\)\. The most recent entrants push toward the same territory K\-Bench occupies from a different direction: GeneBench\-Pro simulates genomics data with a known causal structure so that multistage statistical reasoning can be graded exactly\([Li and Ho 2026](https://arxiv.org/html/2608.21601#bib.bib26)\); FrontierScience assembles expert\-level scientific tasks\([Wang et al\. 2026](https://arxiv.org/html/2608.21601#bib.bib64)\); AIRS\-Bench targets frontier research\-science agents\([Lupidi et al\. 2026](https://arxiv.org/html/2608.21601#bib.bib37)\); and capability\-oriented discovery benchmarks ask directly whether current systems are ready to function as scientists\([Song et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib58);[Shi et al\. 2026](https://arxiv.org/html/2608.21601#bib.bib51)\)\.

### 2\.4 Items drawn from deployment traffic

A separate line of work shifts where tasks come from\. Published interaction logs made it possible: LMSYS\-Chat\-1M and WildChat each released on the order of a million real user conversations\([Zheng et al\. 2023a](https://arxiv.org/html/2608.21601#bib.bib72);[Zhao et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib71)\)\. WildBench then selected 1,024 challenging tasks from over a million chat logs and scored them against task\-specific checklists rather than a key\([Lin et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib30)\), and the BenchBuilder pipeline behind Arena\-Hard automated the curation step, mining hard open\-ended prompts from Chatbot Arena and WildChat and measuring how well the resulting set separates models\([Li et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib28)\)\. Both grade a single response from a general assistant rather than work an agent executed\. The agentic version is recent: RealClawBench reconstructs execution environments for 281 tasks sampled from real developer\-agent sessions, scores them with deterministic verifiers while preserving the source distribution, and reports that the best of 14 models solves 65\.8%\([Lv et al\. 2026](https://arxiv.org/html/2608.21601#bib.bib38)\)\. That design enables automatic scoring at the cost of reconstruction, since an item survives only if a verifier can be written for it\. K\-Bench takes the opposite trade: no reconstruction and no verifier, so the distribution arrives intact and the whole scoring burden moves onto the judges\.

### 2\.5 Judging without a key

Because open\-ended scientific work has no key, K\-Bench inherits the methodological literature on model\-based judging rather than the literature on exact match\. Model judges were shown to track human preference at scale in MT\-Bench and Chatbot Arena\([Zheng et al\. 2023b](https://arxiv.org/html/2608.21601#bib.bib73);[Chiang et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib8)\), and G\-Eval established form\-filling chain\-of\-thought evaluation as a practical protocol\([Liu et al\. 2023b](https://arxiv.org/html/2608.21601#bib.bib34)\)\. Written rubrics are how that protocol reaches expert domains\. HealthBench grades 5,000 open\-ended health conversations against 48,562 criteria written by 262 physicians\([Arora et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib2)\), and its professional edition applies the same machinery to real clinician chats\([Soskin Hicks et al\. 2026](https://arxiv.org/html/2608.21601#bib.bib59)\); ResearchRubrics pairs deep\-research prompts with expert\-written rubrics and finds leading agents below 68% compliance, a result that survives in the adjacent deep\-research suites\([Sharma et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib50);[Du et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib10)\)\. K\-Bench’s rubric is deliberately coarser than these — eight dimensions applied to every task rather than criteria authored per item — because the items are not known before the draw\. The known pathologies are equally well established: judges favor their own generations\([Panickssery et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib46)\), they exhibit position, verbosity and style biases that are separable and measurable\([Ye et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib68);[Hu et al\. 2024b](https://arxiv.org/html/2608.21601#bib.bib18);[Shi et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib52)\), a panel drawn from disjoint model families carries less intra\-model bias than a single large judge\([Verga et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib63)\), and the field now has systematic surveys of both the method and its failure modes\([Gu et al\. 2024a](https://arxiv.org/html/2608.21601#bib.bib14)\)\. Section[3](https://arxiv.org/html/2608.21601#S3)records our controls here\.

### 2\.6 What K\-Bench adds, and what it gives up

K\-Bench adds distributional fidelity in a scientific setting: items are drawn from what scientists sent an agent, with their attachments, their ambiguity and their length, and the grading looks at the delivered artifacts rather than at a reconstructed answer\. It forgoes a reference solution, so absolute correctness is rubric\-anchored rather than key\-anchored\. There is no expert human baseline, so “acceptable” is a standard rather than a measured reference\. Finally, the task set is private\. We return to that trade in Section[7](https://arxiv.org/html/2608.21601#S7)\.

Table 1:K\-Bench relative to representative scientific and agentic benchmarks\. The year in parentheses is the first public posting year of that suite, matching Figure[2](https://arxiv.org/html/2608.21601#S2.F2)\. “Public” refers to release of the task items, not to the existence of a paper\. The table characterizes design choices, not quality: each row buys a different property, and the properties are not substitutes\.

## 3 Methods and harness

### 3\.1 Task set

We drew the task set from approximately 18,000 user sessions logged on K\-Dense Web, a cloud\-hosted scientific\-agent application released in December 2025\([K\-Dense Inc\. 2026](https://arxiv.org/html/2608.21601#bib.bib22);[Li et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib27)\)\. A user sends a request plus files; the production harness then runs the work, including deep research, literature review, and hypothesis generation\([Agarwal et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib1)\)\. K\-Dense Web has API access to more than 200 databases and ships with pre\-installed agent skills that tune the harness for scientific applications\. The items used in this study are those incoming requests\. They were executed on the unmodified LLM without the production K\-Dense Web harness, and they do not receive its database APIs or scientific skills\.

From that population we drew a uniform random subset of 200 sessions, in two batches \(2026\-08\-06\-fulland2026\-08\-07\-batch2\-cpu\)\. We then required that a session run to completion under every one of the nine benchmarked models\. Several models decline some requests on safeguard grounds, and a task attempted by eight models but refused by the ninth would put the per\-model means on different task sets, so we kept only the prompts that ran across all nine\. That complete\-case rule left 178 sessions \(93 and 85 in the two batches\), spanning four scientific domains: life sciences \(n=59n=59\), clinical and health \(n=59n=59\), physical sciences, engineering and computer science \(n=43n=43\), and chemistry, drug and materials \(n=17n=17\)\. Prompting is one\-shot\. Each session was reduced to its first user message, verbatim, together with the files attached to that message\. Follow\-up user turns were discarded\.

Attachments are a defining feature of the distribution: 125 of 178 sessions \(70%\) carry at least one file, the modal non\-zero count is one, and the tail is long\. Prompt length spans two orders of magnitude, from a median of 96 bytes in the shortest quartile to 6,217 bytes in the longest\.

### 3\.2 Models and harness

Nine models ran every task:gpt\-5\.6\-sol,claude\-opus\-5,gpt\-5\.6\-luna,kimi\-k3,grok\-4\.5,gemini\-3\.6\-flash,muse\-spark\-1\.2,gemma\-4\-31b\-it, andnemotron\-3\-ultra\-550b\-a55b\. Each run executed in an isolated Modal sandbox with identical tooling: the stockpi0\.84\.0 harness\([Earendil Works 2026](https://arxiv.org/html/2608.21601#bib.bib12)\), its full built\-in tool set \(shell, file read/write/edit, web search, content fetch, search\-content retrieval, and a source\-checking tool\) and web access\. Thinking levelmaxwas requested where the model exposed one\. We wrote no model\-specific prompts, added no agent skills or sub\-agents, and did not retry failed runs\. Holding the scaffold fixed is not a neutral choice: harness and scaffold move agent results by margins comparable to the model itself\([Kapoor et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib23);[Bragg et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib5)\), so a campaign that varied both would not attribute anything\. All9×178=1,6029\\times 178=1\{,\}602runs completed\. Total inference cost for the generation campaign was $3,649\.18\. All models were accessed through OpenRouter; Table[24](https://arxiv.org/html/2608.21601#A1.T24)records the exact identifier, provider, listing date and context window for each system, so that a reader can tell which artefact was measured\. The campaign ran between 6 and 12 August 2026\.

### 3\.3 Rubric

Judges applied rubric v1\.0 \(7 August 2026\)\. The full judge\-facing instructions are reproduced in Appendix[A\.1](https://arxiv.org/html/2608.21601#A1.SS1)\. Every dimension is an integer from 0 to 10 with written anchors at 0, 3, 5, 8 and 10\. The key anchor is 8:*a domain scientist would accept this work with minor edits*\. Scores of 9 and 10 are reserved for publishable, expert\-grade output\. Judges score what was delivered\.

The eight dimensions aretask\_fulfillment\(coverage of explicit and reasonable implicit requirements at the requested depth\),scientific\_accuracy\(claims, methods, statistics, units, formulas and citations\),reasoning\_quality\(planning, decomposition and error recovery as visible in the transcript\),tool\_use\(tool choice, efficiency and recovery from failure\),data\_handling\(whether attachments were loaded, parsed, sanity\-checked and faithfully represented\),artifact\_quality\(completeness and usefulness of output files\),communication\(structure, length, register and language match with the prompt\), andhonesty\_calibration\(hallucination, overclaiming, and whether failures and limitations are stated\)\. Two dimensions are conditional and are marked not\-applicable where they do not apply:data\_handlingwhen the task had no files and needed no data, andartifact\_qualitywhen a prose answer is the natural deliverable\. In this campaign 33\.5% ofdata\_handlingand 35\.6% ofartifact\_qualityjudgments were marked N/A\.

Two further fields are recorded\.overallis an explicitly holistic 0–10 score, not an average of the dimensions, weighted by what mattered for that particular task\.fully\_successfulis a boolean answering the question: would the scientist who submitted this task be satisfied with no follow\-up at all? Judges also tag every applicable failure mode from a closed 16\-tag taxonomy \(Table[3](https://arxiv.org/html/2608.21601#A1.T3)\) and report a self\-assessedconfidencein\[0,1\]\[0,1\]\.

### 3\.4 Judging protocol

Three judges scored every run independently:gpt\-5\.6\-sol,qwen3\.8\-maxandgrok\-4\.5\. Drawing them from three vendors follows the finding that a panel of disjoint model families carries less intra\-model bias than any single judge\([Verga et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib63)\); Section[5](https://arxiv.org/html/2608.21601#S5)reports how far that held here\. Each judge ran as an agenticpisession with read, bash and write tools rather than as a single scoring call\. The packet each judge received contained the task prompt; the identity\-scrubbed final answer; a deterministic execution digest of the complete transcript, listing every tool call, error and recovery; an artifact inventory; and read\-only access to the run’s actual output tree\. Judges opened those files before scoring, a mean of 4\.0 artifacts per assessment \(gpt\-5\.6\-sol3\.65,qwen3\.8\-max4\.12,grok\-4\.54\.22\)\.

Model identity was removed from every judge\-visible surface\. Directory names are HMAC blind identifiers, vendor and model strings inside agent\-authored text are redacted, and per\-token cost, which fingerprints a vendor, is withheld from the digest\. Blinding of this kind removes explicit self\-identification but not writing style, and we treat the residual effect as a measurable quantity \(Section[5\.4](https://arxiv.org/html/2608.21601#S5.SS4)\)\.

For integrity, benchmark outputs were locked read\-only for the duration of the judging campaign and every run’s output tree was hashed before and after; a post\-campaign verification pass confirmed that the judges mutated nothing\. Of the 4,806 assessments, 4,798 produced schema\-valid scores on the first attempt, 6 required a second attempt and 2 a third\. All 4,806 completed\. Judging consumed 151\.1 hours of judge wall\-clock time at a cost of $996\.99\.

### 3\.5 Analysis conventions

The evaluation produced three tables that constitute the primary record:scores\_wide\.csv\(one row per assessment: 4,806 rows\),scores\_long\.csv\(one row per dimension score: 43,254 rows\) andrun\_metrics\.csv\(one row per run: 1,602 rows\)\. Every quantity in this paper is computed from those three tables, with three exceptions that draw on the run archive rather than the tables and are identified where they appear\.

Five conventions are used throughout:

*Run\-level aggregation\.*A run’soverallis the mean of its three judges’ holistic scores\. Model means are averages over the 178 runs, and confidence intervals are 95% percentile bootstrap intervals resampling sessions \(2,000 draws\), which respects the fact that tasks, not assessments, are the sampling unit\.

*Success rates\.**Majority success*means more than half of the three judges independently setfully\_successful;*unanimous success*means all three did;*unanimous rejection*means none did\.

*The score pool\.*“All scored judgments” means the eight dimension scores plus the holistic overall for every assessment, excluding N/A cells: 39,934 values\. Percentages of dimension scores below a threshold use non\-N/A denominators, so the two conditional dimensions are scored only on the assessments where they applied\.

*Paired comparisons\.*Where models are compared directly, the unit is a pair of runs on the same task scored by the same judge, which cancels both judge calibration and task difficulty\. With 178 tasks and 3 judges this gives 534 paired comparisons per model pair; win rates exclude ties, and the number of decisive pairs is reported where it matters\. Where a paired comparison involves a conditional dimension, pairs in which either run was marked not\-applicable are dropped before the win rate is formed\.

*Single measurement per cell\.*Each of the 1,602 runs was executed once and scored once by each judge\. Nothing in this design separates model capability from run\-to\-run variance, and no quantity below should be read as an expectation over repeated attempts\.

### 3\.6 Manuscript preparation

This paper was written with partial help from K\-Dense Web\([K\-Dense Inc\. 2026](https://arxiv.org/html/2608.21601#bib.bib22)\), the same platform that supplied the task corpus\. It was used for drafting and editing assistance during manuscript preparation\. It was not used to generate, execute, or judge any benchmark run, and it played no part in producing the numbers reported here: all quantities come from the three score tables described above\. The authors verified every claim in the manuscript and take full responsibility for its content\.

## 4 Results

### 4\.1 The tasks are not solved

The best model in the campaign,gpt\-5\.6\-sol, averages 8\.04 out of 10 across the 178 tasks, with a 95% bootstrap interval of \[7\.80, 8\.23\]\. It is the only one of nine models whose pooled point estimate reaches the rubric’s 8\-anchor and the count of models reaching it is judge\-dependent \(Section[5\.2](https://arxiv.org/html/2608.21601#S5.SS2)\)\. The gap to second place is 0\.42 points \(Table[4](https://arxiv.org/html/2608.21601#A1.T4), Figure[3](https://arxiv.org/html/2608.21601#S4.F3)\)\. That margin is smaller than the self\-preferencegpt\-5\.6\-solshows as a judge, and under either judge that is notgpt\-5\.6\-solthe first two places exchange; we develop this in Section[5\.4](https://arxiv.org/html/2608.21601#S5.SS4)and treat the top of the table as unresolved\.

![Refer to caption](https://arxiv.org/html/2608.21601v1/figures/headline.png)Figure 3:Mean overall score by model, pooled over judges\.Blue bars are means over the 178 tasks with a 95% bootstrap interval marked by the white tick; the gray bar spans the strictest to the most lenient judge’s mean for that model\. The shaded region marks scores at or above the rubric’s 8\-anchor\. The panel title printed inside the figure states that one model reaches that line; that count is the pooled\-panel value, and it is zero, one or two depending on which judge is asked \(Section[5\.2](https://arxiv.org/html/2608.21601#S5.SS2)\)\.Three results describe how far the corpus is from solved\. First, the distribution of scores: of all 39,934 scored judgments, 47\.6% fall below 8 and 20\.2% fall below 5\. Second, success rates \(Figure[5](https://arxiv.org/html/2608.21601#S4.F5)\): 40\.1% of runs are called fully successful by a majority of judges and 19\.2% by all three, while 45\.9% are rejected unanimously\. Third, task coverage: 22 of 178 tasks \(12\.4%\) were not majority\-solved by any of the nine models, and the same number had no model reach a mean overall of 8\. Only 6 tasks were majority\-solved by all nine\.

The judges bracket the success rate widely, which is why we report the bracket rather than the midpoint\. The strictest judge,gpt\-5\.6\-sol, marks 22\.2% of runs fully successful; the most lenient,qwen3\.8\-max, marks 48\.9%\. Figure[4](https://arxiv.org/html/2608.21601#S4.F4)shows the effect on levels; the ordering of models is almost unchanged across panels, which is the property the paired analysis of Section[4\.5](https://arxiv.org/html/2608.21601#S4.SS5)rests on\.

![Refer to caption](https://arxiv.org/html/2608.21601v1/figures/overall_by_judge.png)Figure 4:The same 1,602 runs scored by three judges on three different scales\.Each panel gives mean overall score by model for one judge, with 95% bootstrap intervals over sessions; rows are held in the order of the pooled mean\.![Refer to caption](https://arxiv.org/html/2608.21601v1/figures/success_rate.png)Figure 5:How often a run fully satisfies the scientist who asked\.Light bars give the share of that model’s runs marked fully successful by a majority of the three judges; dark bars require unanimity; the gray whisker spans the individual judges\.claude\-opus\-5has the highest majority rate of any model \(75%\) whilegpt\-5\.6\-solhas 71%\. The two differ far more on the unanimous rate \(23% against 56%\), but that gap should not be read as a property of the models: unanimity requires the strictest judge to agree, that judge isgpt\-5\.6\-sol, and it scoresclaude\-opus\-51\.8 points below its peers \(Sections[5\.3](https://arxiv.org/html/2608.21601#S5.SS3)and[5\.4](https://arxiv.org/html/2608.21601#S5.SS4)\)\.Figure[6](https://arxiv.org/html/2608.21601#S4.F6)and Table[5](https://arxiv.org/html/2608.21601#A1.T5)give the full score distribution per judge\. The disagreement is a level shift: all three judges put a large mass at 8 and 9, all three have a substantial low tail, and the share of scores below the acceptable line runs from 38\.8% for the most lenient judge to 61\.1% for the strictest\.

![Refer to caption](https://arxiv.org/html/2608.21601v1/figures/score_distribution.png)Figure 6:Distribution of individual scored judgments by judge\.Every dimension score and the holistic score from every judged run, pooled per judge; the axis labels shorten this to “dimension” scores\. The strictest judge places 61\.1% of scores below the acceptable line and the most lenient 38\.8%, so even the lenient reading leaves well over a third of the work short of acceptable\.
### 4\.2 Accuracy lags communication in every model

The scores are not uniform across the rubric, and the most secure form of the non\-uniformity is a single pairwise comparison\. Scientific accuracy averages 6\.22 and communication 7\.33\. Both are scored on all assessments, so the two rest on identical denominators, and the ordering holds within every one of the nine models, by margins from 0\.21 to 2\.26 points \(Table[6](https://arxiv.org/html/2608.21601#A1.T6)\)\. Every model in the campaign presents its work better than it does the work\.

Sorting the rubric into three execution dimensions — tool use \(6\.79\), communication \(7\.33\) and reasoning quality \(6\.68\), mean 6\.93 — and three substance dimensions — scientific accuracy \(6\.22\), honesty and calibration \(7\.30\) and artifact quality \(5\.50\), mean 6\.34 — gives a gap of 0\.59 points that runs in the same direction for all nine models \(Figures[7](https://arxiv.org/html/2608.21601#S4.F7)and[8](https://arxiv.org/html/2608.21601#S4.F8)\)\. Artificat quality is low partly because many runs write nothing \(Section[4](https://arxiv.org/html/2608.21601#S4)\)\.

![Refer to caption](https://arxiv.org/html/2608.21601v1/figures/dimension_gap.png)Figure 7:Mean score by rubric dimension across every model, judge and task\.Execution dimensions \(blue\) average 6\.93; substance dimensions \(orange\) average 6\.34\. Task fulfillment and data handling \(gray\) belong cleanly to neither group\. The color assignment is a judgment call and the 0\.59\-point gap is sensitive to it \(Section[4\.2](https://arxiv.org/html/2608.21601#S4.SS2)\); the denominator\-matched comparison of scientific accuracy against communication is not\. No dimension reaches the acceptable line on average\.![Refer to caption](https://arxiv.org/html/2608.21601v1/figures/dimension_heatmap.png)Figure 8:Every model against every dimension\.Columns run weakest to strongest, rows strongest to weakest\. Artifact quality is the weakest dimension for seven of the nine models; the exceptions areclaude\-opus\-5, weakest on scientific accuracy \(7\.31 against 7\.62\), andgemini\-3\.6\-flash, weakest on honesty and calibration \(5\.41 against 5\.70\)\. No single dimension is the strongest for a majority: communication tops four models, honesty and calibration four more, and data handling one\.A per\-model view of the same phenomenon makes the point sharply\. Taking each model’s honesty score minus its scientific\-accuracy score, every model exceptgemini\-3\.6\-flashscores itself more honest than it is accurate, by margins from 0\.69 \(claude\-opus\-5\) to 2\.67 \(nemotron\-3\-ultra\-550b\-a55b\)\. Every model without exception scores higher on communication than on artifact quality, by margins from 0\.77 to 4\.45 \(Table[2](https://arxiv.org/html/2608.21601#S4.T2)\)\. An evaluation that scores only the prose will rank these systems in close to the wrong order\.

Table 2:Self\-presentation versus substance, by model\.All values are model means over 4,806 assessments\. The last two columns are the differences honesty−\-accuracy and communication−\-artifact quality; positive values mean the run reads better than it is\. Computed fromscores\_wide\.csv\.
### 4\.3 The leading failure is misrepresentation

Judges tagged every run from a closed 16\-tag taxonomy \(defined in full in Table[3](https://arxiv.org/html/2608.21601#A1.T3)of the rubric, Appendix[A\.1](https://arxiv.org/html/2608.21601#A1.SS1)\), and tags are not exclusive, so columns do not sum to 100%\. The ranking is unambiguous \(Table[7](https://arxiv.org/html/2608.21601#A1.T7), Figure[9](https://arxiv.org/html/2608.21601#S4.F9)\)\. The most frequent tag in the corpus isoverclaiming, on 31\.4% of assessments, followed bymissing\_artifacts\(22\.6%\),shallow\_analysis\(17\.7%\),truncated\_run\(16\.4%\) andpremature\_completion\(12\.5%\)\. Taking the three honesty\-family tags together \(overclaiming,fabricated\_results,fabricated\_citations\), 32\.3% of assessments carry at least one\. Aggregating to runs instead of assessments, 54\.4% of runs are tagged for overclaiming by at least one of their three judges, 26\.9% by at least two, and 12\.9% by all three\. The assessment\-level and run\-level framings differ by a factor of nearly two\.

![Refer to caption](https://arxiv.org/html/2608.21601v1/figures/failure_modes.png)Figure 9:Failure\-mode frequency overall \(left\) and by model \(right\)\.The leading tag across the corpus isoverclaimingat 31\.4% of assessments\. The per\-model panel shows that this is not a property of the weakest systems only:gemini\-3\.6\-flashcarries it on 68% of its assessments andgemma\-4\-31b\-iton 44%, but so doclaude\-opus\-5\(33%\) andkimi\-k3\(32%\), whilegpt\-5\.6\-sol\(6%\) andgpt\-5\.6\-luna\(10%\) are markedly cleaner\.For an evaluation audience this is the most useful axis in the set, because honesty is expensive to measure\. It requires knowing both what a system claimed and what it actually did, which is the pairing K\-Bench’s judging packet supplies\. The dispersion across models is large enough to be a training signal rather than noise\. Two systems in the campaign carry the tag on under 10% of their runs while two others exceed 44%, on the same 178 tasks, under the same rubric, read by the same three judges\.

Overclaiming rises with attachments: 33\.9% of assessments on tasks with attached files versus 25\.4% without\.statistical\_malpracticeshows the sharpest domain structure, from 1\.1% in chemistry and materials to 8\.4% in life sciences, which is what one would expect from a distribution in which the life\-sciences requests are the ones most likely to involve an inferential test\.

### 4\.4 What the transcripts show

Models differ enormously in whether they ever reach for evidence \(Table[8](https://arxiv.org/html/2608.21601#A1.T8), Figure[10](https://arxiv.org/html/2608.21601#S4.F10)\)\. Shell use is common but not universal, from 43% ofgemma\-4\-31b\-itruns to 96% ofclaude\-opus\-5runs, with the other seven models between 66% and 87%\. Evidence\-seeking tools separate the field:source\_checkis used in 34% ofgpt\-5\.6\-solruns and 30% ofclaude\-opus\-5runs but in 0% ofgemini\-3\.6\-flashandnemotron\-3\-ultra\-550b\-a55bruns;fetch\_contentruns from 66% down to 2%; web search from 66% down to 16%\. A model that rarely fetches a page or checks a source is structurally unable to ground a scientific claim, whatever its reasoning quality\.

![Refer to caption](https://arxiv.org/html/2608.21601v1/figures/capability_profile.png)Figure 10:Tool reach \(left\) and empty\-handed runs \(right\)\.Left: the share of runs in which each model used each tool at least once, rows ordered by overall usage rather than by category; the evidence\-seeking tools \(web\_search,fetch\_content,get\_search\_contentandsource\_check\) separate the field most sharply\. Right: the share of runs that end with no file on disk, from 17% forclaude\-opus\-5to 75% fornemotron\-3\-ultra\-550b\-a55b\. A run can score well on prose and still leave a user with nothing\.Deliverable production is the second measurement\. Results show that 767 of 1,602 runs \(47\.9%\) end with no output file at all\. Those runs average 5\.26 overall against 6\.68 for runs that leave something behind\. Per model the empty\-handed rate runs from 17\.4% \(claude\-opus\-5\) to 74\.7% \(nemotron\-3\-ultra\-550b\-a55b\), withgpt\-5\.6\-solat 37\.1% despite leading on every score\-based measure\. Volume beyond the first file adds little: runs leaving 1–2 files average 6\.87 and runs leaving 11 or more average 6\.64, so the discontinuity is between nothing and something \(Table[9](https://arxiv.org/html/2608.21601#A1.T9)\)\.

The third measurement is self\-verification \(Figure[11](https://arxiv.org/html/2608.21601#S4.F11)\)\. Counted from the transcripts,claude\-opus\-5performs 23\.9 verification actions per run against 0\.58 forgemma\-4\-31b\-it, a factor of roughly 41, and writes 70\.2 KB of code per run against 3\.0 KB\. The ordering of models by verification frequency tracks the ordering by score closely, which makes it a cheap, deterministic proxy that a post\-training team can compute on its own transcripts without running a judging campaign at all\. These aggregates are extracted from the transcripts themselves, which are retained in the run archive rather than summarized in the score tables\.

![Refer to caption](https://arxiv.org/html/2608.21601v1/figures/scientific_behavior.png)Figure 11:Measured behavior from the transcripts, no judge involved\.Verification actions per run, code written per run, and distinct scientific packages imported per run\.claude\-opus\-5verifies its own output about 41 times as often per run asgemma\-4\-31b\-itand writes roughly 23 times more code\. Verification frequency is deterministic to compute and tracks the score ordering closely\. Extracted from the 1,602 run transcripts\.
### 4\.5 Paired comparisons

Pairing separates models that the pooled means cannot\.gpt\-5\.6\-solbeatsclaude\-opus\-5in 64% of decisive paired matchups, and the mean paired delta is−0\.42\-0\.42\[−0\.64\-0\.64,−0\.20\-0\.20\] inclaude\-opus\-5’s disfavor\. Of the 534 pairs,gpt\-5\.6\-solwins 39\.0%,claude\-opus\-5wins 21\.9%, and 39\.1% are ties, which is itself a useful number: on two runs out of five the two strongest systems are indistinguishable to the same judge on the same task\. Bradley\-Terry latent strengths fitted by maximum likelihood to the paired outcomes preserve that ordering, and their bootstrap intervals are disjoint for every adjacent pair in the field exceptkimi\-k3andgrok\-4\.5, whose intervals overlap \(Table[11](https://arxiv.org/html/2608.21601#A1.T11), Figure[12](https://arxiv.org/html/2608.21601#S4.F12)\)\.

Both quantities are computed over all three judges and therefore inherit what pairing does not remove\. Thegpt\-5\.6\-sol–claude\-opus\-5cell is the one most exposed: one of its three judges isgpt\-5\.6\-sol, which scoresclaude\-opus\-51\.8 points below the other two \(Section[5\.4](https://arxiv.org/html/2608.21601#S5.SS4)\)\. The 64% win rate and the disjoint Bradley\-Terry intervals at the top of the field should be read with that in mind\. Every comparison below the top two involves at most one contestant judge scoring itself and is correspondingly less affected\. A neutral\-judge\-only refit is the obvious check and is deferred to K\-Bench 02\.

![Refer to caption](https://arxiv.org/html/2608.21601v1/figures/head_to_head.png)Figure 12:Paired comparison matrix \(left\) and fitted Bradley\-Terry strengths with 95% bootstrap intervals \(right\)\.Each cell of the matrix aggregates 534 comparisons on the same task by the same judge\. Pairing removes the additive judge\-calibration term that dominates the pooled means and separates adjacent models that the means leave overlapping\.Broken out by dimension,claude\-opus\-5loses togpt\-5\.6\-solon seven of eight dimensions and wins decisively on one: tool use, where it takes 73% of decisive paired matchups \(Table[12](https://arxiv.org/html/2608.21601#A1.T12)\)\. Its weakest dimensions against the leader are honesty \(17%\) and scientific accuracy \(23%\)\. The two strongest systems in the campaign have different strengths: one is the better engineer, the other the more careful scientist\. For a laboratory choosing between them, the relevant question is not quantified performance, but whether the failure it can least afford is a clumsy pipeline or a confident wrong claim\.

### 4\.6 What makes a task hard

Difficulty is a property of the task distribution rather than of any one specialist area\. Pooled across models, the four domains span 0\.2 points \(chemistry and materials 5\.9, clinical and health 6\.1, life sciences 5\.9, physical sciences and engineering 6\.0\), and each model’s own four domain means span at most±0\.61\\pm 0\.61\(Table[13](https://arxiv.org/html/2608.21601#A1.T13)\)\. Two structural properties of the request itself tend to predict its outcome: how many files came with it, and how long it was\.

##### Attachments\.

Runs on tasks with attached files average 5\.85 against 6\.36 without, and their majority\-success rate falls from 50\.7% to 35\.6%\. The effect is monotone in the number of files: 6\.36 at zero attachments, 6\.28 at one, 5\.66 at two or three, and 5\.39 at four or more, with majority success falling from 50\.7% to 23\.5% across the same bins \(Table[29](https://arxiv.org/html/2608.21601#A1.T29)\)\. What makes this interesting is that the burden is not shared evenly \(Table[14](https://arxiv.org/html/2608.21601#A1.T14)\)\. Attachments cost the weak models 1\.3 to 1\.5 points and the strong models essentially nothing;gpt\-5\.6\-solis 0\.08 points*better*with files than without, andkimi\-k3is 0\.46 better\. Files therefore act as a difficulty amplifier that widens the field\.

##### Request length\.

All of the nine models score lower on the longest quartile of prompts than on the shortest \(Table[15](https://arxiv.org/html/2608.21601#A1.T15), Figure[13](https://arxiv.org/html/2608.21601#S4.F13)\)\. The drop from Q1 \(median 96 bytes\) to Q4 \(median 6,217 bytes\) is 0\.8 points forgpt\-5\.6\-sol, 0\.7 forclaude\-opus\-5, 1\.0 forgrok\-4\.5, 1\.8 formuse\-spark\-1\.2and 3\.4 fornemotron\-3\-ultra\-550b\-a55b\. Long, multi\-part scientific requests are the hardest region of this distribution and the natural sub\-slice to track separately\. The strongest single run\-level correlate of quality in the whole campaign is negative and structural: Spearmanρ=−0\.42\\rho=\-0\.42between input tokens and overall score\. In this distribution, longer prompts are a reliable signal of difficulty\.

![Refer to caption](https://arxiv.org/html/2608.21601v1/figures/task_conditioning.png)Figure 13:What conditions difficulty\.Left: mean overall by attached file type, for file types appearing on at least ten tasks; no format is a soft target and the model spread within each format is 5\.7–6\.8 points\. Right: mean overall by prompt\-length quartile, one line per model; all nine score lower on Q4 than on Q1, most steeply fornemotron\-3\-ultra\-550b\-a55b\. The decline is not monotone for every model:claude\-opus\-5recovers from Q2 to Q3,kimi\-k3from Q1 to Q2 andmuse\-spark\-1\.2from Q3 to Q4\. File extensions and prompt byte counts are task\-registry attributes \(Section[Data, code and availability](https://arxiv.org/html/2608.21601#Sx3)\)\.

### 4\.7 How runs end

How a run terminates is deterministic for its outcome \(Table[16](https://arxiv.org/html/2608.21601#A1.T16)\)\. Of 1,602 runs, 1,354 ended with a clean stop, 224 hit a context window limit, 14 ended in error and 10 ended mid\-tool\-use\. Clean\-stopping runs average 6\.70 and reach majority success 47\.4% of the time\. Every other ending has a majority\-success rate of exactly zero\. Truncated transcripts \(258 runs, 16\.1%\) average 2\.18 against 6\.73 for untruncated, again with no majority successes at all\.

Two consequences follow\. For evaluation, truncation is not a nuisance to be filtered out but a first\-class failure mode that is unevenly distributed across systems: two models truncate on more than half their runs and two never truncate at all, so filtering truncated runs would rescore the weakest systems upward\. For deployment, a length\-limited run is a total loss rather than a partial one, which argues for harness\-level checkpointing of intermediate artifacts rather than for longer limits alone\.

Truncation falls almost entirely on the two lowest\-ranked systems —nemotron\-3\-ultra\-550b\-a55btruncates on 64\.6% of its runs andmuse\-spark\-1\.2on 52\.8%, against 3\.9% or less for the top four — which invites the reading that the bottom of the table is measuring context budgets rather than ability\. The advertised windows are given in Table[24](https://arxiv.org/html/2608.21601#A1.T24)and they do not predict truncation\.gemini\-3\.6\-flashandmuse\-spark\-1\.2run on identical 1,048,576\-token windows and truncate on 0\.0% and 52\.8% of runs respectively;gemma\-4\-31b\-ithas the smallest window in the field at 262,144 and truncates less often \(7\.9%\) thangrok\-4\.5at 500,000 \(11\.2%\)\. What separates them is how fast a run consumes the window, not how large it is:muse\-spark\-1\.2carriestool\_thrashingon 30\.5% of its assessments against 5\.1% across the corpus, at a 26% tool error rate \(Tables[7](https://arxiv.org/html/2608.21601#A1.T7)and[25](https://arxiv.org/html/2608.21601#A1.T25)\)\.

### 4\.8 Effort and cost

Effort correlates with quality across the corpus, moderately and positively: Spearmanρ=0\.27\\rho=0\.27for tool calls,0\.330\.33for wall\-clock time,0\.310\.31for cost and0\.350\.35for thinking characters \(Figure[14](https://arxiv.org/html/2608.21601#S4.F14), Table[17](https://arxiv.org/html/2608.21601#A1.T17)\)\. Some of the spread between models therefore reflects how much work a run did, and comparisons that ignore it are partly measuring budget\.

![Refer to caption](https://arxiv.org/html/2608.21601v1/figures/effort_vs_score.png)Figure 14:Effort explains some of the gap, not most of it\.Hexagonal density of the 1,602 runs against tool calls \(left\) and wall\-clock minutes \(right\), with the median score per effort bin overlaid\. The median trends upward across most of the range, though not monotonically: it dips in the lowest wall\-clock bins before rising and flattens in the highest\. The vertical spread within each bin stays wide throughout: low\-effort runs that score well and long expensive runs that fail are both common\.Within a model, however, the sign flips\. For seven of nine models the within\-model correlation between turns and score is negative, from−0\.37\-0\.37forgemini\-3\.6\-flashto−0\.08\-0\.08fornemotron\-3\-ultra\-550b\-a55b, whilegpt\-5\.6\-sol\(\+0\.01\+0\.01\) andkimi\-k3\(\+0\.02\+0\.02\) are flat\. The between\-model and within\-model relationships answer different questions: across models more effort marks a more capable system, but within a fixed model a long run usually means a stuck one\. Treating turn count as a proxy for diligence is safe only in the first sense\.

Tool error rate has a non\-monotone relationship with quality that is worth stating\. Runs with no tool errors at all average 5\.95, runs with an error rate between 0 and 5% average 7\.47, and runs above 15% average 4\.36\. The best outcomes come from runs that attempted enough to fail occasionally and recovered; zero errors mostly marks a run that never tried anything demanding\.

Cost separates the field by more than three orders of magnitude \(Table[18](https://arxiv.org/html/2608.21601#A1.T18)\)\.gpt\-5\.6\-solcosts $8\.51 per task on average, 411 times the $0\.02 ofgemma\-4\-31b\-it, for\+4\.18\+4\.18points of mean score, and it still lands with its interval straddling the acceptable line\. Three points sit on the Pareto frontier:gemma\-4\-31b\-itat the bottom,gpt\-5\.6\-lunain the middle, andgpt\-5\.6\-solat the top\.gpt\-5\.6\-lunais the notable case: at $0\.15 per task it reaches 7\.46, 93% of the leader’s mean score for 1\.8% of its cost, and it achieves 60% majority success\.

### 4\.9 Task difficulty

Twenty\-three tasks were majority\-solved by exactly one model, and the identity of that model is notably not the leaderboard leader\.claude\-opus\-5is the unique solver on 14 of the 23, against 7 forgpt\-5\.6\-sol, one each forgrok\-4\.5andkimi\-k3, and none for the remaining five systems\. A model that loses the aggregate comparison is therefore the only system that gets a specific piece of work done twice as often as the model that wins it\. For a laboratory this is an argument for portfolio behavior rather than for standardization on the top of a leaderboard\.

The within\-task spread across models is correspondingly large\. The mean standard deviation of the nine model means within a task is 2\.27 points, the mean best\-to\-worst range is 6\.32 and the median is 7\.33, and 113 of 178 tasks have a range greater than 6\. Only two tasks have a range below 1\. Model choice, in other words, is usually the dominant term for any individual scientific request in this distribution\.

### 4\.10 How failures travel together

Failure tags are not independent, and their conditional structure describes recognizable syndromes rather than a list of unrelated defects \(Table[19](https://arxiv.org/html/2608.21601#A1.T19), Figure[15](https://arxiv.org/html/2608.21601#S4.F15)\)\. Three patterns were observed\. When a run is tagged forstatistical\_malpractice, it also carriesoverclaiming92% of the time, and when it is tagged forfabricated\_resultsit carriesoverclaiming94% of the time\. When a run is truncated, it is taggedmissing\_artifacts65% of the time, which is the mechanical signature of being cut off before writing anything out\. And when a run is taggedpremature\_completion, it carriesmissing\_artifacts68% of the time andshallow\_analysis46%\.

![Refer to caption](https://arxiv.org/html/2608.21601v1/figures/failure_cooccurrence.png)Figure 15:Failures travel together\.Conditional co\-occurrence of the ten most frequent failure tags: read a row for the share of runs carrying that tag which also carry the column tag\. The diagonal is 100% by construction\.overclaimingis the darkest off\-diagonal column and exceeds 50% on five of the nine other rows, led byfabricated\_results\(94%\) andstatistical\_malpractice\(92%\); it is a minority companion to the truncation\-driven failures\. Where thetruncated\_runandmissing\_artifactsrows and columns cross, at 65% and 48%, is the mechanical signature of runs cut off before writing output\.A single underlying behavior, finishing the narrative regardless of whether the work finished, produces several errors at once\.

## 5 Judge reliability and alignment

This section treats the judging panel as an object of study: how much of the measurement is reproducible, which part is not, and what happens when two of the three judges are also contestants\.

### 5\.1 The judges agree on order and disagree on level

Across the 1,602 runs, mean pairwise Spearman correlation on the holistic overall score isρ=0\.83\\rho=0\.83, while the judges’ own mean overall scores span 5\.37 \(gpt\-5\.6\-sol\), 6\.24 \(grok\-4\.5\) and 6\.38 \(qwen3\.8\-max\), a range of 1\.0 points \(Table[20](https://arxiv.org/html/2608.21601#A1.T20), Figure[16](https://arxiv.org/html/2608.21601#S5.F16)\)\. Kendall’sWWover the three judges’ rankings of the nine models is 0\.955: the judges essentially agree on the ordering of systems while disagreeing systematically on the level at which to anchor the scale\.

![Refer to caption](https://arxiv.org/html/2608.21601v1/figures/judge_agreement.png)Figure 16:Judges rank runs alike and score them on different scales\.Each point is one of the 1,602 runs, jittered off the integer lattice; points above a diagonal were scored higher by the vertical judge\. Panel headings give the pairwise rank correlation and the mean offset in level\. The largest offset is 1\.0 points\.qwen3\.8\-maxandgrok\-4\.5agree closely with each other \(ρ=0\.89\\rho=0\.89, mean absolute difference 0\.60, 91% within a point\), whilegpt\-5\.6\-solsits about a point below both\. With two judges a disagreement is symmetric and unresolvable; with three, the differently\-calibrated one is identifiable\.gpt\-5\.6\-solis the*strict*judge, so it is the one whose calibration is unusual relative to the panel\. It does not establish that the other two are right, which would take a reference the panel does not contain \(Section[5\.2](https://arxiv.org/html/2608.21601#S5.SS2)\)\.

Agreement also varies by dimension in an interpretable way \(Table[21](https://arxiv.org/html/2608.21601#A1.T21), Figure[17](https://arxiv.org/html/2608.21601#S5.F17)\)\. Communication is the easiest thing to agree on in level \(mean absolute difference 0\.59, 91% within a point\)\. Scientific accuracy is the hardest \(1\.47, 60%\), followed by honesty and calibration \(1\.32, 65%\)\. Those are the two dimensions where a judge must form its own view of the domain rather than assess a surface property\.

![Refer to caption](https://arxiv.org/html/2608.21601v1/figures/dimension_agreement.png)Figure 17:The judges agree on ranking long before they agree on level\.Mean pairwise Spearman correlation \(left\) and mean absolute difference \(right\) for the eight rubric dimensions and the holistic overall score\. Scientific accuracy pairs mid\-field rank agreement \(ρ=0\.76\\rho=0\.76, joint fifth of the nine rows\) with the largest level disagreement anywhere in the rubric \(1\.47\): the judges sort runs on it about as consistently as they sort the rest, and anchor the scale furthest apart\.
### 5\.2 What the panel establishes, and what it does not

Three judges placing 1,602 runs in nearly the same order \(W=0\.955W=0\.955\) establishes that the ranking is reproducible under a change of judge\. It establishes nothing about where the scale sits, because a panel can be reliably wrong about level in the same way it is reliably right about order; this panel visibly disagrees about level, by 1\.0 point on the pooled mean and 1\.47 on scientific accuracy\.

how many of the nine models reach the rubric’s 8\-anchor?qwen3\.8\-maxsays two \(claude\-opus\-58\.29,gpt\-5\.6\-sol8\.20\)\.grok\-4\.5says two \(claude\-opus\-58\.15,gpt\-5\.6\-sol8\.01\)\.gpt\-5\.6\-solsays none: its highest mean for any model, its own included, is 7\.90 \(Table[4](https://arxiv.org/html/2608.21601#A1.T4)\)\.

The published evidence for model judges does not close this gap\. Judges were validated at scale against human preference between responses: MT\-Bench and Chatbot Arena compare judge verdicts to human pairwise choices, and G\-Eval reports rank correlation with human ratings\([Zheng et al\. 2023b](https://arxiv.org/html/2608.21601#bib.bib73);[Chiang et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib8);[Liu et al\. 2023b](https://arxiv.org/html/2608.21601#bib.bib34)\)\. Evidence that a judge puts an absolute threshold where an expert would put it comes from a different design: experts writing the criteria for each item, as in HealthBench with 262 physicians\([Arora et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib2)\), or an expert\-written rubric paired with each prompt, as in ResearchRubrics\([Sharma et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib50)\)\. K\-Bench trades that property away by construction\. Because the items are drawn at random from traffic and are not known before the draw, one rubric has to serve every task\.

### 5\.3 Split decisions

The binary success flag makes the structure of disagreement legible\. Of 1,602 runs, 735 were rejected by all three judges, 307 were accepted by all three, 335 split two\-to\-one in favor and 225 split one\-to\-two, so 35\.0% of runs are split decisions\. The identity of the dissenter is extremely lopsided\. On the 335 runs where two judges accepted and one rejected, the lone rejector isgpt\-5\.6\-sol311 times,grok\-4\.514 times andqwen3\.8\-max10 times\. On the 225 runs where only one judge accepted, the lone acceptor isqwen3\.8\-max152 times,grok\-4\.548 times andgpt\-5\.6\-sol25 times\.

A majority across these three judges is not an independent tie\-break; on split decisions it reports what the two mutually agreeing judges think, and the strict judge is outvoted in 93% of the two\-to\-one cases\. Any headline success rate computed by majority therefore inherits the calibration of that pair\.

### 5\.4 Judges scoring their own runs

Two of the three judges,gpt\-5\.6\-solandgrok\-4\.5, are also benchmarked models, so each scores its own runs\. Blinding removes explicit identity but not writing style, and self\-preference in model judges is a documented effect\([Panickssery et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib46);[Ye et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib68)\)\. The calibration\-adjusted self\-preference is the gap between how a judge scores itself and how it scores everyone else, minus the same gap as the peer judges see it\. Formally, for judgejj,

SPj=\(s¯j→j−s¯j→¬j\)−\(s¯¬j→j−s¯¬j→¬j\),\\text\{SP\}\_\{j\}\\;=\\;\\big\(\\bar\{s\}\_\{j\\rightarrow j\}\-\\bar\{s\}\_\{j\\rightarrow\\neg j\}\\big\)\\;\-\\;\\big\(\\bar\{s\}\_\{\\neg j\\rightarrow j\}\-\\bar\{s\}\_\{\\neg j\\rightarrow\\neg j\}\\big\),wheres¯a→b\\bar\{s\}\_\{a\\rightarrow b\}is the mean overall score given by judge setaato model setbb\. The subtraction removes both the judge’s overall strictness and the model’s actual quality, leaving the excess\.

The result is asymmetric \(Table[22](https://arxiv.org/html/2608.21601#A1.T22)\)\.grok\-4\.5shows a negligible\+0\.11\+0\.11\.gpt\-5\.6\-solshows\+0\.83\+0\.83, which is comparable in magnitude to the gap between the first and second models on the leaderboard\. It is the strictest judge in the panel and it is markedly less strict with itself\.

Table[23](https://arxiv.org/html/2608.21601#A1.T23)givesgpt\-5\.6\-sol’s deviation from peer consensus for each model it scored\. The deviations range from−0\.37\-0\.37to−1\.82\-1\.82points\. An effect appears at the family level rather than the identity level\.gpt\-5\.6\-solas a judge ranks its same\-vendor siblinggpt\-5\.6\-lunasecond of nine, while both other judges rank that model fourth\. On the range\-matched comparison above,gpt\-5\.6\-lunareceives\+1\.10\+1\.10\. We cannot separate a genuine disagreement about quality from a family\-recognition effect with this design, and we therefore treat the pooled ranking ofgpt\-5\.6\-lunaas the least trustworthy number in the campaign\.

Two practical consequences follow: First, the pooled ordering at the top of the table is the product of a single judge \(Table[4](https://arxiv.org/html/2608.21601#A1.T4)\)\.gpt\-5\.6\-solis alone in ranking itself first, and it does so by scoringclaude\-opus\-5at 6\.40 against the 8\.15 and 8\.29 its peers award a gaps nowhere else in the panel\. Where the pooled and the neutral\-judge orderings disagree, the neutral one is preferable\. A future edition should either exclude contestants from the panel entirely or add enough neutral judges that the contaminated ones cannot form a majority\.

### 5\.5 Judge confidence

Judges reported their own confidence and the number of artifacts they inspected\. Mean confidence is 0\.92 forgpt\-5\.6\-sol, 0\.86 forgrok\-4\.5and 0\.83 forqwen3\.8\-max; only 72 of 4,806 assessments \(1\.5%\) were made with confidence below 0\.7\. Confidence correlates negatively with the score awarded \(ρ=−0\.34\\rho=\-0\.34\)\.

Artifact inspection correlates positively with the score awarded \(ρ=0\.22\\rho=0\.22\), for the mechanical reason that runs which produce nothing give a judge nothing to open\. The two judges that opened more files \(grok\-4\.54\.22,qwen3\.8\-max4\.12\) are also the two more lenient ones, andgpt\-5\.6\-sol, which opened the fewest \(3\.65\), is the strictest\. We cannot tell from this design whether reading more artifacts causes a more generous assessment or whether a judge that is already inclined to credit a run reads more of it, and we flag the association without a causal reading\.

## 6 Discussion

### 6\.1 What the corpus says about current systems

The headline number invites a saturation reading and does not support one\. The best system is, on average, borderline acceptable \(Table[4](https://arxiv.org/html/2608.21601#A1.T4)\)\. This distribution is partially solved at the top and materially unsolved in the tail\. The correct object of study going forward is the failing subset rather than the pooled mean\.

The systems present their work better than they*do*the work, as noted in Section[4\.2](https://arxiv.org/html/2608.21601#S4.SS2): scientific accuracy and communication are scored on every assessment, and accuracy sits 1\.11 points below communication in every model in the field\. Honesty and calibration is the second\-highest\-scoring dimension in the rubric on average \(7\.30\), whileoverclaimingis the most frequently applied tag in the taxonomy\([Zhang et al\. 2025](https://arxiv.org/html/2608.21601#bib.bib70)\)\.

### 6\.2 Implications for post\-training

Three specific implications follow from the measurements\.

First, the deficit is scientific judgment\. The top five already use a shell in more than three\-quarters of runs and write a readable answer\. They still lose points on the science\. What separates the models is whether they check a source, whether the claim matches the file they wrote, and whether they flag when it does not\.

Second, honesty is trainable in a way that this corpus makes visible\. The dispersion across models is very large, from 6\.4% to 68\.2% of assessments tagged for overclaiming, and it does not track overall capability:gpt\-5\.6\-lunacarries the tag on 9\.7% of assessments while sitting 0\.16 points belowclaude\-opus\-5on the pooled mean, which carries it on 32\.6%\. The two lowest rates in the field, 6\.4% and 9\.7%, belong to the same model family, while systems that score close to them on overall quality sit above 27%\. We cannot attribute that difference to any particular training choice from this evidence, but it is clearly separable from aggregate quality, and the signal needed to reward it is cheap to compute\.

Third, verification behavior is worth optimizing directly\. Both verification actions and evidence\-tool use are deterministic and immune to judge calibration \(Section[4\.4](https://arxiv.org/html/2608.21601#S4.SS4)\)\.

### 6\.3 Implications for deployment

A leaderboard position is the wrong selection criterion when the top systems differ in kind\. One profile is the better engineer; the other is the more careful scientist \(Sections[4\.5](https://arxiv.org/html/2608.21601#S4.SS5)and[4\.9](https://arxiv.org/html/2608.21601#S4.SS9)\)\. Those profiles suit different work\.

The price of quality is steep and non\-linear \(Section[4\.8](https://arxiv.org/html/2608.21601#S4.SS8)\)\. For high\-volume screening work, the cheap side of the cliff is defensible; for work that will be published, the extra points are concentrated exactly in the dimensions of accuracy and honesty, which a reader would notice\.

A strong score\-based ranking is no guarantee of a low empty rate:gpt\-5\.6\-solleads every score\-based measure in the campaign and still leaves 37\.1% of its runs empty\. Harness\-level enforcement that requires declared deliverables to exist before a run is marked complete would address a failure mode that no amount of model improvement in this campaign eliminated\.

### 6\.4 Implications for benchmark design

The first is that judges must open the files; Section[4\.4](https://arxiv.org/html/2608.21601#S4.SS4)shows why\. The second is that a short, self\-contained prompt set would hide the gaps this distribution exposes \(Section[4\.6](https://arxiv.org/html/2608.21601#S4.SS6)\)\. The third is that a judging panel drawn from the systems under test needs explicit handling \(Section[5](https://arxiv.org/html/2608.21601#S5)\): both the majority vote and the pooled mean carry the panel’s composition inside them\.

### 6\.5 Relation to the published landscape

K\-Bench answers a different question from the suites reviewed in Section[2](https://arxiv.org/html/2608.21601#S2)\. RE\-Bench is the instructive exception: its agents outscore human experts at a two\-hour budget and fall behind them at eight, because humans have better returns to time\([Wijk et al\. 2024](https://arxiv.org/html/2608.21601#bib.bib66)\)\. Our one\-shot design samples the short\-horizon end of that curve, so the headroom we report should not be read as a claim about what these models reach given more turns\.

## 7 Limitations

##### No human calibration, and no expert baseline\.

Two distinct things are missing here\. We never checked the instrument: no human scored any run, so we do not know where a domain scientist would place the rubric’s 8\-anchor relative to where the panel places it, and the panel’s own 1\.0\-point spread shows that the placement is not pinned down even among the three judges we used \(Section[5\.2](https://arxiv.org/html/2608.21601#S5.SS2)\)\. We also do not know what a domain scientist would score on these tasks, so the absolute distance from human performance is unknown\.

##### Judge scores are panel\-dependent\.

Two of three judges are also contestants\.gpt\-5\.6\-solrates its own runs\+0\.83\+0\.83after calibration adjustment and ranks its same\-vendor sibling two places higher than the other judges do\.

##### One harness, one\-shot prompt, one attempt\.

Every run used stockpi0\.84\.0 with no model\-specific prompting, no scientific sub\-agents and no retries\. The model received only the initial user prompt and its attached files; there were no follow\-up user turns\.

##### User content was transmitted without content\-level de\-identification\.

The tasks were replayed to nine external providers as submitted, with attachments\. Zero\-retention endpoints were used, but no redaction pass was applied and no audit for identifiers was carried out \(Section[Data provenance, consent and privacy](https://arxiv.org/html/2608.21601#Sx2)\)\.

##### Private items limit external reproducibility\.

The score tables and the analysis code will be released, so every analysis in this paper can be checked independently, but the items will not be, so no group outside K\-Dense can run a new model against them; see Section[Data, code and availability](https://arxiv.org/html/2608.21601#Sx3)\.

##### Traffic is not a random sample of science\.

The 178 tasks are a uniform random draw from K\-Dense Web traffic in two batches during one week, so they span four broad domains with unequal representation \(chemistry and materials contributes 17 tasks\) as an outcome of the draw rather than by design\. They are also a complete\-case set: the 22 sessions at least one model refused were dropped so that every model is scored on identical tasks, which means the benchmark is silent on requests that sit near a vendor’s safety boundary\.

##### Failure tags are judge\-assigned\.

The failure taxonomy is applied by the same judges that assign the scores, so tag rates inherit judge calibration, and tags are not independent of one another \(Section[4\.10](https://arxiv.org/html/2608.21601#S4.SS10)\)\. Assessment\-level and run\-level tag rates differ by up to a factor of two depending on how many judges must agree, which is why we report several thresholds\.

## Evidence and limits

The 1,602 runs were each executed once, and each was scored once by each of the three judges\. Where we give a confidence interval it is a bootstrap over the 178 sampled sessions and describes sampling uncertainty in the task set alone; it does not cover run\-to\-run variance, judge\-to\-judge variance, or any sensitivity to how the panel was composed\. Point estimates should be read as what these models did on these tasks in this harness on these dates, not as expectations over repeated attempts\.

## Data provenance, consent and privacy

*Basis for use\.*The sessions were drawn from K\-Dense Web logs under the platform’s terms of service, which permit analysis of submitted content for product research and evaluation\. Users were not separately notified of this campaign and did not individually opt in to it\. No session was solicited for the benchmark: all sessions were submitted in the ordinary course of using the product, before the draw was made\.

*What was transmitted\.*Each task was replayed to nine third\-party model providers as the user wrote it, with the attached files intact\. The blinding described in Section[3](https://arxiv.org/html/2608.21601#S3)removes*model*identity from judge\-visible surfaces; it is not a de\-identification of user content, and no content\-level redaction pass was applied to prompts or attachments before the runs\. Requests were executed against vendor endpoints configured for zero retention, so submitted content is excluded from provider retention and from provider training\.

## Data, code and availability

The task prompts, attachments, transcripts and output artifacts are not released\. They are user content\.

A de\-identified form of the scores will be published that produces every table and figure in this paper\. That will be enough to reproduce the analyses and check our arithmetic\. Rubric v1\.0 is reproduced in full in Appendix[A\.1](https://arxiv.org/html/2608.21601#A1.SS1), including every dimension anchor and the 16\-tag failure taxonomy\.

## Author contributions

A\.B\. wrote the manuscript, led the benchmarking effort and set its direction, contributed to rubric development, and ran the internal benchmark campaigns\. D\.P\. built the K\-Dense Web infrastructure and the session management from which the task corpus is drawn\. Y\.H\. led AI engineering for K\-Dense Web and designed the session metadata used to sample and characterize the corpus\. T\.K\. contributed to benchmark and rubric design, to the run and judging engineering, and to the statistical analysis and figures, revised the manuscript, and supervised the project\.

## Competing interests

All authors are employed by K\-Dense, Inc\. The evaluation corpus is traffic from K\-Dense Web, a K\-Dense, Inc\. product, and the benchmark design, the harness configuration and the scoring rubric are all K\-Dense’s\. The nine evaluated systems are third\-party models developed by other organizations, in which K\-Dense had no role\. Readers should weigh the results with the authors’ position in mind\.

## Acknowledgements

We thank the K\-Dense users whose requests constitute this evaluation corpus\.

## References

- Agarwal et al\. \(2025\)Vinayak Agarwal, Orion Li, Christopher A\. Petty, Timothy Kassis, Paul W\. K\. Rothemund, David A\. Sinclair, and Ashwin Gopinath\.Guided multi\-agent AI invents highly accurate, uncertainty\-aware transcriptomic aging clocks\.*bioRxiv*, 2025\.doi:10\.1101/2025\.09\.08\.674588\.URL[https://www\.biorxiv\.org/content/10\.1101/2025\.09\.08\.674588v1](https://www.biorxiv.org/content/10.1101/2025.09.08.674588v1)\.
- Arora et al\. \(2025\)Rahul K\. Arora, Jason Wei, Rebecca Soskin Hicks, et al\.HealthBench: Evaluating Large Language Models Towards Improved Human Health\.*arXiv preprint arXiv:2505\.08775*, 2025\.URL[https://arxiv\.org/abs/2505\.08775](https://arxiv.org/abs/2505.08775)\.
- Bean et al\. \(2025\)Andrew M\. Bean, Ryan Othniel Kearns, Angelika Romanou, et al\.Measuring What Matters: Construct Validity in Large Language Model Benchmarks\.*arXiv preprint arXiv:2511\.04703*, 2025\.URL[https://arxiv\.org/abs/2511\.04703](https://arxiv.org/abs/2511.04703)\.
- Boiko et al\. \(2023\)Daniil A\. Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes\.Autonomous chemical research with large language models\.*Nature*, 624\(7992\):570–578, 2023\.doi:10\.1038/s41586\-023\-06792\-0\.URL[https://doi\.org/10\.1038/s41586\-023\-06792\-0](https://doi.org/10.1038/s41586-023-06792-0)\.Introduces Coscientist; preprint arXiv:2304\.05332 \(2023\)\.
- Bragg et al\. \(2025\)Jonathan Bragg, Mike D’Arcy, Nishant Balepur, et al\.AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite\.*arXiv preprint arXiv:2510\.21652*, 2025\.URL[https://arxiv\.org/abs/2510\.21652](https://arxiv.org/abs/2510.21652)\.Published as a conference paper at ICLR 2026\.
- Chan et al\. \(2024\)Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, et al\.MLE\-bench: Evaluating Machine Learning Agents on Machine Learning Engineering\.*arXiv preprint arXiv:2410\.07095*, 2024\.URL[https://arxiv\.org/abs/2410\.07095](https://arxiv.org/abs/2410.07095)\.
- Chen et al\. \(2024\)Ziru Chen, Shijie Chen, Yuting Ning, et al\.ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data\-Driven Scientific Discovery\.*arXiv preprint arXiv:2410\.05080*, 2024\.URL[https://arxiv\.org/abs/2410\.05080](https://arxiv.org/abs/2410.05080)\.
- Chiang et al\. \(2024\)Wei\-Lin Chiang, Lianmin Zheng, Ying Sheng, et al\.Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference\.*arXiv preprint arXiv:2403\.04132*, 2024\.URL[https://arxiv\.org/abs/2403\.04132](https://arxiv.org/abs/2403.04132)\.
- Dehghani et al\. \(2021\)Mostafa Dehghani, Yi Tay, Alexey A\. Gritsenko, et al\.The Benchmark Lottery\.*arXiv preprint arXiv:2107\.07002*, 2021\.URL[https://arxiv\.org/abs/2107\.07002](https://arxiv.org/abs/2107.07002)\.
- Du et al\. \(2025\)Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, and Zhendong Mao\.DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents\.*arXiv preprint arXiv:2506\.11763*, 2025\.URL[https://arxiv\.org/abs/2506\.11763](https://arxiv.org/abs/2506.11763)\.
- Duan et al\. \(2025\)Haonan Duan, Stephen Zhewen Lu, Caitlin Fiona Harrigan, et al\.Measuring Scientific Capabilities of Language Models with a Systems Biology Dry Lab\.*arXiv preprint arXiv:2507\.02083*, 2025\.URL[https://arxiv\.org/abs/2507\.02083](https://arxiv.org/abs/2507.02083)\.
- Earendil Works \(2026\)Earendil Works\.Pi Agent Harness\.[https://github\.com/earendil\-works/pi](https://github.com/earendil-works/pi), 2026\.Version 0\.84\.0; accessed 18 August 2026\.
- Golchin and Surdeanu \(2023\)Shahriar Golchin and Mihai Surdeanu\.Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models\.*arXiv preprint arXiv:2311\.06233*, 2023\.URL[https://arxiv\.org/abs/2311\.06233](https://arxiv.org/abs/2311.06233)\.
- Gu et al\. \(2024a\)Jiawei Gu, Xuhui Jiang, Zhichao Shi, et al\.A Survey on LLM\-as\-a\-Judge\.*arXiv preprint arXiv:2411\.15594*, 2024a\.URL[https://arxiv\.org/abs/2411\.15594](https://arxiv.org/abs/2411.15594)\.
- Gu et al\. \(2024b\)Ken Gu, Ruoxi Shang, Ruien Jiang, et al\.BLADE: Benchmarking Language Model Agents for Data\-Driven Science\.*arXiv preprint arXiv:2408\.09667*, 2024b\.URL[https://arxiv\.org/abs/2408\.09667](https://arxiv.org/abs/2408.09667)\.
- Hendrycks et al\. \(2020\)Dan Hendrycks, Collin Burns, Steven Basart, et al\.Measuring Massive Multitask Language Understanding\.*arXiv preprint arXiv:2009\.03300*, 2020\.URL[https://arxiv\.org/abs/2009\.03300](https://arxiv.org/abs/2009.03300)\.
- Hu et al\. \(2024a\)Xueyu Hu, Ziyu Zhao, Shuang Wei, et al\.InfiAgent\-DABench: Evaluating Agents on Data Analysis Tasks\.*arXiv preprint arXiv:2401\.05507*, 2024a\.URL[https://arxiv\.org/abs/2401\.05507](https://arxiv.org/abs/2401.05507)\.
- Hu et al\. \(2024b\)Zhengyu Hu, Linxin Song, Jieyu Zhang, et al\.Explaining Length Bias in LLM\-Based Preference Evaluations\.*arXiv preprint arXiv:2407\.01085*, 2024b\.URL[https://arxiv\.org/abs/2407\.01085](https://arxiv.org/abs/2407.01085)\.
- Ivanov \(2024\)Igor Ivanov\.BioLP\-bench: Measuring understanding of biological lab protocols by large language models\.*bioRxiv*, 2024\.doi:10\.1101/2024\.08\.21\.608694\.URL[https://www\.biorxiv\.org/content/10\.1101/2024\.08\.21\.608694](https://www.biorxiv.org/content/10.1101/2024.08.21.608694)\.
- Jimenez et al\. \(2023\)Carlos E\. Jimenez, John Yang, Alexander Wettig, et al\.SWE\-bench: Can Language Models Resolve Real\-World GitHub Issues?*arXiv preprint arXiv:2310\.06770*, 2023\.URL[https://arxiv\.org/abs/2310\.06770](https://arxiv.org/abs/2310.06770)\.
- Jing et al\. \(2024\)Liqiang Jing, Zhehui Huang, Xiaoyang Wang, et al\.DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?*arXiv preprint arXiv:2409\.07703*, 2024\.URL[https://arxiv\.org/abs/2409\.07703](https://arxiv.org/abs/2409.07703)\.
- K\-Dense Inc\. \(2026\)K\-Dense Inc\.K\-Dense Web\.[https://www\.k\-dense\.ai](https://www.k-dense.ai/), 2026\.Accessed 18 August 2026\.
- Kapoor et al\. \(2025\)Sayash Kapoor, Benedikt Stroebl, Peter Kirgis, et al\.Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation\.*arXiv preprint arXiv:2510\.11977*, 2025\.URL[https://arxiv\.org/abs/2510\.11977](https://arxiv.org/abs/2510.11977)\.
- Laurent et al\. \(2024\)Jon M\. Laurent, Joseph D\. Janizek, Michael Ruzo, et al\.LAB\-Bench: Measuring Capabilities of Language Models for Biology Research\.*arXiv preprint arXiv:2407\.10362*, 2024\.URL[https://arxiv\.org/abs/2407\.10362](https://arxiv.org/abs/2407.10362)\.
- Laurent et al\. \(2026\)Jon M Laurent, Albert Bou, Michael Pieler, et al\.LABBench2: An Improved Benchmark for AI Systems Performing Biology Research\.*arXiv preprint arXiv:2604\.09554*, 2026\.URL[https://arxiv\.org/abs/2604\.09554](https://arxiv.org/abs/2604.09554)\.
- Li and Ho \(2026\)Jeremy Li and Andrew Ho\.GeneBench\-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicine\.*bioRxiv*, 2026\.doi:10\.64898/2026\.06\.29\.735386\.URL[https://www\.biorxiv\.org/content/10\.64898/2026\.06\.29\.735386v2](https://www.biorxiv.org/content/10.64898/2026.06.29.735386v2)\.Announced at[https://openai\.com/index/introducing\-genebench\-pro/](https://openai.com/index/introducing-genebench-pro/)\.
- Li et al\. \(2025\)Orion Li, Vinayak Agarwal, Summer Zhou, Ashwin Gopinath, and Timothy Kassis\.K\-Dense Analyst: Towards Fully Automated Scientific Analysis\.*arXiv preprint arXiv:2508\.07043*, 2025\.URL[https://arxiv\.org/abs/2508\.07043](https://arxiv.org/abs/2508.07043)\.
- Li et al\. \(2024\)Tianle Li, Wei\-Lin Chiang, Evan Frick, et al\.From Crowdsourced Data to High\-Quality Benchmarks: Arena\-Hard and BenchBuilder Pipeline\.*arXiv preprint arXiv:2406\.11939*, 2024\.URL[https://arxiv\.org/abs/2406\.11939](https://arxiv.org/abs/2406.11939)\.
- Liang et al\. \(2022\)Percy Liang, Rishi Bommasani, Tony Lee, et al\.Holistic Evaluation of Language Models\.*arXiv preprint arXiv:2211\.09110*, 2022\.URL[https://arxiv\.org/abs/2211\.09110](https://arxiv.org/abs/2211.09110)\.
- Lin et al\. \(2024\)Bill Yuchen Lin, Yuntian Deng, Khyathi Chandu, et al\.WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild\.*arXiv preprint arXiv:2406\.04770*, 2024\.URL[https://arxiv\.org/abs/2406\.04770](https://arxiv.org/abs/2406.04770)\.
- Liu et al\. \(2026a\)Amelia Liu, Andrew Ho, Anne Marie Droste, et al\.LifeSciBench: Evaluating Language Models on Realistic, Expert\-Level Tasks in the Life Sciences\.Technical report, OpenAI and Tacit Labs, jun 2026a\.URL[https://openai\.com/index/introducing\-life\-sci\-bench/](https://openai.com/index/introducing-life-sci-bench/)\.
- Liu et al\. \(2026b\)He Liu, Boyuan Gu, Shuaiqi Cheng, Haiyang Sun, Siyu You, and Xuming Hu\.PhysDox: Benchmarking LLMs on Physical Feasibility Auditing of Physiological Sensing Protocols\.*arXiv preprint arXiv:2606\.05003*, 2026b\.URL[https://arxiv\.org/abs/2606\.05003](https://arxiv.org/abs/2606.05003)\.
- Liu et al\. \(2023a\)Xiao Liu, Hao Yu, Hanchen Zhang, et al\.AgentBench: Evaluating LLMs as Agents\.*arXiv preprint arXiv:2308\.03688*, 2023a\.URL[https://arxiv\.org/abs/2308\.03688](https://arxiv.org/abs/2308.03688)\.
- Liu et al\. \(2023b\)Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu\.G\-Eval: NLG Evaluation using GPT\-4 with Better Human Alignment\.*arXiv preprint arXiv:2303\.16634*, 2023b\.URL[https://arxiv\.org/abs/2303\.16634](https://arxiv.org/abs/2303.16634)\.
- Liu et al\. \(2025\)Yuyang Liu, Liuzhenghao Lv, Xiancheng Zhang, Jingya Wang, Li Yuan, and Yonghong Tian\.BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science\.*arXiv preprint arXiv:2505\.07889*, 2025\.URL[https://arxiv\.org/abs/2505\.07889](https://arxiv.org/abs/2505.07889)\.
- Lu et al\. \(2024\)Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha\.The AI Scientist: Towards Fully Automated Open\-Ended Scientific Discovery\.*arXiv preprint arXiv:2408\.06292*, 2024\.URL[https://arxiv\.org/abs/2408\.06292](https://arxiv.org/abs/2408.06292)\.
- Lupidi et al\. \(2026\)Alisia Lupidi, Bhavul Gauri, Thomas Simon Foster, et al\.AIRS\-Bench: a Suite of Tasks for Frontier AI Research Science Agents\.*arXiv preprint arXiv:2602\.06855*, 2026\.URL[https://arxiv\.org/abs/2602\.06855](https://arxiv.org/abs/2602.06855)\.
- Lv et al\. \(2026\)Zongwei Lv, Zhewen Tan, Yaoming Li, et al\.RealClawBench: Live OpenClaw Benchmarks from Real Developer\-Agent Sessions\.*arXiv preprint arXiv:2606\.03889*, 2026\.URL[https://arxiv\.org/abs/2606\.03889](https://arxiv.org/abs/2606.03889)\.
- M\. Bran et al\. \(2024\)Andres M\. Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D\. White, and Philippe Schwaller\.Augmenting large language models with chemistry tools\.*Nature Machine Intelligence*, 6\(5\):525–535, 2024\.doi:10\.1038/s42256\-024\-00832\-8\.URL[https://doi\.org/10\.1038/s42256\-024\-00832\-8](https://doi.org/10.1038/s42256-024-00832-8)\.Introduces ChemCrow; preprint arXiv:2304\.05376 \(2023\)\.
- Majumder et al\. \(2024\)Bodhisattwa Prasad Majumder, Harshit Surana, Dhruv Agarwal, et al\.DiscoveryBench: Towards Data\-Driven Discovery with Large Language Models\.*arXiv preprint arXiv:2407\.01725*, 2024\.URL[https://arxiv\.org/abs/2407\.01725](https://arxiv.org/abs/2407.01725)\.
- Merrill et al\. \(2026\)Mike A\. Merrill, Alexander G\. Shaw, Nicholas Carlini, et al\.Terminal\-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces\.*arXiv preprint arXiv:2601\.11868*, 2026\.URL[https://arxiv\.org/abs/2601\.11868](https://arxiv.org/abs/2601.11868)\.
- Mialon et al\. \(2023\)Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom\.GAIA: a benchmark for General AI Assistants\.*arXiv preprint arXiv:2311\.12983*, 2023\.URL[https://arxiv\.org/abs/2311\.12983](https://arxiv.org/abs/2311.12983)\.
- Mitchener et al\. \(2025\)Ludovico Mitchener, Jon M Laurent, Alex Andonian, et al\.BixBench: a Comprehensive Benchmark for LLM\-based Agents in Computational Biology\.*arXiv preprint arXiv:2503\.00096*, 2025\.URL[https://arxiv\.org/abs/2503\.00096](https://arxiv.org/abs/2503.00096)\.
- Nathani et al\. \(2025\)Deepak Nathani, Lovish Madaan, Nicholas Roberts, et al\.MLGym: A New Framework and Benchmark for Advancing AI Research Agents\.*arXiv preprint arXiv:2502\.14499*, 2025\.URL[https://arxiv\.org/abs/2502\.14499](https://arxiv.org/abs/2502.14499)\.
- O’Donoghue et al\. \(2023\)Odhran O’Donoghue, Aleksandar Shtedritski, John Ginger, et al\.BioPlanner: Automatic Evaluation of LLMs on Protocol Planning in Biology\.*arXiv preprint arXiv:2310\.10632*, 2023\.URL[https://arxiv\.org/abs/2310\.10632](https://arxiv.org/abs/2310.10632)\.
- Panickssery et al\. \(2024\)Arjun Panickssery, Samuel R\. Bowman, and Shi Feng\.LLM Evaluators Recognize and Favor Their Own Generations\.*arXiv preprint arXiv:2404\.13076*, 2024\.URL[https://arxiv\.org/abs/2404\.13076](https://arxiv.org/abs/2404.13076)\.
- Phan et al\. \(2026\)Long Phan, Alice Gatti, Nathaniel Li, et al\.A benchmark of expert\-level academic questions to assess AI capabilities\.*Nature*, 649\(8099\):1139–1146, 2026\.doi:10\.1038/s41586\-025\-09962\-4\.URL[https://doi\.org/10\.1038/s41586\-025\-09962\-4](https://doi.org/10.1038/s41586-025-09962-4)\.Introduces Humanity’s Last Exam \(HLE\); preprint arXiv:2501\.14249 \(2025\)\.
- Raji et al\. \(2021\)Inioluwa Deborah Raji, Emily M\. Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna\.AI and the Everything in the Whole Wide World Benchmark\.*arXiv preprint arXiv:2111\.15366*, 2021\.URL[https://arxiv\.org/abs/2111\.15366](https://arxiv.org/abs/2111.15366)\.
- Rein et al\. \(2023\)David Rein, Betty Li Hou, Asa Cooper Stickland, et al\.GPQA: A Graduate\-Level Google\-Proof Q&A Benchmark\.*arXiv preprint arXiv:2311\.12022*, 2023\.URL[https://arxiv\.org/abs/2311\.12022](https://arxiv.org/abs/2311.12022)\.
- Sharma et al\. \(2025\)Manasi Sharma, Chen Bo Calvin Zhang, Chaithanya Bandi, et al\.ResearchRubrics: A Benchmark of Prompts and Rubrics for Evaluating Deep Research Agents\.*arXiv preprint arXiv:2511\.07685*, 2025\.URL[https://arxiv\.org/abs/2511\.07685](https://arxiv.org/abs/2511.07685)\.
- Shi et al\. \(2026\)Chuhan Shi, Xiaoquan Ren, Sicheng Song, Haobo Li, Rui Sheng, and Yushi Sun\.Are LLMs Ready for Scientific Discovery? A Capability\-Oriented Benchmark for AI Scientists\.*arXiv preprint arXiv:2607\.11079*, 2026\.URL[https://arxiv\.org/abs/2607\.11079](https://arxiv.org/abs/2607.11079)\.
- Shi et al\. \(2024\)Lin Shi, Chiyu Ma, Wenhua Liang, Xingjian Diao, Weicheng Ma, and Soroush Vosoughi\.Judging the Judges: A Systematic Study of Position Bias in LLM\-as\-a\-Judge\.*arXiv preprint arXiv:2406\.07791*, 2024\.URL[https://arxiv\.org/abs/2406\.07791](https://arxiv.org/abs/2406.07791)\.
- Si et al\. \(2024\)Chenglei Si, Diyi Yang, and Tatsunori Hashimoto\.Can LLMs Generate Novel Research Ideas? A Large\-Scale Human Study with 100\+ NLP Researchers\.*arXiv preprint arXiv:2409\.04109*, 2024\.URL[https://arxiv\.org/abs/2409\.04109](https://arxiv.org/abs/2409.04109)\.
- Si et al\. \(2025\)Chenglei Si, Tatsunori Hashimoto, and Diyi Yang\.The Ideation\-Execution Gap: Execution Outcomes of LLM\-Generated versus Human Research Ideas\.*arXiv preprint arXiv:2506\.20803*, 2025\.URL[https://arxiv\.org/abs/2506\.20803](https://arxiv.org/abs/2506.20803)\.
- Siegel et al\. \(2024\)Zachary S\. Siegel, Sayash Kapoor, Nitya Nadgir, Benedikt Stroebl, and Arvind Narayanan\.CORE\-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark\.*arXiv preprint arXiv:2409\.11363*, 2024\.URL[https://arxiv\.org/abs/2409\.11363](https://arxiv.org/abs/2409.11363)\.
- Singh et al\. \(2025\)Shivalika Singh, Yiyang Nan, Alex Wang, et al\.The Leaderboard Illusion\.*arXiv preprint arXiv:2504\.20879*, 2025\.URL[https://arxiv\.org/abs/2504\.20879](https://arxiv.org/abs/2504.20879)\.
- Sivakumar et al\. \(2026\)Aditya Sivakumar, Ashu Singhal, Nicholas Larus\-Stone, and Nithin Parsan\.BenchBench\-Protocol: Evaluating Real\-World Wet\-Lab Protocol Reasoning and Modification\.Technical report, Benchling, aug 2026\.URL[https://www\.benchling\.com/blog/can\-llms\-work\-in\-the\-wet\-lab](https://www.benchling.com/blog/can-llms-work-in-the-wet-lab)\.
- Song et al\. \(2025\)Zhangde Song, Jieyu Lu, Yuanqi Du, et al\.Evaluating Large Language Models in Scientific Discovery\.*arXiv preprint arXiv:2512\.15567*, 2025\.URL[https://arxiv\.org/abs/2512\.15567](https://arxiv.org/abs/2512.15567)\.
- Soskin Hicks et al\. \(2026\)Rebecca Soskin Hicks, Mikhail Trofimov, Dominick Lim, et al\.HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats\.*arXiv preprint arXiv:2604\.27470*, 2026\.URL[https://arxiv\.org/abs/2604\.27470](https://arxiv.org/abs/2604.27470)\.
- Srivastava et al\. \(2022\)Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, et al\.Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models\.*arXiv preprint arXiv:2206\.04615*, 2022\.URL[https://arxiv\.org/abs/2206\.04615](https://arxiv.org/abs/2206.04615)\.
- Starace et al\. \(2025\)Giulio Starace, Oliver Jaffe, Dane Sherburn, et al\.PaperBench: Evaluating AI’s Ability to Replicate AI Research\.*arXiv preprint arXiv:2504\.01848*, 2025\.URL[https://arxiv\.org/abs/2504\.01848](https://arxiv.org/abs/2504.01848)\.
- Sun et al\. \(2023\)Liangtai Sun, Yang Han, Zihan Zhao, et al\.SciEval: A Multi\-Level Large Language Model Evaluation Benchmark for Scientific Research\.*arXiv preprint arXiv:2308\.13149*, 2023\.URL[https://arxiv\.org/abs/2308\.13149](https://arxiv.org/abs/2308.13149)\.
- Verga et al\. \(2024\)Pat Verga, Sebastian Hofstatter, Sophia Althammer, et al\.Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models\.*arXiv preprint arXiv:2404\.18796*, 2024\.URL[https://arxiv\.org/abs/2404\.18796](https://arxiv.org/abs/2404.18796)\.
- Wang et al\. \(2026\)Miles Wang, Robi Lin, Kat Hu, et al\.FrontierScience: Evaluating AI’s Ability to Perform Expert\-Level Scientific Tasks\.*arXiv preprint arXiv:2601\.21165*, 2026\.URL[https://arxiv\.org/abs/2601\.21165](https://arxiv.org/abs/2601.21165)\.
- Wang et al\. \(2023\)Xiaoxuan Wang, Ziniu Hu, Pan Lu, et al\.SciBench: Evaluating College\-Level Scientific Problem\-Solving Abilities of Large Language Models\.*arXiv preprint arXiv:2307\.10635*, 2023\.URL[https://arxiv\.org/abs/2307\.10635](https://arxiv.org/abs/2307.10635)\.
- Wijk et al\. \(2024\)Hjalmar Wijk, Tao Lin, Joel Becker, et al\.RE\-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts\.*arXiv preprint arXiv:2411\.15114*, 2024\.URL[https://arxiv\.org/abs/2411\.15114](https://arxiv.org/abs/2411.15114)\.
- Yao et al\. \(2024\)Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan\.τ\\tau\-bench: A Benchmark for Tool\-Agent\-User Interaction in Real\-World Domains\.*arXiv preprint arXiv:2406\.12045*, 2024\.URL[https://arxiv\.org/abs/2406\.12045](https://arxiv.org/abs/2406.12045)\.
- Ye et al\. \(2024\)Jiayi Ye, Yanbo Wang, Yue Huang, et al\.Justice or Prejudice? Quantifying Biases in LLM\-as\-a\-Judge\.*arXiv preprint arXiv:2410\.02736*, 2024\.URL[https://arxiv\.org/abs/2410\.02736](https://arxiv.org/abs/2410.02736)\.
- Zhang et al\. \(2026\)Jinwei Zhang, Xucheng Liang, Yu Zhang, Ruijie Yu, Xiaokang Yang, Yaohui Jin, and Yanyan Xu\.ChemReason\-Bench: Benchmarking Large Language Models for Procedural Reasoning in Experimental Chemistry\.In*Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 33211–33248, San Diego, California, United States, jul 2026\. Association for Computational Linguistics\.doi:10\.18653/v1/2026\.acl\-long\.1535\.URL[https://aclanthology\.org/2026\.acl\-long\.1535/](https://aclanthology.org/2026.acl-long.1535/)\.
- Zhang et al\. \(2025\)Weichen Zhang, Yiyou Sun, Pohao Huang, Jiayue Pu, Heyue Lin, and Dawn Song\.MIRAGE\-Bench: LLM Agent is Hallucinating and Where to Find Them\.*arXiv preprint arXiv:2507\.21017*, 2025\.URL[https://arxiv\.org/abs/2507\.21017](https://arxiv.org/abs/2507.21017)\.
- Zhao et al\. \(2024\)Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng\.WildChat: 1M ChatGPT Interaction Logs in the Wild\.*arXiv preprint arXiv:2405\.01470*, 2024\.URL[https://arxiv\.org/abs/2405\.01470](https://arxiv.org/abs/2405.01470)\.
- Zheng et al\. \(2023a\)Lianmin Zheng, Wei\-Lin Chiang, Ying Sheng, et al\.LMSYS\-Chat\-1M: A Large\-Scale Real\-World LLM Conversation Dataset\.*arXiv preprint arXiv:2309\.11998*, 2023a\.URL[https://arxiv\.org/abs/2309\.11998](https://arxiv.org/abs/2309.11998)\.
- Zheng et al\. \(2023b\)Lianmin Zheng, Wei\-Lin Chiang, Ying Sheng, et al\.Judging LLM\-as\-a\-Judge with MT\-Bench and Chatbot Arena\.*arXiv preprint arXiv:2306\.05685*, 2023b\.URL[https://arxiv\.org/abs/2306\.05685](https://arxiv.org/abs/2306.05685)\.

## Appendix AAppendix

### A\.1 Scoring rubric v1\.0

The judge\-facing rubric is reproduced below in full, matching version 1\.0 \(2026\-08\-07\)\.

Scoring Rubric — Agentic Scientific Task Benchmark Version: 1\.0\(2026\-08\-07\)

You are evaluating the work of an anonymous AI research agent on a real scientific task submitted by a real user\. The agent ran in a sandbox with shell access, file tools, and web search, and produced a final answer plus \(possibly\) output files\.

#### How to score

- •Every dimension is aninteger 0–10\. Written anchors are given at 0, 3, 5, 8, 10; interpolate for in\-between scores\.
- •Scores must discriminate\.The purpose of this evaluation is to measure how far current models are from expert\-level scientific work\. A score of8 means a domain scientist would accept the work with minor edits\. Reserve 9–10 for genuinely publishable, expert\-grade output\. Do not cluster scores in the 6–7 comfort zone: if the work has real gaps, score it in the 3–5 range; if it is superficial or wrong, score lower still\.
- •Judge what wasactually delivered, not what was promised\. A plan to do an analysis is not an analysis\. A script that was never run produces no results\.
- •Ground every score in evidence: the task prompt, the final answer, the transcript digest, and the artifacts you inspect\. Spot\-check claims against artifacts wherever possible \(e\.g\., does the number quoted in the answer appear in the results file?\)\.
- •If the run was cut off mid\-execution \(a truncation banner will say so\), score the work that exists — do not extrapolate credit for what might have followed — and tagtruncated\_run\.

#### Dimensions

##### 1\.task\_fulfillment— Did it do what was asked?

Coverage of every explicit request and every reasonable implicit requirement, at the depth the user asked for\.

- •10— Every explicit and reasonable implicit requirement fully met at the requested depth; nothing the user would need to ask again for\.
- •8— All major requirements met; at most one minor sub\-request shallow or missing\.
- •5— The core question is addressed, but sub\-requests are dropped, depth is below what was asked, or a deliverable \(e\.g\., “make me a PPT/figure/table”\) is missing\.
- •3— Only a fraction of the request addressed; major deliverables absent\.
- •0— Off\-task, no substantive answer, or answered a different question\.

##### 2\.scientific\_accuracy— Is the science right?

Correctness of scientific claims, methods, statistics, units, formulas, and citations\.

- •10— Methodologically defensible throughout; claims accurate; statistics appropriate and correctly executed; citations real and relevant\.
- •8— Sound overall; minor imprecision that would not change conclusions\.
- •5— Broadly plausible but with unchecked assumptions, questionable method choices, or minor errors that a reviewer would flag\.
- •3— Material scientific errors: wrong method for the question, misused statistics, incorrect units/conversions, or misinterpreted results\.
- •0— Fabricated results or citations, pseudo\-science, or fundamentally wrong\.

##### 3\.reasoning\_quality— Was the approach intelligent?

Planning, problem decomposition, hypothesis\-driven exploration, and error recovery, as visible in the transcript digest\.

- •10— Clear plan, sensible decomposition, adapts intelligently to what it finds, verifies intermediate results before building on them\.
- •8— Good plan and adaptation with occasional inefficiency\.
- •5— Some structure but linear/mechanical; misses obvious checks; recovers from errors slowly or by trial\-and\-error\.
- •3— Little visible planning; flails between approaches; builds on unverified intermediate results\.
- •0— Incoherent; no discernible strategy\.

##### 4\.tool\_use— Were the tools used competently?

Effective use of shell, file tools, and web search: right tool for the job, efficient sequences, graceful recovery from failures\.

- •10— Fluent: efficient commands, sensible environment setup, quick diagnosis and recovery from failures, no wasted cycles\.
- •8— Competent with minor waste \(redundant reads, an avoidable dead end\)\.
- •5— Gets there but inefficiently: repeated failed commands with small tweaks, clumsy environment management, ignores informative error messages\.
- •3— Substantial thrashing: loops of near\-identical failing commands, abandons tools that would have worked, fights the environment instead of adapting\.
- •0— Tool use actively counterproductive or essentially absent when clearly needed\.

##### 5\.data\_handling— Were the user’s files used correctly?

Whether attached files were actually loaded, parsed correctly, sanity\-checked, and faithfully represented\.N/A if the task had no attachments and needed no data\.

- •10— All attachments loaded and parsed correctly; contents sanity\-checked \(shapes, ranges, missingness\); analysis faithful to the actual data\.
- •8— Data used correctly; light on sanity checks\.
- •5— Data loaded but partially used, or used without verification; minor misreadings that don’t invalidate the main result\.
- •3— Attachments ignored, misparsed, or replaced with assumed/simulated stand\-ins without telling the user\.
- •0— Claims to have analyzed data it never opened; invents columns or values\.

##### 6\.artifact\_quality— Are the output files any good?

Completeness, correctness, and usefulness of files produced \(reports, figures, tables, code, processed data\)\.N/A if a prose answer is the natural deliverable and none were expected\.

- •10— Artifacts are complete, correct, well\-organized, and directly usable; code runs as\-is; figures/tables are publication\-quality\.
- •8— Usable with minor cleanup; code essentially correct\.
- •5— Artifacts exist but are incomplete, partially broken, or need real work to use; or key promised artifacts are missing while others exist\.
- •3— Mostly broken, placeholder, or trivial artifacts; promised deliverables absent\.
- •0— No artifacts despite the task clearly requiring them, or artifacts are misleading \(e\.g\., a “results” file of fabricated numbers\)\.

##### 7\.communication— Is the answer well communicated?

Structure, clarity, appropriate length, andlanguage match: the response must be in the language of the user’s prompt\(you must report this explicitly\)\.

- •10— Clear, well\-structured, right length and register for the request; in the prompt’s language; a scientist could act on it immediately\.
- •8— Clear and complete with minor verbosity or structural noise\.
- •5— Understandable but bloated, poorly organized, or padded with boilerplate; or partially in the wrong language\.
- •3— Hard to follow; buries the answer; wrong language for substantial portions\.
- •0— Incoherent or entirely in the wrong language\.

##### 8\.honesty\_calibration— Is the agent honest about what it did?

Absence of hallucination and overclaiming; failures acknowledged rather than papered over; limitations stated\.

- •10— Every claim traceable to work actually done; failures and limitations stated plainly; uncertainty appropriately flagged\.
- •8— Honest overall; minor unflagged uncertainty\.
- •5— Overstates polish or completeness; glosses over steps that failed; presents assumptions as facts\.
- •3— Presents placeholder or simulated numbers as computed results; claims success on visibly failed steps\.
- •0— Systematic fabrication: invented results, citations, or a false narrative of what was done\.

#### Overall score and success flag

- •overall\(0–10\) — Holistic quality of this run as a response to this user’s request\.Not an averageof the dimensions: weight what mattered most for this particular task\.
- •fully\_successful\(boolean\) — Would the user who submitted this task be satisfied with this responsewithout needing any follow\-up? Apply a demanding standard: this is the bar of a paying scientist\-user, not a benevolent grader\.

#### Failure\-mode taxonomy

Tag every failure mode that applies \(empty list if none\)\. Use only these tags:

Table 3:Failure\-mode taxonomyfrom rubric v1\.0\. Judges apply every tag that fits and may apply none\.
#### Confidence

Reportconfidencein\[0,1\]\[0,1\]: how confident you are in your own scores given what you could inspect\. Lower it when artifacts were too large to verify, the run was truncated, or the domain is outside what you could check\.

### A\.2 Results and judge tables

The tables below support Sections[4](https://arxiv.org/html/2608.21601#S4)and[5](https://arxiv.org/html/2608.21601#S5)\. They are numbered in the order they are cited in the main text\.

Table 4:Headline results by model\.Overall is the mean across 178 tasks of the three\-judge mean, with a 95% percentile bootstrap interval over sessions\. The three judge columns give the same quantity computed from that judge alone\. Majority success requires more than half of the three judges to independently mark the run fully successful; unanimous requires all three\. “Scores≥8\\geq 8” is the share of that model’s individual scored judgments at or above the acceptable line\. Computed fromscores\_wide\.csv\.Table 5:Score distribution by judge, counts over the eight rubric dimensions plus the holistic overall, N/A excluded\. Computed fromscores\_wide\.csv\.Table 6:Rubric dimensions, hardest to easiest\.Mean is pooled over all models, judges and tasks\. Percentages use non\-N/A denominators, sonis the applicable count for that dimension and the two conditional dimensions are scored only where they applied\. Per\-model columns use abbreviated names, left to right in leaderboard order\. Computed fromscores\_wide\.csv\.Table 7:Failure\-mode frequencies, as a percentage of judged runs \(assessments\)\.Tags are not exclusive\. Computed from thefailure\_modesfield ofscores\_wide\.csv\.Table 8:Share of runs that used each tool at least once \(%\)\.Computed from thetool\_calls\_by\_toolfield ofrun\_metrics\.csvover all 1,602 runs\.Table 9:Deliverables on disk versus outcome\.Left: runs grouped by whether any output file exists\. Right: runs grouped by file count\. Computed by joiningrun\_metrics\.csvto run\-level means fromscores\_wide\.csv\.Table 10:Paired win rate of the row model against the column model \(%\), same task and same judge, ties excluded\.534 paired comparisons per cell\. Computed fromscores\_wide\.csv\.Table 11:Bradley\-Terry strengths fitted from the paired outcomes\(log scale, centered, 95% bootstrap CI over sessions\), and mean paired score delta against the leader with win and loss shares\. Strengths are identified up to an additive constant, so only differences are interpretable\.Table 12:Paired win rate againstgpt\-5\.6\-solby rubric dimension \(%\), same task and same judge, ties excluded\.For the two conditional dimensions, pairs in which either run was marked not\-applicable are dropped, so those rows rest on fewer decisive pairs than the other six\. Computed fromscores\_wide\.csv\.Table 13:Mean overall by domain and model\.Computed fromscores\_wide\.csvjoined to the domain field ofrun\_metrics\.csv\.Table 14:Effect of attachments on mean overall, by model\.125 of 178 sessions carry at least one file\. Computed by joiningrun\_metrics\.csvto run\-level means\.Table 15:Mean overall by prompt\-length quartile\.Quartiles are formed over the 178 tasks by the byte length of the first user message, which is held in the task registry rather than in the score tables \(Section[Data, code and availability](https://arxiv.org/html/2608.21601#Sx3)\)\. Eight of the nine models are tabulated here for width;gemma\-4\-31b\-itis plotted alongside them in Figure[13](https://arxiv.org/html/2608.21601#S4.F13)and also scores lower on Q4 than on Q1\.Table 16:Run endings and outcomes\.Left: mean overall score and majority\-success rate by final stop reason\. Right: truncation rate by model, with that model’s mean overall for reference\. Computed fromrun\_metrics\.csvjoined to run\-level means\.EndingnOverallMajorityModelTruncatedOverallstop1,3546\.7047\.4%nemotron\-3\-ultra\-550b\-a55b64\.6%2\.78length2242\.220\.0%muse\-spark\-1\.252\.8%4\.27toolUse103\.600\.0%grok\-4\.511\.2%6\.96error140\.000\.0%gemma\-4\-31b\-it7\.9%3\.86Not truncated1,3446\.7347\.8%claude\-opus\-53\.9%7\.61Truncated2582\.180\.0%kimi\-k32\.8%7\.17gpt\-5\.6\-luna1\.7%7\.46gpt\-5\.6\-sol0\.0%8\.04gemini\-3\.6\-flash0\.0%5\.84Table 17:Run\-level Spearman correlates of overall score\(n=1,602n=1\{,\}602\)\. Computed fromrun\_metrics\.csvjoined to run\-level means\.Table 18:Cost and efficiency\.Mean and median inference cost per task, total campaign cost, and two efficiency ratios\. Computed fromrun\_metrics\.csv; total generation cost across all 1,602 runs was $3,649\.18\.Table 19:Failure co\-occurrence,P⁡\(column∣row\)P\(\\text\{column\}\\mid\\text\{row\}\)in percent\.Read a row as: when this failure occurs, how often the column failure also occurs\. Computed from thefailure\_modesfield ofscores\_wide\.csv\.Table 20:Pairwise judge agreement on the holistic overall score\(n=1,602n=1\{,\}602runs\)\.A−BA\-Bis the mean difference in level\. Computed fromscores\_wide\.csv\.Table 21:Judge agreement by rubric dimension\.Means over the three judge pairs, computed on runs where all three judges scored the dimension\. Computed fromscores\_wide\.csv;nnis smaller for the two conditional dimensions because all three judges must have marked them applicable\.Table 22:Calibration\-adjusted self\-preference for the two judges that are also contestants\.All values are mean overall scores\. Computed fromscores\_wide\.csv\.Table 23:gpt\-5\.6\-sol’s deviation from peer consensus, by model judged\.“Peer mean” is the mean of the two other judges’ means for that model\. A single additive strictness term would make the deviation column constant; it is not\. The two smallest magnitudes belong to models the panel already places near 3, where downward deviation is bounded by the floor of the scale, which biases the pooled baseline — and hence the self\-preference estimate of Table[22](https://arxiv.org/html/2608.21601#A1.T22)— toward zero\. Computed from the per\-judge columns of Table[4](https://arxiv.org/html/2608.21601#A1.T4)\.
### A\.3 Supporting tables

Table 24:Exact systems evaluated\.All models were accessed through OpenRouter; the identifier column is the slug passed to the API and the date is the model’s OpenRouter listing date, not a vendor snapshot date\. Context length is the window advertised by the endpoint\. All nine benchmarked models ran under the stockpi0\.84\.0 harness in Modal sandboxes with the same tool set, with thinking levelmaxrequested where the model exposed one, no model\-specific prompting, no sub\-agents and no retries, against zero\-retention endpoints \(Section[Data provenance, consent and privacy](https://arxiv.org/html/2608.21601#Sx2)\)\. The campaign ran between 6 and 12 August 2026\. Context length does not predict truncation \(Section[4\.7](https://arxiv.org/html/2608.21601#S4.SS7)\)\.Table 25:Objective run metrics by model\.Medians over 178 runs except where noted\. “No output files” is the share of runs ending with an empty output tree\. Computed fromrun\_metrics\.csv\.Table 26:Score consistency by model, over all 4,806 assessment\-level holistic scores\. No model is a high\-variance gambler: the standard deviations are similar across the top five, so the ranking reflects level rather than luck\. Computed fromscores\_wide\.csv\.Table 27:Conditional\-dimension applicability and language compliance\.N/A shares are the proportion of assessments on which the judge marked the dimension inapplicable; language compliance is the share of assessments in which the judge recorded that the answer was in the language of the prompt\. Computed fromscores\_wide\.csv\.Table 28:Unsolved tasks and unique solvers by domain\.A task is unsolved if no model’s run was called fully successful by a majority of judges; it has a unique solver if exactly one of the nine models cleared that bar\. Computed fromscores\_wide\.csvandrun\_metrics\.csv\.Table 29:Outcome by tool error rate and by attachment count\.Both are run\-level bins\. The tool\-error relationship is non\-monotone: the best outcomes come from runs that attempted enough to fail occasionally and recovered\. Computed fromrun\_metrics\.csvjoined to run\-level means\.Table 30:Mean overall by attached file type, for file types appearing on at least ten tasks\. “Spread” is best model minus worst model\. File extensions are task\-registry attributes rather than fields of the score tables \(Section[Data, code and availability](https://arxiv.org/html/2608.21601#Sx3)\)\. A task attaching several formats contributes to each, so the rows are not disjoint\.Table 31:Mean overall by sampling batch\.The two batches were drawn one day apart and sampled independently; model ordering is stable across them\. Computed fromrun\_metrics\.csvjoined to run\-level means\.Table 32:Judging operations\.All 4,806 assessments completed; 4,798 produced schema\-valid scores on the first attempt, 6 required two attempts and 2 required three\. Computed fromscores\_wide\.csv\.Total judging cost $996\.99 over 151\.1 hours of judge wall\-clock time; 72 assessments \(1\.5%\) were made with confidence below 0\.7\.

Similar Articles