Knowledge Index of Noah's Ark

arXiv cs.AI Papers

Summary

KINA (Knowledge Index of Noah's Ark) is an 899-item LLM benchmark spanning 261 fine-grained disciplines, introducing formal guarantees for disciplinary representativeness, incentive-aligned annotation via bonus-on-bar tournaments, and bootstrap ranking-stability reporting. Evaluating 42 models, top performers include Gemini-3.1-Pro-Preview (53.17%), Claude-Opus-4.6 (49.92%), and GPT-5.4 (48.55%), revealing a tiered rather than smooth leaderboard structure.

arXiv:2606.05104v1 Announce Type: new Abstract: Knowledge benchmarks for LLMs face three issues: scaling-driven designs that do not operationalize disciplinary representativeness; flat-payment annotation that permits lazy consensus; and unaudited ranking instability under bounded test budgets. We introduce KINA, an 899-item benchmark across 261 fine-grained disciplines, with two formal results. First, we cast representativeness as a coverage-style objective over expert-elicited anchors and operationalize disciplinary representativeness through a proxy, yielding a (1-1/e) greedy approximation (Proposition 1); the guarantee applies to the proxy, not to population representativeness. Second, we prove a bonus-on-bar tournament weakly FOSD-dominates flat payment in released-review quality, with incentive-compatibility threshold B > Delta C / Delta p_min (Theorem 1). Evaluating 42 models from 13 labs, the top model, Gemini-3.1-Pro-Preview, reaches 53.17%, followed by Claude-Opus-4.6 at 49.92% and GPT-5.4 at 48.55%, leaving substantial headroom below saturation. The full leaderboard shows a tiered structure rather than a smooth total order: a small frontier tier lies above 48%, a dense strong-model tier spans roughly 38-45%, and low-performing models remain only modestly above the 10% chance baseline. Tool augmentation adds up to 5.17 points across the five tool-use evaluations, with gains varying substantially across models. We report bootstrap ranking-stability statistics to make bounded-budget variance explicit and to discourage over-interpretation of adjacent ranks.
Original Article
View Cached Full Text

Cached at: 06/05/26, 02:10 AM

# Knowledge Index of Noah’s Ark
Source: [https://arxiv.org/html/2606.05104](https://arxiv.org/html/2606.05104)
Minghao Liu1,2,\*,†Yunze Xiao3Zeqi Zhou4Heli Qi5Yifan Yao2 Meishu Song1,6Kaijing Ma2Xuan Zhang1Sicong Jiang1Yizhe Li1Ningshan Ma7Jie Wei1 Ziniu Li2Minglai Yang1,8Bangya Liu1Yiming Liang2Xiao Fang7Qingcheng Zeng9Jiarui Liu3 Rui Yang10Shen Yan2Wenhao Huang2Jiaheng Liu2Zihan Wang1Weihao Xuan6,†Ge Zhang2,† 12077AI2M\-A\-P3Carnegie Mellon University4Brown University5Waseda University 6The University of Tokyo7Massachusetts Institute of Technology8University of Arizona 9Northwestern University10Duke\-NUS Medical School \*Equal contribution\.†Corresponding authors

###### Abstract

Knowledge benchmarks for LLMs face three issues: scaling\-driven designs that do not operationalize disciplinary representativeness; flat\-payment annotation that permits lazy consensus; and unaudited ranking instability under bounded test budgets\. We introduceKINA111[https://www\.2077ai\.com/datasets/dataset\-kina](https://www.2077ai.com/datasets/dataset-kina), an899899\-item benchmark across261261fine\-grained disciplines, with two formal results\. First, we cast representativeness as a coverage\-style objective over expert\-elicited anchors and operationalize disciplinary representativeness through a proxy, yielding a\(1−1/e\)\(1\-1/e\)greedy approximation \(Proposition[1](https://arxiv.org/html/2606.05104#Thmproposition1)\); the guarantee applies to the proxy, not to population representativeness\. Second, we prove a bonus\-on\-bar tournament weakly FOSD\-dominates flat payment in released\-review quality, with incentive\-compatibility thresholdB\>Δ​C/Δ​pminB\>\\Delta C/\\Delta p\_\{\\min\}\(Theorem[1](https://arxiv.org/html/2606.05104#Thmtheorem1)\)\. Evaluating4242models from1313labs, the top model, Gemini\-3\.1\-Pro\-Preview, reaches53\.17%53\.17\\%, followed by Claude\-Opus\-4\.6 at49\.92%49\.92\\%and GPT\-5\.4 at48\.55%48\.55\\%, leaving substantial headroom below saturation\. The full leaderboard shows a tiered structure rather than a smooth total order: a small frontier tier lies above48%48\\%, a dense strong\-model tier spans roughly3838–45%45\\%, and low\-performing models remain only modestly above the10%10\\%chance baseline\. Tool\-augmentation adds up to5\.175\.17points across the five tool\-use evaluations, with gains varying substantially across models\. We report bootstrap ranking\-stability statistics to make bounded\-budget variance explicit and to discourage over\-interpretation of adjacent ranks\.

## 1Introduction

A knowledge benchmark for LLMs can serve as either a difficulty thermometer or a diagnostic instrument, and current designs lean predominantly toward the former\. The dominant strategies, namely scaling the test set toward hundreds of subfields\[[7](https://arxiv.org/html/2606.05104#bib.bib4)\]and pushing each item toward research\-frontier difficulty\[[22](https://arxiv.org/html/2606.05104#bib.bib5)\], tell us*whether*frontier models struggle, but say less about*where*and*why*they do\. Existing pipelines do not formalize disciplinary representativeness as a selection criterion, do not align reviewer compensation with effort under formal guarantees, and do not routinely report whether observed model rankings remain stable under resampling\. We argue that representativeness, incentive\-aligned review, and ranking stability deserve to be treated as primary design considerations of a knowledge benchmark, and we developKINA, the Knowledge Index of Noah’s Ark, to make that case concrete\.

We motivate each of the three concerns in turn\. Items in current benchmarks are typically organized by subject taxonomy rather than selected by whether they elicit the core competencies of a discipline; aggregate difficulty can therefore be high while the coverage of theoretical pivots remains uneven\. Most review pipelines also compensate reviewers at a flat per\-item rate, so accepting a borderline item incurs no cost and rational reviewers tend to converge on*lazy consensus*; bonus\-heavy designs such as the expert pipeline of GPQA\[[23](https://arxiv.org/html/2606.05104#bib.bib2)\]mitigate but do not formally rule out this failure mode\. Finally, although a compact test set of about10310^\{3\}items is cheap to iterate and resistant to contamination,55\-point gaps near the top of the leaderboard between frontier models on such a set may not survive resampling, and bootstrap\-stable rankings are rarely reported alongside the headline numbers\.

KINAis a benchmark of899899items spanning261261fine\-grained disciplines, designed around two formal results that target the first two of these gaps\. To address representativeness, we operationalize the criterion as*budgeted support centrality*, an operational proxy under which each candidate item is scored against domain\-aligned anchors elicited from experts and items are then selected greedily under capacity constraints\. Proposition[1](https://arxiv.org/html/2606.05104#Thmproposition1)shows that the resulting selection objective is monotone submodular, so greedy selection attains a\(1−1/e\)\(1\-1/e\)approximation of the proxy optimum\. We view the proxy as a tractable surrogate for representativeness rather than a global guarantee on representativeness itself\. To address review quality, we replace flat payment with a*bonus\-on\-bar tournament*: two reviewers evaluate each item, the higher\-scoring reviewer receives a bonusBBprovided that the winning score clears a barτ\\tau, and the principal performs stochastic audits to deter collusive approval\. Under standard assumptions on effort\-induced FOSD, score\-noise independence, and monotonicity of the aggregator \(Assumptions[1](https://arxiv.org/html/2606.05104#Thmassumption1)–[3](https://arxiv.org/html/2606.05104#Thmassumption3)\), Theorem[1](https://arxiv.org/html/2606.05104#Thmtheorem1)establishes that this mechanism strictly improves the quality of released reviews relative to the flat\-payment baseline in the FOSD sense, with a closed\-form bonus calibrationB\>Δ​C/Δ​pminB\>\\Delta C/\\Delta p\_\{\\min\}\. The third gap, ranking stability, is addressed empirically rather than formally, by reporting bootstrap\-based ranking statistics as a standard column on the leaderboard ofKINA\.

We evaluate4242frontier models from1313labs onKINA\. Gemini\-3\.1\-Pro\-Preview leads with an overall accuracy of53\.17%53\.17\\%, followed by Claude\-Opus\-4\.6 and GPT\-5\.4 at49\.92%49\.92\\%and48\.55%48\.55\\%respectively, and the leaderboard remains far from saturation\. Web\-search augmentation yields positive but non\-uniform gains across the five tool\-use evaluations, ranging from\+1\.50\+1\.50to\+5\.17\+5\.17points, which suggests that retrieval contributes in different ways to weak and strong base models\. Per\-discipline analysis reveals heterogeneous spread within the top\-1010models that aggregate accuracy alone would hide: the spread is only9\.839\.83points in Science, while it reaches38\.1638\.16points in Sociology\. We read this contrast as suggestive rather than conclusive, since the disciplines with the largest spread also have the smallest item counts; what we draw from it is that humanities and social\-science content deserves separate reporting at the frontier, not that it is the principal source of differentiation\. We further report bootstrap\-based ranking\-stability statistics in §[5\.2](https://arxiv.org/html/2606.05104#S5.SS2), in order to make the variance under bounded test budgets visible to readers rather than implicit\.

Our contributions are threefold\.

1. 1\.\(§[3\.1](https://arxiv.org/html/2606.05104#S3.SS1)\) formalizes disciplinary representativeness as budgeted support centrality and proves that the resulting selection objective is monotone submodular, which gives a\(1−1/e\)\(1\-1/e\)approximation guarantee on the proxy under greedy selection\.
2. 2\.\(§[3\.2](https://arxiv.org/html/2606.05104#S3.SS2)\) shows that, under standard assumptions on effort\-induced FOSD, score\-noise independence, and monotonicity of the aggregator, a bonus\-on\-bar tournament strictly improves the quality of released reviews relative to the flat\-payment baseline, with a closed\-form bonus calibration\.
3. 3\.\(§[5](https://arxiv.org/html/2606.05104#S5)\) releasesKINAtogether with an evaluation of4242frontier models, per\-discipline scores, web\-search ablations, an analysis of parameter scaling, and bootstrap\-based ranking\-stability statistics\.

Alongside the dataset, we release the annotation and reviewer manuals, the LLM\-judge rubrics, and the evaluation code, to enable replication and extension\.

## 2Related Work

##### Knowledge benchmarks\.

The standard paradigm was set by SuperGLUE\[[30](https://arxiv.org/html/2606.05104#bib.bib11)\]and MMLU\[[11](https://arxiv.org/html/2606.05104#bib.bib9)\], both of which are now saturated by frontier models and have known quality issues\. Post\-hoc analysis of MMLU\[[9](https://arxiv.org/html/2606.05104#bib.bib45)\]estimates a6\.5%6\.5\\%overall error rate, with per\-subject rates exceeding50%50\\%in several domains and rank changes on error\-corrected subsets\. Subsequent work pursues two trajectories\. The*depth\-first*line includes ScienceQA\[[16](https://arxiv.org/html/2606.05104#bib.bib12)\], ARC\-AGI\[[4](https://arxiv.org/html/2606.05104#bib.bib10),[5](https://arxiv.org/html/2606.05104#bib.bib13)\], GPQA\[[23](https://arxiv.org/html/2606.05104#bib.bib2)\], and HLE\[[22](https://arxiv.org/html/2606.05104#bib.bib5)\]\. The*breadth\-first*line includes MMLU\-Pro\[[31](https://arxiv.org/html/2606.05104#bib.bib3)\]and SuperGPQA\[[7](https://arxiv.org/html/2606.05104#bib.bib4)\]\. None of these designs makes*disciplinary representativeness*an explicit selection criterion: items are scored on difficulty or sourced by availability, not on whether they probe central theoretical pivots\.

##### Annotation methodology and incentive design\.

GPQA pioneered an expert\-in\-the\-loop pipeline with∼\\sim$95/hr compensation and three\-stage validation, but covers only three domains\. SuperGPQA scales the same expert\-review template to∼\\sim8080annotators and26,52926\{,\}529items; an initial broader\-crowdsourcing attempt rejected63%63\\%of submissions, motivating a shift to expert\-only annotation\. HLE used an open\-call solicitation with a$​500,000\\mathdollar 500\{,\}000prize pool and a∼\\sim55\-minute verification cap per item\. The post\-release audit\[[36](https://arxiv.org/html/2606.05104#bib.bib8),[29](https://arxiv.org/html/2606.05104#bib.bib33)\]found that29±3\.7%29\\pm 3\.7\\%of biology and chemistry items conflicted with peer\-reviewed literature, and only26%26\\%of audited items passed verification unmodified\. The structural cause is incentive: a flat\-payment cap on review effort, combined with a “stump\-the\-LLM” contributor incentive, drives reviewers toward lazy consensus\. To our knowledge, no prior knowledge benchmark formally analyzes the incentive structure of its review pipeline\.

##### Mechanism design for elicitation\.

Tournament and contest\-based mechanisms have a long history in labor economics\[[14](https://arxiv.org/html/2606.05104#bib.bib46)\]and have been used in NLP for adversarial data collection\[[19](https://arxiv.org/html/2606.05104#bib.bib49)\]\. Our use of a bonus\-on\-bar tournament for annotation review is, to our knowledge, novel in the LLM\-benchmark setting\. Our analysis \(§[3\.2](https://arxiv.org/html/2606.05104#S3.SS2)\) is a comparative claim relative to flat payment under standard FOSD assumptions; it is not a full equilibrium characterization\.

##### Comparison\.

Table[1](https://arxiv.org/html/2606.05104#S2.T1)situatesKINAagainst representative prior benchmarks along five axes\.

Table 1:Comparison ofKINAwith prior knowledge benchmarks\.“Disc\.” = number of disciplines; “Rep\.”==explicit representativeness criterion; “Inc\.”==explicit incentive\-aligned review; “Tool”==tool\-use evaluation reported in original paper\. Best\-saturated overall accuracy of frontier models is approximate\.

## 3Two Formal Guarantees for the Data Pipeline

The data pipeline ofKINAmakes two decisions whose quality dominates everything downstream: which items to keep from a much larger candidate pool, and how to compensate the reviewers who certify the survivors\. We give a formal guarantee at each decision point\. The first guarantee says greedy selection is good enough at covering a discipline, in a sense we make precise below\. The second says a tournament\-style payment scheme is good enough at eliciting reviewer effort, in a similar sense\. The two arguments are independent, but together they pin down what a reader can take from the pipeline ofKINAas a formal claim, as opposed to a methodological choice\. Section[4](https://arxiv.org/html/2606.05104#S4)describes how each guarantee is realized in the actual workflow, and Section[6](https://arxiv.org/html/2606.05104#S6)states what neither guarantee establishes\.

### 3\.1Selection: greedy coverage of a disciplinary prototype

Representativeness is a property of the selected subset, not of items in isolation\. An item is well\-positioned if it materially supports at least one canonical anchor of its discipline; a subset is representative if every important anchor is supported by at least one item it contains\. To make this concrete, domain experts elicit a compact disciplinary prototypeΣd\\Sigma\_\{d\}of methods, problems, theorems, concepts, and applications, and an LLM\-judge scores each candidate itemqqfor how strongly it supports each anchoruu, producingS^dsp​\(q,u\)∈\[0,1\]\\hat\{S\}\_\{d\}^\{\\mathrm\{sp\}\}\(q,u\)\\in\[0,1\]\. We aggregate these scores into a coverage objective:

Fdsp​\(𝒮\)≜∑u∈B¯dμd​\(u\)​maxq∈𝒮⁡S^dsp​\(q,u\),F\_\{d\}^\{\\mathrm\{sp\}\}\(\\mathcal\{S\}\)\\;\\triangleq\\;\\sum\_\{u\\in\\bar\{B\}\_\{d\}\}\\mu\_\{d\}\(u\)\\,\\max\_\{q\\in\\mathcal\{S\}\}\\,\\hat\{S\}\_\{d\}^\{\\mathrm\{sp\}\}\(q,u\),\(1\)with anchor weightsμd​\(u\)≥0\\mu\_\{d\}\(u\)\\geq 0\. Selection picks𝒮d\\mathcal\{S\}\_\{d\}of sizeKdK\_\{d\}to maximizeFdspF\_\{d\}^\{\\mathrm\{sp\}\}, subject to per\-subfield quotas and a near\-duplicate constraint we describe in Appendix[B](https://arxiv.org/html/2606.05104#A2)\.

The objectiveFdspF\_\{d\}^\{\\mathrm\{sp\}\}is monotone and submodular: for each anchoruu, the per\-anchor coveragemaxq∈𝒮⁡S^dsp​\(q,u\)\\max\_\{q\\in\\mathcal\{S\}\}\\hat\{S\}\_\{d\}^\{\\mathrm\{sp\}\}\(q,u\)is monotone and submodular in𝒮\\mathcal\{S\}by a standard max\-coverage argument;FdspF\_\{d\}^\{\\mathrm\{sp\}\}is a nonnegative weighted sum of such functions and inherits both properties\. Greedy maximization under the cardinality constraint\|𝒮d\|=Kd\|\\mathcal\{S\}\_\{d\}\|=K\_\{d\}therefore attains a\(1−1/e\)\(1\-1/e\)approximation of the proxy optimum \(Proposition[1](https://arxiv.org/html/2606.05104#Thmproposition1); full statement and proof in Appendix[B](https://arxiv.org/html/2606.05104#A2)\)\. The lazy\-greedy variant we use is described there as well\.

### 3\.2Review: bonus\-on\-bar tournament

Two reviewers evaluate each item independently\. Under flat per\-item payment, the reviewer’s optimal effort is whatever minimizes private disutility regardless of item quality, which is the failure mode we want to rule out\. We instead pay a base wage plus a single bonus to the reviewer with the higher validated score, conditional on the winning score clearing a minimum barτ\\tau\. The bar exists to deter collusive approval: two reviewers cannot agree to put in low effort and split the bonus, because a low\-quality winner forfeits it\.

We want to show that this scheme strictly raises the quality of released reviews relative to the flat baseline, under transparent assumptions\. We assume effort raises latent review quality in the sense of first\-order stochastic dominance \(Assumption[1](https://arxiv.org/html/2606.05104#Thmassumption1)\), that the principal observes a noisy monotone score whose noise is independent across reviewers \(Assumption[2](https://arxiv.org/html/2606.05104#Thmassumption2)\), and that reviewer cost types are drawn i\.i\.d\. from a continuous distributionGG\(Assumption[3](https://arxiv.org/html/2606.05104#Thmassumption3)\)\. LetΔ​pmin\\Delta p\_\{\\min\}be the minimum gain in winning probability when a reviewer switches from low to high effort, taken over the opponent’s effort choice; positivity ofΔ​pmin\\Delta p\_\{\\min\}is the substantive precondition\.

Under these assumptions, the bonus\-on\-bar tournament is FOSD\-improving: the equilibrium high\-effort rate satisfiesπtour≥G​\(B​Δ​pmin\)≥G​\(0\)=πflat\\pi^\{\\mathrm\{tour\}\}\\geq G\(B\\,\\Delta p\_\{\\min\}\)\\geq G\(0\)=\\pi^\{\\mathrm\{flat\}\}, and released review quality satisfiesYtour⪰FOSDYflatY^\{\\mathrm\{tour\}\}\\succeq\_\{\\mathrm\{FOSD\}\}Y^\{\\mathrm\{flat\}\}\(Theorem[1](https://arxiv.org/html/2606.05104#Thmtheorem1); full statement and proof in Appendix[C](https://arxiv.org/html/2606.05104#A3)\)\. The argument has two pieces: a reviewer with cost typeκi\\kappa\_\{i\}prefers high effort under tournament wheneverB​Δ​p​\(e−i\)≥κiB\\,\\Delta p\(e\_\{\-i\}\)\\geq\\kappa\_\{i\}, so all types withκi≤B​Δ​pmin\\kappa\_\{i\}\\leq B\\,\\Delta p\_\{\\min\}choose high effort regardless of the opponent’s behavior, while under flat payment no positive\-cost type does; the induced quality CDF of a randomly drawn review is therefore stochastically lower under tournament, and a quantile\-coupling step lifts marginal FOSD to any coordinatewise nondecreasing aggregator of the two reviews\.

A useful consequence is that the bonus needed to dominate the high\-effort branch has a closed form: settingB\>Δ​C/Δ​pminB\>\\Delta C/\\Delta p\_\{\\min\}, whereΔ​C\\Delta Cbounds the cost of high effort, makes high effort strictly preferred for every reviewer type \(Corollary[1](https://arxiv.org/html/2606.05104#Thmcorollary1)\)\. More generally,B≥G−1​\(π⋆\)/Δ​pminB\\geq G^\{\-1\}\(\\pi^\{\\star\}\)/\\Delta p\_\{\\min\}delivers any target high\-effort rateπ⋆∈\[0,1\)\\pi^\{\\star\}\\in\[0,1\)\. We use the closed\-form bound in pilot calibration and the more general form when targeting a specific high\-effort fraction in a niche discipline\.

## 4KINAConstruction

This section describes how the formal criteria of §[3\.1](https://arxiv.org/html/2606.05104#S3.SS1)and §[3\.2](https://arxiv.org/html/2606.05104#S3.SS2)are realized in a concrete data\-collection pipeline\. The taxonomy backbone is the U\.S\. Classification of Instructional Programs \(CIP\), which we refine to a three\-level hierarchy of1212disciplines,7070fields, and261261fine\-grained subfields; the full taxonomy with per\-subfield item counts is in Appendix[D](https://arxiv.org/html/2606.05104#A4)\. Among compact knowledge benchmarks \(≤104\\leq 10^\{4\}items\), this is the broadest taxonomy we are aware of, and only SuperGPQA \(285285subfields\) is wider among very large benchmarks, at roughly30×30\\timesthe item count\. EachKINAitem is a pseudo\-multiple\-choice question: the stem contains several substantive statements, and each of the1010options is a combinatorial selection over those statements\. This format reduces chance accuracy from25%25\\%\(44\-option MCQ\) to10%10\\%, weakens process\-of\-elimination shortcuts, and forces joint reasoning over all statements\. Each item ships with the question, the1010options, an option\-level explanation with sources, and the originating material; examples are in Appendix[A](https://arxiv.org/html/2606.05104#A1)\.

### 4\.1Pipeline

![Refer to caption](https://arxiv.org/html/2606.05104v2/x1.png)Figure 1:KINAdata\-collection pipeline\. Topic pre\-approval enforces representativeness via the proxy of Proposition[1](https://arxiv.org/html/2606.05104#Thmproposition1); the double\-blind expert\-review stage instantiates the bonus\-on\-bar tournament of Theorem[1](https://arxiv.org/html/2606.05104#Thmtheorem1); LLM\-as\-judge consensus filters residual ambiguity; an agentic refinement loop addresses boundary defects\.Each candidate item flows through four stages \(Figure[1](https://arxiv.org/html/2606.05104#S4.F1), Table[2](https://arxiv.org/html/2606.05104#S4.T2)\)\.

##### Rule\-based screening\.

Each item passes a uniqueness check \(cosine similarity below0\.80\.8against the existing pool, no duplicate options\), a formatting check \(LaTeX compiles in Markdown\), and a difficulty filter that requires at least three of five flagship LLMs to answer the item incorrectly\. The 3\-of\-5 floor preserves both difficulty headroom and topic diversity at the rate observed during pilot collection\.

##### Expert review under the bonus\-on\-bar tournament\.

Reviewers are recruited from two pools, graduate students at top\-tier global universities and senior industry experts, and each candidate passes a two\-round examination, namely a discipline\-specific depth test in Round 1 and a one\-shot formal item simulation in Round 2; the full annotation and reviewer manuals are in Appendix[E](https://arxiv.org/html/2606.05104#A5)\. Each item is assigned to two independent reviewers under double\-blind allocation, who score it on six rubrics: representativeness and depth, factuality, source reliability, logical rigor, distractor plausibility, and combinatorial validity \(a pseudo\-MCQ must contain at least66statements with no subset relations among options\)\. The validated scoreSiS\_\{i\}is the rubric sum normalized by the rubric weight schedule, and the bonusBBis awarded to the reviewer with the higherSiS\_\{i\}providedmaxi⁡Si≥τ\\max\_\{i\}S\_\{i\}\\geq\\tau\. The principal performs stochastic audits on a random55to10%10\\%of approved items, and an item flagged as flawed triggers a joint penalty on both reviewers, which deters collusive lazy consensus\. We calibrateBBper discipline using Corollary[1](https://arxiv.org/html/2606.05104#Thmcorollary1)from a pilot estimate ofΔ​pmin\\Delta p\_\{\\min\}and per\-discipline costΔ​C\\Delta C, cap reviewer work\-in\-progress at33to55items, and use scarcity\-aware pricing for niche disciplines\.

##### LLM\-as\-judge consensus\.

Three independent LLM judges score each item along four feature axes, namely knowledge coverage, disciplinary uniqueness, socio\-economic impact, and practical value, and analyze the failure pattern of the five flagship LLMs on the item\. An item is admitted if at least two of three judges vote yes; otherwise it is returned to the annotator for revision or discarded\.

##### Agentic refinement of flagged items\.

A manual spot\-check after the consensus stage identified a residual class of boundary defects, primarily context drift between stem and options and ambiguity in the implied causal chain of an explanation\. To address these systematically, items flagged in this spot\-check or returned by the LLM\-judge consensus enter a refinement loop: an upstream agent \(with web access\) searches for counter\-evidence to the explanation, and a downstream agent revises the stem or supplies missing premises in response\. Each revised item is re\-scored by a human reviewer under the Stage 2 rubric and must clear the same bar before re\-entering the pool\. Both agents are instantiated with GPT\-5\.2\-Pro in our run; the loop is model\-agnostic\. Approximately one\-third of items required at least one round of refinement before final acceptance, yielding the final899899items\.

Table 2:Four\-stage construction pipeline\.

### 4\.2Statistics

![Refer to caption](https://arxiv.org/html/2606.05104v2/x2.png)Figure 2:Distribution ofKINAitems across the1212top\-level disciplines\.Figure[2](https://arxiv.org/html/2606.05104#S4.F2)shows the distribution\. Engineering \(34\.26%34\.26\\%\) and Science \(22\.36%22\.36\\%\) dominate, which reflects the larger number of sub\-branches in those areas, butKINAretains substantial coverage in Medicine \(9\.01%9\.01\\%\), Law \(7\.23%7\.23\\%\), Literature and Arts \(7\.23%7\.23\\%\), and Economics, Education, Sociology, History, and Philosophy\. Mean stem length is around9595tokens, and mean explanation length is around210210tokens\.

## 5Experiments

We evaluate4242models from1313AI labs onKINA, spanning closed\-source flagship APIs, open\-source dense and Mixture\-of\-Experts \(MoE\) checkpoints, and reasoning models\. The full model list is in Appendix[F\.1](https://arxiv.org/html/2606.05104#A6.SS1)\. All results are reported asavg@4accuracy at default temperature, with3232K maximum new tokens and a reasoning budget of1616K \(or “Medium”\); prompts are in Appendix[F\.2](https://arxiv.org/html/2606.05104#A6.SS2)\. We organize the discussion around four questions: how the leaderboard looks \(§[5\.1](https://arxiv.org/html/2606.05104#S5.SS1)\), whether the resulting ranking is stable \(§[5\.2](https://arxiv.org/html/2606.05104#S5.SS2)\), whether the tournament\-review prediction holds in our own pipeline \(§[5\.3](https://arxiv.org/html/2606.05104#S5.SS3)\), and what the per\-discipline and tool\-use breakdowns reveal about the diagnostic content ofKINA\(§[5\.4](https://arxiv.org/html/2606.05104#S5.SS4), §[5\.5](https://arxiv.org/html/2606.05104#S5.SS5)\)\.

### 5\.1Overall accuracy

Table[3](https://arxiv.org/html/2606.05104#S5.T3)reports the full4242models onKINA\. Gemini\-3\.1\-Pro\-Preview leads at53\.17%53\.17\\%overall, securing the top score in88of1212disciplines\. Claude\-Opus\-4\.6 \(49\.92%49\.92\\%\) and GPT\-5\.4 \(48\.55%48\.55\\%\) follow\. Within the open\-source ecosystem, Qwen3\.5\-397B\-A17B \(42\.99%42\.99\\%\) outperforms several closed\-source systems, including GPT\-5\.2 \(39\.52%39\.52\\%\) and Doubao\-Seed\-2\.0\-Lite \(41\.49%41\.49\\%\)\. The leaderboard is therefore far from saturation: current model performance remains below55%55\\%, and even the strongest closed\-source frontier system fails on close to half of the items\.

A closer look suggests that the full leaderboard is better understood as a tiered structure rather than a smooth total order\. The top three models form a small frontier tier above48%48\\%\. Below them, a dense strong\-model tier spans roughly38\.01%38\.01\\%–44\.99%44\.99\\%, from DeepSeek\-V3\.2\-Thinking to Doubao\-Seed\-2\.0\-Pro\-260215\. This band contains both closed\-source systems and large open\-source or MoE models, suggesting thatKINAseparates the frontier from the strong non\-frontier group while still preserving resolution within the latter\. However, the small gaps inside this band should not be over\-interpreted as precise rank differences\. The lower end of the leaderboard is also informative\. Since eachKINAitem has1010options, chance accuracy is10%10\\%\. Models in the14%14\\%–25%25\\%range are therefore only modestly above random guessing in raw accuracy terms\. For example, Qwen3\-0\.6B obtains14\.49%14\.49\\%, Mixtral\-8x7B\-Instruct obtains17\.83%17\.83\\%, Qwen3\-1\.7B obtains18\.33%18\.33\\%, and Qwen3\.5\-2B obtains20\.52%20\.52\\%\. These results indicate that small or older models can occasionally solve discipline\-specific items, but their performance is still far from robust knowledge mastery under the pseudo\-MCQ format\.

Two additional patterns are visible from the full table\. Within the Qwen3 family, thinking\-augmented reasoning is scale\-dependent: at8080B\-A3B scale the Thinking variant \(28\.28%28\.28\\%\) trails the Instruct variant \(30\.09%30\.09\\%\), while at3030B and44B the Thinking variants lead\. We hypothesize that extended chain\-of\-thought may introduce noise on items requiring precise factual recall at large scale while compensating for weaker parametric knowledge at small scale\. Claude\-Opus\-4\.6 also shows a markedly skewed disciplinary profile: it ties for first in Philosophy \(36\.54%36\.54\\%\) and Agronomy \(61\.88%61\.88\\%\) but reaches only15\.38%15\.38\\%in History, illustrating why a per\-discipline breakdown is more diagnostic than aggregate accuracy alone\.

Table 3:Full leaderboard of 42 models onKINA, sorted by overall accuracy\. Cell shading scales with score \(darker==higher\)\. Abbreviations: Agr\. \(Agronomy\), Econ\. \(Economics\), Edu\. \(Education\), Eng\. \(Engineering\), Hist\. \(History\), Arts \(Lit\. & Arts\), Mgt\. \(Management\), Med\. \(Medicine\), Phil\. \(Philosophy\), Sci\. \(Science\), Soc\. \(Sociology\)\. Bold==best per column; underline==second\-best\. Evaluation stability table in Appendix[F](https://arxiv.org/html/2606.05104#A6)\.
### 5\.2Ranking stability under bounded test budgets

A central concern with compact benchmarks is whether observed gaps between frontier models survive resampling\. We address this directly with a nonparametric bootstrap: for each ofB=1,000B=1\{,\}000replications, we draw a stratified random subsample ofρ⋅899\\rho\\cdot 899items \(stratified by discipline,ρ∈\{0\.5,0\.7,0\.9\}\\rho\\in\\\{0\.5,0\.7,0\.9\\\}\), recompute every model’s accuracy, and recompute the top\-1010ranking\. We report the average Kendallτ\\taubetween the bootstrap ranking and the full\-sample ranking, and the rank\-11retention rate\.

Table 4:Ranking stability under stratified subsampling\. Top\-1010Kendallτ\\tauaveraged over1,0001\{,\}000replications, with rank\-11retention\. Numbers reflect a preliminary run on the released test set\.Table[4](https://arxiv.org/html/2606.05104#S5.T4)shows that top\-1010rankings remain highly stable even atρ=0\.5\\rho=0\.5, and that Gemini\-3\.1\-Pro\-Preview retains rank\-11in at least94%94\\%of replications\. Pairwise gaps below roughly22percentage points are not statistically resolvable at the95%95\\%level, so we caution against fine\-grained rank claims among models in the compressed mid\-tier \(e\.g\., Doubao\-Lite vs\. Kimi\-K2\.5\)\. The stability analysis provides a quantitative warrant for the compact\-benchmark design without overclaiming resolution that a899899\-item budget cannot support\.

### 5\.3Empirical validation of the tournament mechanism

We test whether the prediction of Theorem[1](https://arxiv.org/html/2606.05104#Thmtheorem1)—that tournament review increases caught\-flaw rates over flat\-payment review—holds in our own pipeline\. We compare two operating regimes from internal logs: a flat\-payment pilot phase \(the first∼\\sim1,0001\{,\}000submitted items, no tournament, no audit\) and a subsequent tournament phase \(with bonus and audit\)\. We measure three quantities\. The reviewer\-asymmetric catch rate is the fraction of flaws identified by reviewer A but missed by reviewer B; under flat payment with identical\-effort baselines this should sit near50%50\\%of all flaws under random allocation, while under tournament the asymmetric incentive should drive both reviewers toward effort, narrowing the gap\. The audit\-flagged rate is the fraction of items approved by both reviewers but later flagged in the principal’s stochastic audit; tournament should reduce this rate\. The total caught\-flaw rate at Stage 2 is the fraction of items revised or rejected\.

Table 5:Tournament vs\. flat\-payment review \(in\-house logs\)\. Higher caught\-flaw rate is better\. Audit\-flagged rate is the fraction of twice\-approved items that fail audit; lower is better\. Preliminary numbers from internal logs\.Table[5](https://arxiv.org/html/2606.05104#S5.T5)shows that the tournament phase increases the caught\-flaw rate from41%41\\%to58%58\\%and reduces the audit\-flagged rate by roughly a factor of2\.52\.5\. The reviewer\-asymmetric catch rate falls under tournament, consistent with both reviewers exerting more effort\. We emphasize that this comparison is observational rather than randomized: the two phases are not interleaved, and we cannot fully rule out time trends or annotator\-pool composition shifts\. A randomized A/B design and a wider replication across disciplines are left to future work; the present result is best read as evidence consistent with Theorem[1](https://arxiv.org/html/2606.05104#Thmtheorem1)rather than as a causal verification\.

### 5\.4Diagnostic findings: tool use and discipline structure

The accuracy table alone hides two findings that we view as the principal diagnostic value ofKINA: how much each model gains from web search, and where in the disciplinary spectrum frontier models actually separate from one another\. We unpack each in turn\.

##### Tool\-use yields consistent but non\-uniform gains\.

Table[6](https://arxiv.org/html/2606.05104#S5.T6)reports tool\-use evaluation on five frontier models, each equipped with the provider’s native web search under unlimited interaction turns\. Web search yields universally positive gains, ranging from\+1\.50\+1\.50to\+5\.17\+5\.17points\. The pattern is non\-monotonic in base capability: both the weakest model in the cohort \(GPT\-5\.2\) and the strongest \(Gemini\-3\.1\-Pro\-Preview\) gain the most, while the GPT\-5\.4 variants gain the least\. We read this as evidence that retrieval serves two distinct functions onKINA—filling parametric gaps in lower\-capability models, and supplying grounding evidence that highly capable models can synthesize into verified reasoning chains—rather than as a strict capability multiplier\. The cohort is small, so the U\-shape should be confirmed at scale before being treated as universal; we restrict the tool\-use cohort to five models because each tool\-use run is roughly4×4\\timesthe inference cost of a base run, and broader sweeps are reserved for the public leaderboard\. Future leaderboards may benefit from reporting tool\-use efficiency \(accuracy gained per search query\) as an independent capability dimension\.

Table 6:Performance comparison of direct inference and tool\-use inference\.
##### Humanities and social\-science content drives discrimination at the frontier\.

Table[7](https://arxiv.org/html/2606.05104#S5.T7)reports per\-discipline statistics over the top\-1010models\. The STEM\-oriented disciplines—Science and Engineering—show the smallest spreads, withΔ=9\.83\\Delta=9\.83and14\.2914\.29respectively; this is consistent with the hypothesis that their content is densely represented in pre\-training corpora and amenable to systematic reasoning, so frontier models converge in performance\. The humanities and social\-science contents show the opposite pattern: Sociology \(Δ=38\.16\\Delta=38\.16\), Management \(32\.7632\.76\), and Literature & Arts \(32\.3132\.31\) display spreads three to four times larger\. History combines very low mean accuracy \(25\.96%25\.96\\%\) with high CV \(34\.6334\.63\), suggesting it remains a persistent challenge across the model spectrum rather than a low\-tier\-only problem\. At the frontier, performance differentiation onKINAis dominated by humanities and social\-science content rather than STEM\-oriented content\. The phenomenon is suggestive rather than conclusive—disciplines with the largestΔ\\Deltaalso have the smallest item counts \(Sociology has1919items, Philosophy1313, History1313\), so part of the spread is sample variance—but the contrast is large enough, and one\-sided enough, that it is worth surfacing as an empirical pattern that compact benchmarks can discover\.

Table 7:Discrimination indices over the top\-1010models\. Mean, SD, CV computed over the top\-1010models ranked by overall accuracy\.Δ=Max−Min\\Delta=\\mathrm\{Max\}\-\\mathrm\{Min\}\. LargestΔ\\Deltain bold\.

### 5\.5Parameter scaling

Figure[3](https://arxiv.org/html/2606.05104#S5.F3)shows scaling curves for the Qwen3 and Qwen3\.5 families, dense and MoE\. Three observations stand out\. First, Qwen3\.5 dense scales at roughly14\.814\.8points per decade of total parameters, nearly double the7\.67\.6points\-per\-decade slope of Qwen3; this is consistent with generational improvements in pre\-training data and recipe contributing substantially toKINAperformance, although confounders such as post\-training pipeline differences cannot be ruled out from this figure alone\. Second, MoE models exceed dense models at matched active parameter counts, suggesting a benefit of sparse activation for knowledge\-intensive tasks of the kindKINAprobes\. Third, the scaling curve shows mild log\-linear saturation at the upper end—Qwen3\.5\-397B\-A17B at42\.99%42\.99\\%versus Qwen3\.5\-122B\-A10B at38\.88%38\.88\\%—indicating that parameter scaling alone may face diminishing returns onKINA, and that the humanities and social\-science gap surfaced in Table[7](https://arxiv.org/html/2606.05104#S5.T7)will likely require methodological rather than purely scale\-based progress\.

![Refer to caption](https://arxiv.org/html/2606.05104v2/kina_scaling_law.png)Figure 3:Parameter scaling onKINA\. Left: dense models, accuracy vs\. total parameters \(log scale\)\. Right: MoE models, accuracy vs\. active parameters \(log scale\)\. Generation\-over\-generation slope increases from roughly7\.67\.6\(Qwen3\) to14\.814\.8\(Qwen3\.5\) points per decade\.

## 6Limitations

We highlight five limitations\.

##### Sample\-size variance\.

KINAis intentionally compact\. While §[5\.2](https://arxiv.org/html/2606.05104#S5.SS2)shows that top\-1010rankings remain stable under50%50\\%subsampling \(Kendallτ≈0\.89\\tau\\approx 0\.89\), pairwise gaps below∼\\sim22percentage points are not statistically resolvable at the95%95\\%level\. Fine\-grained per\-discipline claims with denominators<<3030items should be treated as suggestive rather than confirmatory\.

##### Difficulty drift\.

The 3\-of\-5 flagship\-failure filter \(§[4](https://arxiv.org/html/2606.05104#S4)\) couplesKINA’s difficulty distribution to the capabilities of the five LLMs at construction time\. As frontier models improve, the absolute difficulty ofKINAwill erode and the benchmark will require periodic recalibration\. We commit to releasing aKINA\-vx​\.0x\.0refresh whenever the strongest evaluated model exceeds70%70\\%overall accuracy\.

##### Subjectivity in representativeness elicitation\.

The disciplinary prototypeΣd\\Sigma\_\{d\}used to anchor support centrality \(§[3\.1](https://arxiv.org/html/2606.05104#S3.SS1)\) is elicited from human experts\. Different experts within the same discipline may produce non\-identicalΣd\\Sigma\_\{d\}, and we do not currently audit inter\-expert agreement onΣd\\Sigma\_\{d\}itself\. Proposition[1](https://arxiv.org/html/2606.05104#Thmproposition1)guarantees a\(1−1/e\)\(1\-1/e\)approximation of the operational proxyFdspF\_\{d\}^\{\\mathrm\{sp\}\}, not of an underlying population representativeness quantity\.

##### Tournament theory scope\.

Theorem[1](https://arxiv.org/html/2606.05104#Thmtheorem1)compares tournament with flat payment under FOSD assumptions\. It does not establish truthfulness in the mechanism\-design sense, nor does it rule out collusion in the repeated game; collusion deterrence in our pipeline relies on stochastic principal audits, whose effectiveness is empirical\. The empirical validation in §[5\.3](https://arxiv.org/html/2606.05104#S5.SS3)is observational, not randomized\.

##### Tool\-use evaluation confounds\.

We use each model provider’s native web\-search tool\. Search engine indices, backend processing, and retrieval policies vary across providers and over time\. Tool\-use accuracy should therefore be interpreted as effectiveness of the integrated system, not the foundation model’s intrinsic reasoning capacity\.

##### Cultural and linguistic scope\.

KINA’s factuality bar privileges English\-language Q1 journals, monographs, CSSCI core journals, and authoritative domain sources\. Disciplines with regionally divergent canons \(e\.g\., comparative law, vernacular literature\) may be under\-represented in this anchoring, even when the resulting items are technically correct\. We invite community contributions for non\-English discipline expansions\.

## 7Conclusion

We presentedKINA, a knowledge benchmark designed to be diagnostic rather than merely difficult\. Its design is anchored by two formal results: a\(1−1/e\)\(1\-1/e\)\-approximation guarantee for greedy selection under budgeted support centrality \(Proposition[1](https://arxiv.org/html/2606.05104#Thmproposition1)\), and a first\-order stochastic dominance improvement of released review quality under a bonus\-on\-bar tournament mechanism over flat payment \(Theorem[1](https://arxiv.org/html/2606.05104#Thmtheorem1)\)\. The empirical evaluation of4242frontier models surfaces three findings useful for benchmark methodology: \(i\) ranking stability remains reasonably consistent under50%50\\%stratified subsampling, providing a quantitative warrant for the compact design; \(ii\) tool\-use yields universally positive but non\-uniform gains, with a non\-monotonic relationship between base capability and gain magnitude; \(iii\) discrimination at the frontier is dominated by humanities and social\-science content, not hard sciences—a phenomenon visible only under fine\-grained disciplinary breakdown\.

We releaseKINA, the annotation and reviewer manuals, the LLM\-judge rubrics, and the evaluation framework\. We hope the formal\-incentive methodology travels to other benchmark efforts and that the representativeness\-as\-submodular\-coverage framing finds use beyond the LLM\-evaluation context\.

## 8Contributions and Acknowledgements

KINA is developed through a collaboration among 2077AI, Multimodal Art Projection \(M\-A\-P\), the University of Tokyo, and Carnegie Mellon University\. The collaboration combines community\-driven benchmark development, open\-source AI data infrastructure, and academic research expertise in evaluation methodology and empirical analysis\.

Our team members contribute to the development ofKINAfrom the following perspectives:

- •Benchmark design and disciplinary taxonomy
- •Data collection and annotation management
- •Data quality inspection and audit

- •Incentive mechanism design
- •Model evaluation and result analysis
- •Paper writing and project release

## References

- \[1\]Anthropic\(2026\-02\)Introducing Claude Opus 4\.6\.External Links:[Link](https://www.anthropic.com/news/claude-opus-4-6)Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.4.3.2.1.1)\.
- \[2\]Anthropic\(2026\-02\)Introducing Claude Sonnet 4\.6\.External Links:[Link](https://www.anthropic.com/news/claude-sonnet-4-6)Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.4.3.2.1.1)\.
- \[3\]R\. by arXiv\(2026\)The Llama 4 herd: architecture, training, evaluation, and deployment notes\.External Links:2601\.11659,[Link](https://arxiv.org/abs/2601.11659)Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.10.9.2.1.1)\.
- \[4\]F\. Chollet, M\. Knoop, G\. Kamradt,et al\.\(2025\)ARC prize 2024: technical report\.External Links:2412\.04604,[Link](https://arxiv.org/abs/2412.04604)Cited by:[§2](https://arxiv.org/html/2606.05104#S2.SS0.SSS0.Px1.p1.2)\.
- \[5\]F\. Chollet, M\. Knoop, G\. Kamradt,et al\.\(2026\)ARC\-AGI\-2: a new challenge for frontier AI reasoning systems\.External Links:2505\.11831,[Link](https://arxiv.org/abs/2505.11831)Cited by:[§2](https://arxiv.org/html/2606.05104#S2.SS0.SSS0.Px1.p1.2)\.
- \[6\]G\. DeepMind\(2026\-02\)Gemini\.External Links:[Link](https://deepmind.google/models/gemini/)Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.3.2.2.1.1)\.
- \[7\]X\. Du, Y\. Yao, K\. Ma,et al\.\(2025\)SuperGPQA: scaling LLM evaluation across 285 graduate disciplines\.External Links:2502\.14739,[Link](https://arxiv.org/abs/2502.14739)Cited by:[§1](https://arxiv.org/html/2606.05104#S1.p1.1),[§2](https://arxiv.org/html/2606.05104#S2.SS0.SSS0.Px1.p1.2),[Table 1](https://arxiv.org/html/2606.05104#S2.T1.10.4.2)\.
- \[8\]T\. Gebru, J\. Morgenstern, B\. Vecchione,et al\.\(2021\)Datasheets for datasets\.Communications of the ACM64\(12\),pp\. 86–92\.Cited by:[Appendix G](https://arxiv.org/html/2606.05104#A7.p1.1)\.
- \[9\]A\. P\. Gema, J\. O\. J\. Leang, G\. Hong,et al\.\(2025\)Are we done with mmlu?\.External Links:2406\.04127,[Link](https://arxiv.org/abs/2406.04127)Cited by:[§2](https://arxiv.org/html/2606.05104#S2.SS0.SSS0.Px1.p1.2)\.
- \[10\]A\. Grattafiori, A\. Dubey, A\. Jauhri,et al\.\(2024\)The Llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.10.9.2.1.1)\.
- \[11\]D\. Hendrycks, C\. Burns, S\. Basart,et al\.\(2021\)Measuring massive multitask language understanding\.External Links:2009\.03300,[Link](https://arxiv.org/abs/2009.03300)Cited by:[§2](https://arxiv.org/html/2606.05104#S2.SS0.SSS0.Px1.p1.2),[Table 1](https://arxiv.org/html/2606.05104#S2.T1.8.2.2)\.
- \[12\]A\. Huang, A\. Li, A\. Kong,et al\.\(2026\)Step 3\.5 Flash: open frontier\-level intelligence with 11b active parameters\.External Links:2602\.10604,[Link](https://arxiv.org/abs/2602.10604)Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.16.15.2.1.1)\.
- \[13\]A\. Q\. Jiang, A\. Sablayrolles, A\. Roux,et al\.\(2024\)Mixtral of experts\.External Links:2401\.04088,[Link](https://arxiv.org/abs/2401.04088)Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.17.16.2.1.1)\.
- \[14\]E\. P\. Lazear and S\. Rosen\(1981\)Rank\-order tournaments as optimum labor contracts\.Journal of Political Economy89\(5\),pp\. 841–864\.Cited by:[§2](https://arxiv.org/html/2606.05104#S2.SS0.SSS0.Px3.p1.1)\.
- \[15\]A\. Liu, A\. Mei, B\. Lin,et al\.\(2025\)DeepSeek\-V3\.2: pushing the frontier of open large language models\.External Links:2512\.02556,[Link](https://arxiv.org/abs/2512.02556)Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.11.10.2.1.1)\.
- \[16\]P\. Lu, S\. Mishra, T\. Xia,et al\.\(2022\)Learn to explain: multimodal reasoning via thought chains for science question answering\.InThe 36th Conference on Neural Information Processing Systems \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2606.05104#S2.SS0.SSS0.Px1.p1.2)\.
- \[17\]MiniMax\(2026\-02\)MiniMax M2\.5: built for real\-world productivity\.External Links:[Link](https://www.minimax.io/news/minimax-m25)Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.15.14.2.1.1)\.
- \[18\]G\. L\. Nemhauser, L\. A\. Wolsey, and M\. L\. Fisher\(1978\)An analysis of approximations for maximizing submodular set functions—i\.Mathematical Programming14\(1\),pp\. 265–294\.Cited by:[Appendix B](https://arxiv.org/html/2606.05104#A2.SS0.SSS0.Px6.1.p1.9),[Proposition 1](https://arxiv.org/html/2606.05104#Thmproposition1.p1.6.6)\.
- \[19\]Y\. Nie, A\. Williams, E\. Dinan,et al\.\(2020\)Adversarial nli: a new benchmark for natural language understanding\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,pp\. 4885–4901\.Cited by:[§2](https://arxiv.org/html/2606.05104#S2.SS0.SSS0.Px3.p1.1)\.
- \[20\]OpenAI\(2025\-12\)Introducing GPT\-5\.2\.External Links:[Link](https://openai.com/index/introducing-gpt-5-2/)Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.2.1.2.1.1)\.
- \[21\]OpenAI\(2026\-03\)Introducing GPT\-5\.4\.External Links:[Link](https://openai.com/index/introducing-gpt-5-4/)Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.2.1.2.1.1)\.
- \[22\]L\. Phan, A\. Gatti, N\. Li,et al\.\(2026\)A benchmark of expert\-level academic questions to assess AI capabilities\.Nature649\(8099\),pp\. 1139–1146\.Cited by:[§1](https://arxiv.org/html/2606.05104#S1.p1.1),[§2](https://arxiv.org/html/2606.05104#S2.SS0.SSS0.Px1.p1.2),[Table 1](https://arxiv.org/html/2606.05104#S2.T1.10.6.2.1)\.
- \[23\]D\. Rein, B\. L\. Hou, Stickland,et al\.\(2023\)GPQA: a graduate\-level google\-proof q&a benchmark\.External Links:2311\.12022,[Link](https://arxiv.org/abs/2311.12022)Cited by:[§1](https://arxiv.org/html/2606.05104#S1.p2.2),[§2](https://arxiv.org/html/2606.05104#S2.SS0.SSS0.Px1.p1.2),[Table 1](https://arxiv.org/html/2606.05104#S2.T1.9.3.2)\.
- \[24\]B\. Seed\(2026\-02\)Seed 2\.0 official launch\.External Links:[Link](https://seed.bytedance.com/en/blog/seed-2-0-official-launch)Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.5.4.2.1.1)\.
- \[25\]K\. Team, T\. Bai, Y\. Bai,et al\.\(2026\)Kimi k2\.5: visual agentic intelligence\.arXiv preprint arXiv:2602\.02276\.Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.12.11.2.1.1)\.
- \[26\]Q\. Team, A\. Yang, B\. Yang,et al\.\(2024\)Qwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.9.8.2.1.1)\.
- \[27\]Q\. Team\(2025\-09\)Qwen3\-Max: just scale it\.Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.7.6.2.1.1)\.
- \[28\]Q\. Team\(2026\-02\)Qwen3\.5: accelerating productivity with native multimodal agents\.External Links:[Link](https://qwen.ai/blog?id=qwen3.5)Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.6.5.2.1.1)\.
- \[29\]A\. Tu, W\. Xuan, H\. Qi,et al\.\(2025\)Position: the hidden costs and measurement gaps of reinforcement learning with verifiable rewards\.External Links:2509\.21882,[Link](https://arxiv.org/abs/2509.21882)Cited by:[§2](https://arxiv.org/html/2606.05104#S2.SS0.SSS0.Px2.p1.10)\.
- \[30\]A\. Wang, Y\. Pruksachatkun, N\. Nangia,et al\.\(2019\)Superglue: a stickier benchmark for general\-purpose language understanding systems\.Advances in neural information processing systems32\.Cited by:[§2](https://arxiv.org/html/2606.05104#S2.SS0.SSS0.Px1.p1.2)\.
- \[31\]Y\. Wang, X\. Ma, G\. Zhang,et al\.\(2024\)MMLU\-Pro: a more robust and challenging multi\-task language understanding benchmark\.Advances in Neural Information Processing Systems37,pp\. 95266–95290\.Cited by:[§2](https://arxiv.org/html/2606.05104#S2.SS0.SSS0.Px1.p1.2),[Table 1](https://arxiv.org/html/2606.05104#S2.T1.10.5.1.1)\.
- \[32\]xAI Team\(2025\-11\)Grok\-4\-1 model card\.External Links:[Link](https://data.x.ai/2025-11-17-grok-4-1-model-card.pdf)Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.13.12.2.1.1)\.
- \[33\]A\. Yang, A\. Li, B\. Yang,et al\.\(2025\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.7.6.2.1.1),[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.8.7.2.1.1)\.
- \[34\]A\. Yang, B\. Yang, B\. Hui,et al\.\(2024\)Qwen2 technical report\.External Links:2407\.10671,[Link](https://arxiv.org/abs/2407.10671)Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.9.8.2.1.1)\.
- \[35\]A\. Zeng, X\. Lv, Z\. Hou,et al\.\(2026\)GLM\-5: from vibe coding to agentic engineering\.arXiv preprint arXiv:2602\.15763\.Cited by:[Table 13](https://arxiv.org/html/2606.05104#A6.T13.1.14.13.2.1.1)\.
- \[36\]W\. Zhai, Z\. Wang, J\. Wang,et al\.\(2026\)HLE\-verified: a systematic verification and structured revision of humanity’s last exam\.External Links:2602\.13964,[Link](https://arxiv.org/abs/2602.13964)Cited by:[§2](https://arxiv.org/html/2606.05104#S2.SS0.SSS0.Px2.p1.10)\.

## Appendix AData Samples

Pseudo\-Multi\-Choice Sample \(Some Content Omitted \)Discipline: EngineeringQuestion: Which of the following options are correct?1\.D∈ℝ6​M×3​NcD\\in\\mathbb\{R\}^\{6M\\times 3N\_\{c\}\}andMMis a diagonal matrix\. By definingN=D⊤​M−1​DN=D^\{\\top\}M^\{\-1\}D, it follows thatN∈ℝ3​Nc×3​NcN\\in\\mathbb\{R\}^\{3N\_\{c\}\\times 3N\_\{c\}\}\. The Schur complement is given byS=N\+M^−B​E−1​CS=N\+\\hat\{M\}\-BE^\{\-1\}C\. SinceM^\\hat\{M\}is diagonal andB​E−1​CBE^\{\-1\}Cis a block\-diagonal matrix with3×33\\times 3blocks,SSpossesses the identical sparsity pattern asNN\.2\. In the contact pairs of granular dynamics, the friction cone constraint requires that the magnitude of the tangential force impulse vector does not exceed the product of the static friction coefficient and the normal force impulse\. When the static friction coefficient is 0\.3 and the normal force impulse is 10 N⋅\\cdots, the maximum possible magnitude of the tangential force impulse vector is 3 N⋅\\cdots, and this value satisfies the friction cone constraint\.3\. The AMEN Cross complexity bound in this setting scales as𝒪​\(r3​N\)\\mathcal\{O\}\(r^\{3\}N\)\(linear in the number of unknownsNN\)\. ForN=106N=10^\{6\}and maximum TT\-rankr=10r=10, the computational cost is approximately10910^\{9\}operations; however, this contradicts the claim of "sublinear inNN" \(a linear scaling inNNcannot be sublinear\)\.4\. Using the time\-stepping update formula:M​\(vk\+1−vk\)=Δ​t​fB\+∑i∈A​\(qk,δ\)Di​γiM\\left\(v^\{k\+1\}\-v^\{k\}\\right\)=\\Delta t\\,f\_\{B\}\+\\sum\_\{i\\in A\(q^\{k\},\\delta\)\}D\_\{i\}\\gamma\_\{i\}\. GivenΔ​t=0\.1​s\\Delta t=0\.1\\ \\text\{s\},M=2​IM=2I,vk=3​m/sv^\{k\}=3\\ \\text\{m/s\},Δ​t​fB=4​N s\\Delta tf\_\{B\}=4\\ \\text\{N s\}, and∑Di​γi=6​N s\\sum D\_\{i\}\\gamma\_\{i\}=6\\ \\text\{N s\}, we calculatevk\+1v^\{k\+1\}as follows:vk\+1\\displaystyle v^\{k\+1\}=vk\+1M​\(Δ​t​fB\+∑Di​γi\)=3\+12​\(4\+6\)=8​m/s\.\\displaystyle=v^\{k\}\+\\frac\{1\}\{M\}\\left\(\\Delta t\\,f\_\{B\}\+\\sum D\_\{i\}\\gamma\_\{i\}\\right\)=3\+\\frac\{1\}\{2\}\\left\(4\+6\\right\)=8\\ \\text\{m/s\}\.5\. WithN=D⊤​M−1​DN=D^\{\\top\}M^\{\-1\}D, forM=2M=2rigid bodies \(yielding6​M=126M=12degrees of freedom, DOF\) andNc=1N\_\{c\}=1contact \(yielding3​Nc=33N\_\{c\}=3multipliers\),NNis a3×33\\times 3matrix, as it is defined by the productD⊤​M−1​DD^\{\\top\}M^\{\-1\}DwhereD∈ℝ12×3D\\in\\mathbb\{R\}^\{12\\times 3\},M−1∈ℝ12×12M^\{\-1\}\\in\\mathbb\{R\}^\{12\\times 12\}, and thus acts on the contact multiplier space \(not generalized\-velocity space\)\.6\. For the position\-based normal complementarity constraint:γi,n≥0,Φi​\(q\)≥0,Φi​\(q\)​γi,n=0\\gamma\_\{i,n\}\\geq 0,\\ \\Phi\_\{i\}\(q\)\\geq 0,\\ \\Phi\_\{i\}\(q\)\\gamma\_\{i,n\}=0\. IfΦi​\(q\)=0\.5​m\\Phi\_\{i\}\(q\)=0\.5\\ \\text\{m\}andγi,n=2​N s\\gamma\_\{i,n\}=2\\ \\text\{N s\}, the constraint isnotsatisfied: while both quantities are nonnegative, their product is1​m N s≠01\\ \\text\{m N s\}\\neq 0, violating the complementarity condition \(Φi​\(q\)​γi,n=0\\Phi\_\{i\}\(q\)\\gamma\_\{i,n\}=0\)\.7\. For TT matrix\-vector products used in the paper’s TT\-based preconditioner, the complexity is𝒪​\(r2​N​log⁡N\)\\mathcal\{O\}\(r^\{2\}N\\log N\)\. While this complexity scales linearly inNN\(up to a logarithmic factorlog⁡N\\log N\), the cost is still considered \*asymptotically sublinear inNN\* when normalized byNN—or more precisely, \*sublinear in the sense of superlinear scaling avoidance\*—becauselog⁡N\\log Ngrows much slower thanNNitself\. WhenNNdoubles, the cost increases by a factor of approximately2​log⁡\(2​N\)/log⁡N≈2​\(1\+log⁡2/log⁡N\)2\\log\(2N\)/\\log N\\approx 2\(1\+\\log 2/\\log N\), which remains close to 2 \(linear scaling\) but avoids the superlinear growth that would violate sublinearity claims\.Options: A\) 1,2B\) 2,3C\) 3,4D\) 4,5E\) 5,6F\) 6,7G\) 2,4H\) 3,5I\) 4,6J\) 5,7Explanations of A\-J A\) Statement 1 is correct in the Corona et al\. setting: the excerpt definesD∈ℝ6​M×3​NcD\\in\\mathbb\{R\}^\{6M\\times 3N\_\{c\}\}, uses a diagonal matrixMM, and formsN=D⊤​M−1​DN=D^\{\\top\}M^\{\-1\}Dsuch thatN∈ℝ3​Nc×3​NcN\\in\\mathbb\{R\}^\{3N\_\{c\}\\times 3N\_\{c\}\}\. The excerpt further notesh​a​t​M\\\\ hat\{M\}is diagonal andB​E−1​CBE^\{\-1\}Cis block\-diagonal \(with3×33\\times 3blocks\), implyingSSshares the sparsity pattern ofNN\.Statement 2: The friction cone constraint establishes that the magnitude of the tangential force impulse vector cannot exceed the product of the static friction coefficient \(μi\\mu\_\{i\}\) and the normal force impulse \(γi​n\\gamma\_\{in\}\)\. Forμi\\mu\_\{i\}= 0\.3 andγi​n\\gamma\_\{in\}= 10 N·s, the maximum allowable tangential force impulse magnitude is 0\.3 \* 10 = 3 N·s, which directly satisfies the constraint’s requirement\.Explanations of other options are omittedSources of A\-J: A\) Corona, E\., Gorsich, D\., Jayakumar, P\., & Veerapaneni, S\. \(2019\)\. Tensor train accelerated solvers for nonsmooth rigid body dynamics\.Applied Mechanics Reviews, 71\(5\), 050804\.![[Uncaptioned image]](https://arxiv.org/html/2606.05104v2/sample1_fig1.png)Sources of other options are omittedQuestion Source: Corona, E\., Gorsich, D\., Jayakumar, P\., & Veerapaneni, S\. \(2019\)\. Tensor train accelerated solvers for nonsmooth rigid body dynamics\.Applied Mechanics Reviews, 71\(5\), 050804\.Question Materail: Corona, E\., Gorsich, D\., Jayakumar, P\., & Veerapaneni, S\. \(2019\)\. Tensor train accelerated solvers for nonsmooth rigid body dynamics\.Applied Mechanics Reviews, 71\(5\), 050804\.![[Uncaptioned image]](https://arxiv.org/html/2606.05104v2/sample1_fig1.png)![[Uncaptioned image]](https://arxiv.org/html/2606.05104v2/sample1_fig2.png)![[Uncaptioned image]](https://arxiv.org/html/2606.05104v2/sample1_fig3.png)![[Uncaptioned image]](https://arxiv.org/html/2606.05104v2/sample1_fig4.png)![[Uncaptioned image]](https://arxiv.org/html/2606.05104v2/sample1_fig5.png)![[Uncaptioned image]](https://arxiv.org/html/2606.05104v2/sample1_fig6.png)![[Uncaptioned image]](https://arxiv.org/html/2606.05104v2/sample1_fig7.png)Correct Answer:A

Standard Sample \(Some Content Omitted \)Discipline: ScienceQuestion:In the extremely complex physical environment of the Galactic Bulge, observed Planetary Nebulae \(P​NPN\) are typically subjected to severe interstellar reddening\. Furthermore, high\-excitation central stars \(Teff\>150,000T\_\{\\mathrm\{eff\}\}\>150,000K\) can produce hard ultraviolet photons capable of doubly ionizing helium \(He\+\+\\mathrm\{He^\{\+\+\}\}\)\. Accurately determining the ionizing photon production rateQ​\(H0\)Q\(\\mathrm\{H^\{0\}\}\)for such objects requires not only correcting for line\-of\-sight extinction but also verifying the ionization structure through radio continuum observations\.Observational Data and Analysis Clues:Data 1 \(Spectroscopy\): The observed Hβ\\betaflux isFH​βobs=2\.40×10−13F\_\{\\mathrm\{H\\beta\}\}^\{\\mathrm\{obs\}\}=2\.40\\times 10^\{\-13\}erg s\-1cm\-2\. The observed Balmer decrement isFH​α/FH​β=12\.0F\_\{\\mathrm\{H\\alpha\}\}/F\_\{\\mathrm\{H\\beta\}\}=12\.0\. Under the condition ofTe=12,000T\_\{e\}=12,000K, the theoretical Case B ratio is2\.852\.85\. The extinction law followsf​\(H​α\)=−0\.35f\(\\mathrm\{H\\alpha\}\)=\-0\.35\(reddening curve offset relative to Hβ\\beta\)\.Data 2 \(Ionization Structure\): The nebula shows high excitation\. Spectral fitting yields a helium abundance ofy=nHe/nH=0\.12y=n\_\{\\mathrm\{He\}\}/n\_\{\\mathrm\{H\}\}=0\.12, with40%40\\%of helium in the doubly ionized state \(He\+\+\\mathrm\{He^\{\+\+\}\}\) and the remainder in the singly ionized state \(He\+\\mathrm\{He^\{\+\}\}\)\.Data 3 \(Radio Observation\): The Very Large Array \(V​L​AVLA\) measured a radio flux density ofSν=15\.2S\_\{\\nu\}=15\.2mJy at55GHz\. The nebula is optically thin in the radio band\.Data 4 \(Geometry and Environment\): The distance to the nebula isd=8\.1d=8\.1kpc\. The electron temperature is measured atTe=12,000T\_\{e\}=12,000K\. Due to gravitational potential and evolutionary constraints, the total ionized gas massMgasM\_\{\\mathrm\{gas\}\}must be less than0\.6​M⊙0\.6M\_\{\\odot\}\.Physical Constants Reference:AtTe=12,000T\_\{e\}=12,000K,αB=2\.22×10−13\\alpha\_\{B\}=2\.22\\times 10^\{\-13\}cm3s\-1AtTe=12,000T\_\{e\}=12,000K,αH​βeff=2\.58×10−14\\alpha\_\{\\mathrm\{H\\beta\}\}^\{\\mathrm\{eff\}\}=2\.58\\times 10^\{\-14\}cm3s\-1Radio\-to\-Hβ\\betaflux conversion formula \(forTe≈10,000T\_\{e\}\\approx 10,000\-15,00015,000K\):Q​\(H0\)\\displaystyle Q\(\\mathrm\{H^\{0\}\}\)≈7\.54×1046×Sν×d2×Te−0\.45\\displaystyle\\approx 54\\times 0^\{46\}\\times S\_\{\\nu\}\\times d^\{2\}\\times T\_\{e\}^\{\-0\.45\}Integrating the extinction correction and radio observation data, calculate the most robust hydrogen ionizing photon production rateQ​\(H0\)Q\(\\mathrm\{H^\{0\}\}\)for the central star of this nebula\.Options: A\)5\.60×1049​s−15\.60\\times 10^\{49\}s^\{\-1\}B\)1\.09×1045​s−11\.09\\times 10^\{45\}s^\{\-1\}C\)3\.38\.×1045s−13\.38\.\\times 10^\{45\}s^\{\-1\}D\)1\.38×1034​s−11\.38\\times 10^\{34\}s^\{\-1\}E\)1\.36×1048​s−11\.36\\times 10^\{48\}s^\{\-1\}F\)6\.15×1049​s−16\.15\\times 10^\{49\}s^\{\-1\}G\)3\.94×1044​s−13\.94\\times 10^\{44\}s^\{\-1\}H\)8\.22×1044​s−18\.22\\times 10^\{44\}s^\{\-1\}I\)3\.17×1049​s−13\.17\\times 10^\{49\}s^\{\-1\}J\)4\.52×1047​s−14\.52\\times 10^\{47\}s^\{\-1\}Explanations of A\-J B\) Step 1: Calculate the Extinction Correction\.We use the difference between the observed and theoretical Balmer decrements to derive the reddening constantc​\(H​β\)c\(\\mathrm\{H\\beta\}\)\. According to the formula:FH​αFH​β\\displaystyle\\frac\{F\_\{\\mathrm\{H\\alpha\}\}\}\{F\_\{\\mathrm\{H\\beta\}\}\}=\(FH​αFH​β\)theo×10c​\(H​β\)\\displaystyle=\\left\(\\frac\{F\_\{\\mathrm\{H\\alpha\}\}\}\{F\_\{\\mathrm\{H\\beta\}\}\}\\right\)\_\{\\text\{theo\}\}\\times 0^\{c\(\\mathrm\{H\\beta\}\)\}Substituting the data:12\.0=2\.85×10c​\(H​β\)×\(0−\(−0\.35\)\)\\displaystyle 20=85\\times 0^\{c\(\\mathrm\{H\\beta\}\)\\times\(0\-\(\-0\.35\)\)\}4\.21=100\.35×c​\(H​β\)4\.21=10^\{0\.35\\times c\(\\mathrm\{H\\beta\}\)\}Taking the logarithm:0\.35×c​\(H​β\)=log10⁡\(4\.21\)≈0\.6240\.35\\times c\(\\mathrm\{H\\beta\}\)=\\log\_\{10\}\(4\.21\)\\approx 0\.624Solving this givesc​\(H​β\)≈1\.78c\(\\mathrm\{H\\beta\}\)\\approx 1\.78\. Thus, the dereddened \(true\) Hβ\\betaflux is:FH​βdered=FH​βobs×101\.78=2\.40×10−13×60\.26\\displaystyle F\_\{\\mathrm\{H\\beta\}\}^\{\\text\{dered\}\}=F\_\{\\mathrm\{H\\beta\}\}^\{\\mathrm\{obs\}\}\\times 0^\{1\.78\}=40\\times 0^\{\-13\}\\times 026≈1\.446×10−11​erg s−1​cm−2\\displaystyle\\approx 446\\times 0^\{\-11\}\\text\{ erg s\}^\{\-1\}\\text\{ cm\}^\{\-2\}Step 2: EstimateQ​\(H0\)Q\(\\mathrm\{H^\{0\}\}\)Based on Optical Data\.First, calculate the Hβ\\betaluminosity:LH​β=4​π​d2​FH​βdered\\displaystyle L\_\{\\mathrm\{H\\beta\}\}=4\\pi d^\{2\}F\_\{\\mathrm\{H\\beta\}\}^\{\\text\{dered\}\}≈4​π×\(2\.5×1022​cm\)2×1\.446×10−11≈1\.135×1035​erg s−1\\displaystyle\\approx 4\\pi\\times\(5\\times 0^\{22\}\\text\{ cm\}\)^\{2\}\\times 446\\times 0^\{\-11\}\\approx 135\\times 0^\{35\}\\text\{ erg s\}^\{\-1\}The corresponding Hβ\\betaphoton emission rate:NH​β=LH​β/h​νH​β≈2\.78×1046​s−1\\displaystyle N\_\{\\mathrm\{H\\beta\}\}=L\_\{\\mathrm\{H\\beta\}\}/h\\nu\_\{\\mathrm\{H\\beta\}\}\\approx 78\\times 0^\{46\}\\text\{ s\}^\{\-1\}\(whereh​νH​β≈4\.08×10−12h\\nu\_\{\\mathrm\{H\\beta\}\}\\approx 4\.08\\times 10^\{\-12\}erg\)\.Using the ratio of recombination coefficients:Q​\(H0\)opt=NH​β×\(αB/αH​βeff\)≈2\.78×1046×8\.60≈2\.39×1047​s−1\\displaystyle Q\(\\mathrm\{H^\{0\}\}\)\_\{\\text\{opt\}\}=N\_\{\\mathrm\{H\\beta\}\}\\times\(\\alpha\_\{B\}/\\alpha\_\{\\mathrm\{H\\beta\}\}^\{\\mathrm\{eff\}\}\)\\approx 78\\times 0^\{46\}\\times 60\\approx 39\\times 0^\{47\}\\text\{ s\}^\{\-1\}Step 3: EstimateQ​\(H0\)Q\(\\mathrm\{H^\{0\}\}\)Based on Radio Data\.Radio continuum radiation originates from free\-free emission and is unaffected by dust extinction, making it much more robust for objects buried in the Galactic Bulge\. Using the provided radio conversion formula:Q​\(H0\)radio=7\.54×1046×Sν×d2×Te−0\.45\\displaystyle Q\(\\mathrm\{H^\{0\}\}\)\_\{\\text\{radio\}\}=54\\times 0^\{46\}\\times S\_\{\\nu\}\\times d^\{2\}\\times T\_\{e\}^\{\-0\.45\}Substituting the values \(noteSν=15\.2S\_\{\\nu\}=15\.2mJy=0\.0152=0\.0152Jy\):Q​\(H0\)radio=7\.54×1046×0\.0152×8\.12×\(12,000\)−0\.45\\displaystyle Q\(\\mathrm\{H^\{0\}\}\)\_\{\\text\{radio\}\}=54\\times 0^\{46\}\\times 0152\\times 1^\{2\}\\times\(2,00\)^\{\-0\.45\}Calculating the temperature term:\(12,000\)−0\.45≈0\.0145\(12,000\)^\{\-0\.45\}\\approx 0\.0145\.Q​\(H0\)radio=7\.54×1046×0\.0152×65\.61×0\.0145≈1\.09×1045​s−1\\displaystyle Q\(\\mathrm\{H^\{0\}\}\)\_\{\\text\{radio\}\}=54\\times 0^\{46\}\\times 0152\\times 561\\times 0145\\approx 09\\times 0^\{45\}\\text\{ s\}^\{\-1\}Step 4: Physical Robustness Assessment and Conclusion\.A significant discrepancy is found:Q​\(H0\)optQ\(\\mathrm\{H^\{0\}\}\)\_\{\\text\{opt\}\}is much larger thanQ​\(H0\)radioQ\(\\mathrm\{H^\{0\}\}\)\_\{\\text\{radio\}\}\. In the direction of the Galactic Bulge, due to variations in the local extinction law \(f​\(λ\)f\(\\lambda\)\) and complex background scattering, fluxes corrected via the Balmer decrement often carry huge systematic errors\. Conversely, the optically thin radio flux is directly proportional to the total number of ionizing photons and is independent of the filling factor or extinction models\. Therefore, physically, the result derived from radio data,1\.09×10451\.09\\times 10^\{45\}s\-1, is considered reliable\.Explanations of other options are omittedSources of A\-J: B\) Aksaker, N\., Demirci, A\., Erzincan, N\., & Akyuz, A\. \(2025\)\. Photoionization Modeling of Planetary Nebulae in the Galactic Bulge\.Advances in Space Research ![[Uncaptioned image]](https://arxiv.org/html/2606.05104v2/sample2_fig1.png)![[Uncaptioned image]](https://arxiv.org/html/2606.05104v2/sample2_fig2.png)![[Uncaptioned image]](https://arxiv.org/html/2606.05104v2/sample2_fig3.png)![[Uncaptioned image]](https://arxiv.org/html/2606.05104v2/sample2_fig4.png)Sources of other options are omittedQuestion Source: Aksaker, N\., Demirci, A\., Erzincan, N\., & Akyuz, A\. \(2025\)\. Photoionization Modeling of Planetary Nebulae in the Galactic Bulge\.Advances in Space Research [https://www\.sciencedirect\.com/science/article/abs/pii/S0273117725011755](https://www.sciencedirect.com/science/article/abs/pii/S0273117725011755)Question Materail: ![[Uncaptioned image]](https://arxiv.org/html/2606.05104v2/sample2_fig1.png)![[Uncaptioned image]](https://arxiv.org/html/2606.05104v2/sample2_fig2.png)![[Uncaptioned image]](https://arxiv.org/html/2606.05104v2/sample2_fig3.png)![[Uncaptioned image]](https://arxiv.org/html/2606.05104v2/sample2_fig4.png)Correct Answer:B

## Appendix BBudgeted support centrality under domain alignment

##### Scope\.

This appendix specifies a budgeted, operational proxy for domain support centrality used for filtering and final selection\. Its role is to privilege items that are well supported by domain\-aligned local structure, while reducing reliance on items that are idiosyncratic, weakly grounded, ambiguously specified, or near\-duplicative\. The proxy is not claimed to identify an underlying population quantity without additional modeling assumptions\. Bootstrap lower percentiles are used as robustness\-oriented summaries rather than formal lower confidence bounds\. Complexity statements refer to LLM\-call counts and sparse pair scoring under fixed hyperparameters\. The structured signatureψ​\(q\)\\psi\(q\)is used as a low\-cost operational representation; we do not assume that it exhausts the latent disciplinary structure\.

##### Notation\.

For a domaindd, let the raw pool be𝒬~d=\{q1,…,qNd\}\\widetilde\{\\mathcal\{Q\}\}\_\{d\}=\\\{q\_\{1\},\\dots,q\_\{N\_\{d\}\}\\\}\. Experts provide a compact prototype

Σd=\(ℳd,𝒫d,𝒯d,𝒞d,𝒜d\),\\Sigma\_\{d\}=\(\\mathcal\{M\}\_\{d\},\\mathcal\{P\}\_\{d\},\\mathcal\{T\}\_\{d\},\\mathcal\{C\}\_\{d\},\\mathcal\{A\}\_\{d\}\),where𝒜d⊂𝒬~d\\mathcal\{A\}\_\{d\}\\subset\\widetilde\{\\mathcal\{Q\}\}\_\{d\}is a small anchor set used only as reference items \(not as final candidates\)\. Leth:𝒬~d→ℋdh:\\widetilde\{\\mathcal\{Q\}\}\_\{d\}\\to\\mathcal\{H\}\_\{d\}be a coarse stratum map,B¯d\\bar\{B\}\_\{d\}a deduplicated reference bank,μd\\mu\_\{d\}a distribution onB¯d\\bar\{B\}\_\{d\},𝒩k​\(q\)⊆B¯d∖\{q\}\\mathcal\{N\}\_\{k\}\(q\)\\subseteq\\bar\{B\}\_\{d\}\\setminus\\\{q\\\}a retrieved neighborhood \(truncated if fewer thankkitems are available\),𝒢​\(q\):=𝒩k​\(q\)∪𝒜d\\mathcal\{G\}\(q\):=\\mathcal\{N\}\_\{k\}\(q\)\\cup\\mathcal\{A\}\_\{d\}the sparse comparison set,𝒦d\\mathcal\{K\}\_\{d\}the accepted shortlist, and𝒮d⊆𝒦d\\mathcal\{S\}\_\{d\}\\subseteq\\mathcal\{K\}\_\{d\}the final set of sizeKdK\_\{d\}\. IfB¯d=∅\\bar\{B\}\_\{d\}=\\varnothingor𝒜d=∅\\mathcal\{A\}\_\{d\}=\\varnothing, we skip automatic selection and send the domain to manual review\.

##### Item alignment and reference bank\.

A single item\-level LLM call returns

\(ψ​\(q\),A~d​\(q\)\):=Γθ,d​\(q\),ψ​\(q\)\\displaystyle\(\\psi\(q\),\\widetilde\{A\}\_\{d\}\(q\)\)=\\Gamma\_\{\\theta,d\}\(q\),\\psi\(q\)=\(mq,pq,tq,cq\),\\displaystyle=\(m\_\{q\},p\_\{q\},t\_\{q\},c\_\{q\}\),wheremqm\_\{q\}is the method family,pqp\_\{q\}the core principle,tqt\_\{q\}the answer type, andcqc\_\{q\}salient constraints\. With a small item label set

ℒditem=\{\(q,yA​\(q\)\)\},yA​\(q\)∈\{0,1\},\\mathcal\{L\}\_\{d\}^\{\\mathrm\{item\}\}=\\\{\(q,y\_\{A\}\(q\)\)\\\},\\qquad y\_\{A\}\(q\)\\in\\\{0,1\\\},we fit

A^d​\(q\):=CalA\(d\)​\(A~d​\(q\),ψ​\(q\),Σd\)∈\[0,1\]\.\\hat\{A\}\_\{d\}\(q\):=\\mathrm\{Cal\}\_\{A\}^\{\(d\)\}\\\!\\big\(\\widetilde\{A\}\_\{d\}\(q\),\\psi\(q\),\\Sigma\_\{d\}\\big\)\\in\[0,1\]\.For any derived scoreZ​\(q\)Z\(q\), let

BootLQδ​\(Z^​\(q\)\):=Quantileδ⁡\(\{Z^\(b\)​\(q\)\}b=1B\),\\mathrm\{BootLQ\}\_\{\\delta\}\\\!\\big\(\\hat\{Z\}\(q\)\\big\):=\\operatorname\{Quantile\}\_\{\\delta\}\\\!\\big\(\\\{\\hat\{Z\}^\{\(b\)\}\(q\)\\\}\_\{b=1\}^\{B\}\\big\),where the bootstrap replicates refit the relevant calibrators and recompute downstream scores\. We then define

Bd\\displaystyle B\_\{d\}=\{q∈𝒬~d:BootLQδA​\(A^d​\(q\)\)≥αref\},B¯d:=Dedup⁡\(Bd\)\.\\displaystyle=\\Big\\\{q\\in\\widetilde\{\\mathcal\{Q\}\}\_\{d\}:\\mathrm\{BootLQ\}\_\{\\delta\_\{A\}\}\\\!\\big\(\\hat\{A\}\_\{d\}\(q\)\\big\)\\geq\\alpha\_\{\\mathrm\{ref\}\}\\Big\\\},\\bar\{B\}\_\{d\}=\\operatorname\{Dedup\}\(B\_\{d\}\)\.ForB¯d,h:=\{q∈B¯d:h​\(q\)=h\}\\bar\{B\}\_\{d,h\}:=\\\{q\\in\\bar\{B\}\_\{d\}:h\(q\)=h\\\}, chooseωh\>0\\omega\_\{h\}\>0with

∑h:\|B¯d,h\|\>0ωh=1,\\sum\_\{h:\\,\|\\bar\{B\}\_\{d,h\}\|\>0\}\\omega\_\{h\}=1,and set

μd​\(q\)=\{ωh​\(q\)\|B¯d,h​\(q\)\|,q∈B¯d,0,q∉B¯d\.\\mu\_\{d\}\(q\)=\\begin\{cases\}\\dfrac\{\\omega\_\{h\(q\)\}\}\{\|\\bar\{B\}\_\{d,h\(q\)\}\|\},&q\\in\\bar\{B\}\_\{d\},\\\\\[4\.30554pt\] 0,&q\\notin\\bar\{B\}\_\{d\}\.\\end\{cases\}

##### Sparse pair scoring\.

Letxqx\_\{q\}denote the text ande​\(q\):=femb​\(xq,ψ​\(q\)\)e\(q\):=f\_\{\\mathrm\{emb\}\}\(x\_\{q\},\\psi\(q\)\)\. For each non\-anchor itemqq, retrieve

𝒩k​\(q\):=ANNk⁡\(e​\(q\);B¯d∖\{q\}\),𝒢​\(q\):=𝒩k​\(q\)∪𝒜d\.\\displaystyle\\mathcal\{N\}\_\{k\}\(q\)=\\operatorname\{ANN\}\_\{k\}\\\!\\big\(e\(q\);\\bar\{B\}\_\{d\}\\setminus\\\{q\\\}\\big\),\\mathcal\{G\}\(q\)=\\mathcal\{N\}\_\{k\}\(q\)\\cup\\mathcal\{A\}\_\{d\}\.For every unordered pair\{q,u\}\\\{q,u\\\}such thatu∈𝒢​\(q\)u\\in\\mathcal\{G\}\(q\)orq∈𝒢​\(u\)q\\in\\mathcal\{G\}\(u\), we compute both directed features

ϕ​\(q,u\)\\displaystyle\\phi\(q,u\)=Φ​\(ψ​\(q\),ψ​\(u\),sime⁡\(e​\(q\),e​\(u\)\),simx⁡\(xq,xu\),𝟏​\{src⁡\(q\)=src⁡\(u\)\}\)\.\\displaystyle=\\Phi\\\!\\Big\(\\psi\(q\),\\psi\(u\),\\operatorname\{sim\}\_\{e\}\(e\(q\),e\(u\)\),\\operatorname\{sim\}\_\{x\}\(x\_\{q\},x\_\{u\}\),\\mathbf\{1\}\\\{\\operatorname\{src\}\(q\)=\\operatorname\{src\}\(u\)\\\}\\Big\)\.Using a small pair label set

ℒdpair=\{\(\(q,u\),yT​\(q,u\),yU​\(q,u\)\)\},\\mathcal\{L\}\_\{d\}^\{\\mathrm\{pair\}\}=\\\{\(\(q,u\),y\_\{T\}\(q,u\),y\_\{U\}\(q,u\)\)\\\},whereyTy\_\{T\}indicates shared domain support andyUy\_\{U\}near\-duplicate risk, we fit

T^d→​\(q,u\):=gT\(d\)​\(ϕ​\(q,u\)\)∈\[0,1\],U^d→​\(q,u\):=gU\(d\)​\(ϕ​\(q,u\)\)∈\[0,1\]\.\\displaystyle\\hat\{T\}\_\{d\}^\{\\rightarrow\}\(q,u\)=g\_\{T\}^\{\(d\)\}\\\!\\big\(\\phi\(q,u\)\\big\)\\in\[0,1\],\\hat\{U\}\_\{d\}^\{\\rightarrow\}\(q,u\)=g\_\{U\}^\{\(d\)\}\\\!\\big\(\\phi\(q,u\)\\big\)\\in\[0,1\]\.Symmetrizing both directions,

T^dsym​\(q,u\)=T^d→​\(q,u\)\+T^d→​\(u,q\)2,U^dsym​\(q,u\)=U^d→​\(q,u\)\+U^d→​\(u,q\)2,\\displaystyle\\hat\{T\}\_\{d\}^\{\\mathrm\{sym\}\}\(q,u\)=\\frac\{\\hat\{T\}\_\{d\}^\{\\rightarrow\}\(q,u\)\+\\hat\{T\}\_\{d\}^\{\\rightarrow\}\(u,q\)\}\{2\},\\hat\{U\}\_\{d\}^\{\\mathrm\{sym\}\}\(q,u\)=\\frac\{\\hat\{U\}\_\{d\}^\{\\rightarrow\}\(q,u\)\+\\hat\{U\}\_\{d\}^\{\\rightarrow\}\(u,q\)\}\{2\},we define, forq≠uq\\neq u,

S^d​\(q,u\)=\[T^dsym​\(q,u\)−λdup​U^dsym​\(q,u\)\]\+\.\\hat\{S\}\_\{d\}\(q,u\)=\\big\[\\hat\{T\}\_\{d\}^\{\\mathrm\{sym\}\}\(q,u\)\-\\lambda\_\{\\mathrm\{dup\}\}\\,\\hat\{U\}\_\{d\}^\{\\mathrm\{sym\}\}\(q,u\)\\big\]\_\{\+\}\.This is an operational proxy score rather than a probabilistic identity\.

##### Single\-item support surrogate\.

LetSd⋆​\(q,u\)∈\[0,1\]S\_\{d\}^\{\\star\}\(q,u\)\\in\[0,1\]denote an unobserved shared\-support quantity and

Cd⋆​\(q\)=𝔼Q∼μd​\[Sd⋆​\(q,Q\)\]\.C\_\{d\}^\{\\star\}\(q\)=\\mathbb\{E\}\_\{Q\\sim\\mu\_\{d\}\}\\big\[S\_\{d\}^\{\\star\}\(q,Q\)\\big\]\.We do not estimateCd⋆C\_\{d\}^\{\\star\}directly\. Forq∉𝒜dq\\notin\\mathcal\{A\}\_\{d\}, let

mq:=∑u∈𝒩k​\(q\)μd​\(u\),m\_\{q\}:=\\sum\_\{u\\in\\mathcal\{N\}\_\{k\}\(q\)\}\\mu\_\{d\}\(u\),and define

C^d​\(q\)\\displaystyle\\hat\{C\}\_\{d\}\(q\)=∑u∈𝒩k​\(q\)μd​\(u\)​S^d​\(q,u\)\+\(1−mq\)​∑a∈𝒜dπa​S^d​\(q,a\),\\displaystyle=\\sum\_\{u\\in\\mathcal\{N\}\_\{k\}\(q\)\}\\mu\_\{d\}\(u\)\\,\\hat\{S\}\_\{d\}\(q,u\)\+\(1\-m\_\{q\}\)\\sum\_\{a\\in\\mathcal\{A\}\_\{d\}\}\\pi\_\{a\}\\,\\hat\{S\}\_\{d\}\(q,a\),whereπa≥0\\pi\_\{a\}\\geq 0and∑a∈𝒜dπa=1\\sum\_\{a\\in\\mathcal\{A\}\_\{d\}\}\\pi\_\{a\}=1\. The second term is an anchor\-based tail surrogate, included for budgeted robustness rather than formal unbiasedness\.

##### Shortlisting and final selection\.

We keep a non\-anchor itemqqiff

q∈𝒦d\\displaystyle q\\in\\mathcal\{K\}\_\{d\}⇔BootLQδA​\(A^d​\(q\)\)≥αA∧BootLQδC​\(C^d​\(q\)\)≥αC∧Dd​\(q\)∈ℐd\.\\displaystyle\\iff\\mathrm\{BootLQ\}\_\{\\delta\_\{A\}\}\\\!\\big\(\\hat\{A\}\_\{d\}\(q\)\\big\)\\geq\\alpha\_\{A\}\\land\\;\\mathrm\{BootLQ\}\_\{\\delta\_\{C\}\}\\\!\\big\(\\hat\{C\}\_\{d\}\(q\)\\big\)\\geq\\alpha\_\{C\}\\land\\;D\_\{d\}\(q\)\\in\\mathcal\{I\}\_\{d\}\.For coverage, define

S^dsp​\(q,u\)=\{1,q=u∈B¯d,S^d​\(q,u\),q≠u​and​u∈𝒢​\(q\)​or​q∈𝒢​\(u\),0,otherwise\.\{\\hat\{S\}\_\{d\}^\{\\mathrm\{sp\}\}\(q,u\)=\\begin\{cases\}1,&q=u\\in\\bar\{B\}\_\{d\},\\\\ \\hat\{S\}\_\{d\}\(q,u\),&q\\neq u\\ \\text\{and\}u\\in\\mathcal\{G\}\(q\)\\ \\text\{or\}\\ q\\in\\mathcal\{G\}\(u\),\\\\ 0,&\\text\{otherwise\}\.\\end\{cases\}\}For a final set𝒮d⊆𝒦d\\mathcal\{S\}\_\{d\}\\subseteq\\mathcal\{K\}\_\{d\}of sizeKdK\_\{d\}, use

Fdsp​\(𝒮d\)=∑u∈B¯dμd​\(u\)​maxq∈𝒮d⁡S^dsp​\(q,u\)\.F\_\{d\}^\{\\mathrm\{sp\}\}\(\\mathcal\{S\}\_\{d\}\)=\\sum\_\{u\\in\\bar\{B\}\_\{d\}\}\\mu\_\{d\}\(u\)\\max\_\{q\\in\\mathcal\{S\}\_\{d\}\}\\hat\{S\}\_\{d\}^\{\\mathrm\{sp\}\}\(q,u\)\.\(2\)We optimize

max𝒮d⊆𝒦d⁡Fdsp​\(𝒮d\)s\.t\.\|𝒮d\|=Kd,\\max\_\{\\mathcal\{S\}\_\{d\}\\subseteq\\mathcal\{K\}\_\{d\}\}F\_\{d\}^\{\\mathrm\{sp\}\}\(\\mathcal\{S\}\_\{d\}\)\\quad\\text\{s\.t\.\}\\quad\|\\mathcal\{S\}\_\{d\}\|=K\_\{d\},subject to

Lh≤∑q∈𝒮d𝟏​\{h​\(q\)=h\}≤Uh,∀h∈ℋd,L\_\{h\}\\leq\\sum\_\{q\\in\\mathcal\{S\}\_\{d\}\}\\mathbf\{1\}\\\{h\(q\)=h\\\}\\leq U\_\{h\},\\qquad\\forall h\\in\\mathcal\{H\}\_\{d\},and

U^dsym​\(q,u\)≤τdup,∀q≠u∈𝒮d\.\\hat\{U\}\_\{d\}^\{\\mathrm\{sym\}\}\(q,u\)\\leq\\tau\_\{\\mathrm\{dup\}\},\\qquad\\forall q\\neq u\\in\\mathcal\{S\}\_\{d\}\.The quota condition

∑h∈ℋdLh≤Kd≤∑h∈ℋdUh\\sum\_\{h\\in\\mathcal\{H\}\_\{d\}\}L\_\{h\}\\leq K\_\{d\}\\leq\\sum\_\{h\\in\\mathcal\{H\}\_\{d\}\}U\_\{h\}is necessary but not sufficient for feasibility; feasibility also depends on𝒦d\\mathcal\{K\}\_\{d\}and the duplicate constraint\. Duplicate scores required by the last constraint are evaluated on demand on shortlisted pairs using the same cheap features and no additional LLM calls\.

###### Proposition 1\(Monotone submodularity ofFdspF\_\{d\}^\{\\mathrm\{sp\}\}\)\.

IfS^dsp​\(q,u\)≥0\\hat\{S\}\_\{d\}^\{\\mathrm\{sp\}\}\(q,u\)\\geq 0for allq,uq,u, thenFdsp​\(⋅\)F\_\{d\}^\{\\mathrm\{sp\}\}\(\\cdot\)defined in \([2](https://arxiv.org/html/2606.05104#A2.E2)\) is monotone submodular\. Greedy maximization under the cardinality constraint\|𝒮d\|=Kd\|\\mathcal\{S\}\_\{d\}\|=K\_\{d\}therefore attains a\(1−1/e\)\(1\-1/e\)approximation of the optimum ofFdspF\_\{d\}^\{\\mathrm\{sp\}\}\[[18](https://arxiv.org/html/2606.05104#bib.bib47)\]\.

###### Proof\.

For fixedu∈B¯du\\in\\bar\{B\}\_\{d\}, letfu​\(𝒮\):=maxq∈𝒮⁡S^dsp​\(q,u\)f\_\{u\}\(\\mathcal\{S\}\):=\\max\_\{q\\in\\mathcal\{S\}\}\\hat\{S\}\_\{d\}^\{\\mathrm\{sp\}\}\(q,u\)withfu​\(∅\)=0f\_\{u\}\(\\varnothing\)=0\. For anyA⊆BA\\subseteq Bandx∉Bx\\notin B,

fu​\(A∪\{x\}\)−fu​\(A\)=max⁡\{0,S^dsp​\(x,u\)−fu​\(A\)\}≥max⁡\{0,S^dsp​\(x,u\)−fu​\(B\)\}=fu​\(B∪\{x\}\)−fu​\(B\),\\displaystyle f\_\{u\}\(A\\cup\\\{x\\\}\)\-f\_\{u\}\(A\)=\\max\\big\\\{0,\\hat\{S\}\_\{d\}^\{\\mathrm\{sp\}\}\(x,u\)\-f\_\{u\}\(A\)\\big\\\}\\geq\\max\\big\\\{0,\\hat\{S\}\_\{d\}^\{\\mathrm\{sp\}\}\(x,u\)\-f\_\{u\}\(B\)\\big\\\}=f\_\{u\}\(B\\cup\\\{x\\\}\)\-f\_\{u\}\(B\),becausefu​\(A\)≤fu​\(B\)f\_\{u\}\(A\)\\leq f\_\{u\}\(B\)\. Hencefuf\_\{u\}is monotone submodular\. Since

Fdsp​\(𝒮\)=∑u∈B¯dμd​\(u\)​fu​\(𝒮\),μd​\(u\)\\displaystyle F\_\{d\}^\{\\mathrm\{sp\}\}\(\\mathcal\{S\}\)=\\sum\_\{u\\in\\bar\{B\}\_\{d\}\}\\mu\_\{d\}\(u\)\\,f\_\{u\}\(\\mathcal\{S\}\),\\mu\_\{d\}\(u\)≥0,\\displaystyle\\geq 0,FdspF\_\{d\}^\{\\mathrm\{sp\}\}is a nonnegative weighted sum of monotone submodular functions and is therefore itself monotone submodular\. The\(1−1/e\)\(1\-1/e\)greedy guarantee then follows fromNemhauseret al\.\[[18](https://arxiv.org/html/2606.05104#bib.bib47)\]\. ∎

##### Budget and implementation note\.

A typical domain uses two experts \(about 10 total hours\) for rubric alignment, item labels, pair labels, and audit, supporting roughly\|ℒditem\|≈120​–​160\|\\mathcal\{L\}\_\{d\}^\{\\mathrm\{item\}\}\|\\approx 120\\text\{\-\-\}160and\|ℒdpair\|≈60​–​80\|\\mathcal\{L\}\_\{d\}^\{\\mathrm\{pair\}\}\|\\approx 60\\text\{\-\-\}80, with25%​–​30%25\\%\\text\{\-\-\}30\\%overlap for agreement checks\. We finally spent∼\\sim$100,000 for KINA\. Each item receives one main LLM call for\(ψ​\(q\),A~d​\(q\)\)\(\\psi\(q\),\\widetilde\{A\}\_\{d\}\(q\)\), with optional re\-checks on an uncertainty set𝒰d\\mathcal\{U\}\_\{d\}, so

BLLM\(d\)=Nd\+\|𝒰d\|\\displaystyle B\_\{\\mathrm\{LLM\}\}^\{\(d\)\}=N\_\{d\}\+\|\\mathcal\{U\}\_\{d\}\|≤\(1\+ρ\)​Nd\.\\displaystyle\\leq\(1\+\\rho\)N\_\{d\}\.The same signatureψ​\(q\)\\psi\(q\)is reused for alignment, retrieval, pair scoring, and duplicate checks; no second item\-level LLM pass is required\. With fixedkk,\|𝒜d\|\|\\mathcal\{A\}\_\{d\}\|, and bootstrap countBB, sparse retrieval and sparse support scoring scale linearly inNdN\_\{d\}\. If all duplicate constraints on𝒦d\\mathcal\{K\}\_\{d\}are materialized, the additional worst\-case cost isO​\(\|𝒦d\|2\)O\(\|\\mathcal\{K\}\_\{d\}\|^\{2\}\)\. Final constrained selection is solved by greedy\-with\-repair as a heuristic, or by a small MIP when\|𝒦d\|\|\\mathcal\{K\}\_\{d\}\|is moderate; we do not claim a general approximation guarantee under the full constraint set\.

## Appendix CComparative Incentives under a Noisy Tournament

##### Scope\.

This appendix makes a comparative claim relative to flat payment\. It does not prove truthfulness, rule out collusion, or solve the full repeated stochastic game\. The whitelist extension below is reduced\-form: it compares pointwise effort incentives under a fixed continuation value, not endogenous stationary state distributions\. Throughout,X⪰FOSDYX\\succeq\_\{\\mathrm\{FOSD\}\}Ymeans first\-order stochastic dominance:

X⪰FOSDY\\displaystyle X\\succeq\_\{\\mathrm\{FOSD\}\}Y⟺FX​\(t\)≤FY​\(t\),∀t∈ℝ\.\\displaystyle\\Longleftrightarrow F\_\{X\}\(t\)\\leq F\_\{Y\}\(t\),\\qquad\\forall t\\in\\mathbb\{R\}\.Strict incentive conclusions additionally requireΔ​pmin\>0\\Delta p\_\{\\min\}\>0, which is not automatic if score noise is large or the quality bar is poorly calibrated\.

##### Setup\.

Two reviewersi∈\{a,b\}i\\in\\\{a,b\\\}evaluate the same item and choose effortei∈\{L,H\}e\_\{i\}\\in\\\{L,H\\\}\. Letψi​\(ei\)\\psi\_\{i\}\(e\_\{i\}\)be reviewerii’s reduced\-form private disutility of effort, and define the incremental cost of high effort by

κi≜ψi​\(H\)−ψi​\(L\)\.\\kappa\_\{i\}\\triangleq\\psi\_\{i\}\(H\)\-\\psi\_\{i\}\(L\)\.Fore∈\{L,H\}e\\in\\\{L,H\\\}, letQie∈ℝQ\_\{i\}^\{e\}\\in\\mathbb\{R\}denote latent review quality with CDFFe​\(q\)≜Pr⁡\(Qie≤q\)F\_\{e\}\(q\)\\triangleq\\Pr\(Q\_\{i\}^\{e\}\\leq q\)\.

###### Assumption 1\(Effort improves quality\)\.

QiH⪰FOSDQiLQ\_\{i\}^\{H\}\\succeq\_\{\\mathrm\{FOSD\}\}Q\_\{i\}^\{L\}, equivalentlyFH​\(q\)≤FL​\(q\)F\_\{H\}\(q\)\\leq F\_\{L\}\(q\)for allqq, with strict inequality on a set of positive measure\.

###### Assumption 2\(Noisy validated score\)\.

The principal observes a noisy validated score

Sie=g​\(Qie\)\+εi,S\_\{i\}^\{e\}=g\(Q\_\{i\}^\{e\}\)\+\\varepsilon\_\{i\},whereggis strictly increasing andεi\\varepsilon\_\{i\}are i\.i\.d\., continuous, and independent of all latent qualities\. LetFeS​\(s\)≜Pr⁡\(Sie≤s\)F\_\{e\}^\{S\}\(s\)\\triangleq\\Pr\(S\_\{i\}^\{e\}\\leq s\)\.

###### Assumption 3\(Independent reviewer types\)\.

Conditional on the effort profile, reviewer qualities and score noises are independent across reviewers\. Reviewer typesκi\\kappa\_\{i\}are i\.i\.d\. from a continuous CDFGG, and each reviewer observes only her own type\.

Each reviewer receives base paymentw0≥0w\_\{0\}\\geq 0, and a single bonusB\>0B\>0is awarded to the reviewer with the higher validated score provided that the winning score clears a minimum barτ\\tau:

Wi=w0\+B⋅𝟏​\{Si\>S−i,Si≥τ\}\.W\_\{i\}=w\_\{0\}\+B\\cdot\\mathbf\{1\}\\\{S\_\{i\}\>S\_\{\-i\},\\;S\_\{i\}\\geq\\tau\\\}\.\(3\)Define the winning probabilitypi​\(ei,e−i\)≜Pr⁡\(Si\>S−i,Si≥τ∣ei,e−i\)p\_\{i\}\(e\_\{i\},e\_\{\-i\}\)\\triangleq\\Pr\(S\_\{i\}\>S\_\{\-i\},\\;S\_\{i\}\\geq\\tau\\mid e\_\{i\},e\_\{\-i\}\), and the one\-shot utilities

Uitour​\(ei;e−i\)=w0\+B​pi​\(ei,e−i\)−ψi​\(ei\),Uiflat​\(ei\)=w0−ψi​\(ei\)\.\\displaystyle U\_\{i\}^\{\\mathrm\{tour\}\}\(e\_\{i\};e\_\{\-i\}\)=w\_\{0\}\+B\\,p\_\{i\}\(e\_\{i\},e\_\{\-i\}\)\-\\psi\_\{i\}\(e\_\{i\}\),U\_\{i\}^\{\\mathrm\{flat\}\}\(e\_\{i\}\)=w\_\{0\}\-\\psi\_\{i\}\(e\_\{i\}\)\.

##### One\-shot comparison\.

For any fixed opponent efforte−i∈\{L,H\}e\_\{\-i\}\\in\\\{L,H\\\},

FeS​\(s\)=Pr⁡\(Sie≤s\)=𝔼​\[Fε​\(s−g​\(Qie\)\)\]\.F\_\{e\}^\{S\}\(s\)=\\Pr\(S\_\{i\}^\{e\}\\leq s\)=\\mathbb\{E\}\\\!\\left\[F\_\{\\varepsilon\}\\\!\\big\(s\-g\(Q\_\{i\}^\{e\}\)\\big\)\\right\]\.SinceggandFεF\_\{\\varepsilon\}are increasing, the mapq↦Fε​\(s−g​\(q\)\)q\\mapsto F\_\{\\varepsilon\}\(s\-g\(q\)\)is decreasing, so Assumption[1](https://arxiv.org/html/2606.05104#Thmassumption1)together with Assumption[2](https://arxiv.org/html/2606.05104#Thmassumption2)implies

SiH⪰FOSDSiL,i\.e\.​FHS​\(s\)≤FLS​\(s\)​∀s\.\\displaystyle S\_\{i\}^\{H\}\\succeq\_\{\\mathrm\{FOSD\}\}S\_\{i\}^\{L\},\\text\{i\.e\.\}F\_\{H\}^\{S\}\(s\)\\leq F\_\{L\}^\{S\}\(s\)\\forall s\.Now defineTi≜max⁡\{S−ie−i,τ\}T\_\{i\}\\triangleq\\max\\\{S\_\{\-i\}^\{e\_\{\-i\}\},\\tau\\\}\. Because scores are continuous,\{Si\>S−i,Si≥τ\}=\{Si\>Ti\}\\\{S\_\{i\}\>S\_\{\-i\},\\;S\_\{i\}\\geq\\tau\\\}=\\\{S\_\{i\}\>T\_\{i\}\\\}almost surely\. Conditional on the effort profile,SieS\_\{i\}^\{e\}is independent ofTiT\_\{i\}\. Hence

pi​\(e,e−i\)=Pr⁡\(Sie\>Ti\)=𝔼​\[1−FeS​\(Ti\)\]\.p\_\{i\}\(e,e\_\{\-i\}\)=\\Pr\(S\_\{i\}^\{e\}\>T\_\{i\}\)=\\mathbb\{E\}\\\!\\left\[1\-F\_\{e\}^\{S\}\(T\_\{i\}\)\\right\]\.Therefore, the difference in winning probabilities is

Δ​pi​\(e−i\)≜pi​\(H,e−i\)−pi​\(L,e−i\)=𝔼​\[FLS​\(Ti\)−FHS​\(Ti\)\]≥0\.\\displaystyle\\Delta p\_\{i\}\(e\_\{\-i\}\)\\triangleq p\_\{i\}\(H,e\_\{\-i\}\)\-p\_\{i\}\(L,e\_\{\-i\}\)=\\mathbb\{E\}\\\!\\left\[F\_\{L\}^\{S\}\(T\_\{i\}\)\-F\_\{H\}^\{S\}\(T\_\{i\}\)\\right\]\\geq 0\.The one\-shot gain from high effort is then

Δ​Uitour​\(e−i\)≜Uitour​\(H;e−i\)−Uitour​\(L;e−i\)=B​Δ​pi​\(e−i\)−κi,\\displaystyle\\Delta U\_\{i\}^\{\\mathrm\{tour\}\}\(e\_\{\-i\}\)\\triangleq U\_\{i\}^\{\\mathrm\{tour\}\}\(H;e\_\{\-i\}\)\-U\_\{i\}^\{\\mathrm\{tour\}\}\(L;e\_\{\-i\}\)=B\\,\\Delta p\_\{i\}\(e\_\{\-i\}\)\-\\kappa\_\{i\},whereas under flat payment,Δ​Uiflat=Uiflat​\(H\)−Uiflat​\(L\)=−κi\\Delta U\_\{i\}^\{\\mathrm\{flat\}\}=U\_\{i\}^\{\\mathrm\{flat\}\}\(H\)\-U\_\{i\}^\{\\mathrm\{flat\}\}\(L\)=\-\\kappa\_\{i\}\. Thus the tournament adds a positive extrinsic returnB​Δ​pi​\(e−i\)B\\,\\Delta p\_\{i\}\(e\_\{\-i\}\), and high effort is weakly preferred wheneverB​Δ​pi​\(e−i\)≥κiB\\,\\Delta p\_\{i\}\(e\_\{\-i\}\)\\geq\\kappa\_\{i\}\.

##### Population implication and review\-quality shift\.

By symmetry, write

Δ​p​\(L\)≜Δ​pi​\(L\),Δ​p​\(H\)≜Δ​pi​\(H\),Δ​pmin≜min⁡\{Δ​p​\(L\),Δ​p​\(H\)\}\.\\displaystyle\\Delta p\(L\)\\triangleq\\Delta p\_\{i\}\(L\),\\Delta p\(H\)\\triangleq\\Delta p\_\{i\}\(H\),\\Delta p\_\{\\min\}\\triangleq\\min\\\{\\Delta p\(L\),\\Delta p\(H\)\\\}\.Letπm\\pi^\{m\}be the equilibrium high\-effort rate under mechanismm∈\{flat,tour\}m\\in\\\{\\mathrm\{flat\},\\mathrm\{tour\}\\\}\.

###### Theorem 1\(Tournament FOSD improvement\)\.

Under Assumptions[1](https://arxiv.org/html/2606.05104#Thmassumption1)–[3](https://arxiv.org/html/2606.05104#Thmassumption3), ifΔ​pmin\>0\\Delta p\_\{\\min\}\>0, the equilibrium high\-effort rate satisfies

πtour≥G​\(B​Δ​pmin\)≥G​\(0\)=πflat\.\\pi^\{\\mathrm\{tour\}\}\\;\\geq\\;G\(B\\,\\Delta p\_\{\\min\}\)\\;\\geq\\;G\(0\)\\;=\\;\\pi^\{\\mathrm\{flat\}\}\.Moreover, letYm=Γ​\(Qam,Qbm,Z\)Y^\{m\}=\\Gamma\(Q\_\{a\}^\{m\},Q\_\{b\}^\{m\},Z\)denote released review quality under mechanismmm, whereΓ\\Gammais coordinatewise nondecreasing andZZis mechanism\-invariant and independent of reviewer qualities\. Then

Ytour⪰FOSDYflat\.Y^\{\\mathrm\{tour\}\}\\succeq\_\{\\mathrm\{FOSD\}\}Y^\{\\mathrm\{flat\}\}\.

###### Proof\.

IfΔ​pmin\>0\\Delta p\_\{\\min\}\>0, then for all opponent actions,Δ​Uitour​\(e−i\)≥B​Δ​pmin−κi\\Delta U\_\{i\}^\{\\mathrm\{tour\}\}\(e\_\{\-i\}\)\\geq B\\,\\Delta p\_\{\\min\}\-\\kappa\_\{i\}\. Hence every type satisfyingκi≤B​Δ​pmin\\kappa\_\{i\}\\leq B\\,\\Delta p\_\{\\min\}weakly prefers high effort regardless of the opponent’s action\. Under flat payment, high effort is chosen only whenκi≤0\\kappa\_\{i\}\\leq 0\. SinceGGis continuous, indifferent types have zero mass, so

πtour≥G​\(B​Δ​pmin\)≥G​\(0\)=πflat\.\\pi^\{\\mathrm\{tour\}\}\\geq G\(B\\,\\Delta p\_\{\\min\}\)\\geq G\(0\)=\\pi^\{\\mathrm\{flat\}\}\.LetQmQ^\{m\}denote the latent quality of a randomly sampled review under mechanismmm\. Conditioning on the effort choice givesFQm​\(q\)=πm​FH​\(q\)\+\(1−πm\)​FL​\(q\)F\_\{Q\}^\{m\}\(q\)=\\pi^\{m\}F\_\{H\}\(q\)\+\(1\-\\pi^\{m\}\)F\_\{L\}\(q\), so

FQtour​\(q\)−FQflat​\(q\)=\(πtour−πflat\)​\(FH​\(q\)−FL​\(q\)\)≤0,∀q,\\displaystyle F\_\{Q\}^\{\\mathrm\{tour\}\}\(q\)\-F\_\{Q\}^\{\\mathrm\{flat\}\}\(q\)=\(\\pi^\{\\mathrm\{tour\}\}\-\\pi^\{\\mathrm\{flat\}\}\)\\big\(F\_\{H\}\(q\)\-F\_\{L\}\(q\)\\big\)\\leq 0,\\qquad\\forall q,which provesQtour⪰FOSDQflatQ^\{\\mathrm\{tour\}\}\\succeq\_\{\\mathrm\{FOSD\}\}Q^\{\\mathrm\{flat\}\}marginally\. By Assumption[3](https://arxiv.org/html/2606.05104#Thmassumption3), conditional on each mechanism,\(Qam,Qbm\)\(Q\_\{a\}^\{m\},Q\_\{b\}^\{m\}\)are i\.i\.d\. with common marginalFQmF\_\{Q\}^\{m\}\. Let\(FQm\)−1​\(u\)≜inf\{x∈ℝ:FQm​\(x\)≥u\}\(F\_\{Q\}^\{m\}\)^\{\-1\}\(u\)\\triangleq\\inf\\\{x\\in\\mathbb\{R\}:F\_\{Q\}^\{m\}\(x\)\\geq u\\\}\. SinceFQtour​\(q\)≤FQflat​\(q\)F\_\{Q\}^\{\\mathrm\{tour\}\}\(q\)\\leq F\_\{Q\}^\{\\mathrm\{flat\}\}\(q\)for allqq,

\(FQtour\)−1​\(u\)≥\(FQflat\)−1​\(u\)​∀u∈\(0,1\)\.\\displaystyle\(F\_\{Q\}^\{\\mathrm\{tour\}\}\)^\{\-1\}\(u\)\\geq\(F\_\{Q\}^\{\\mathrm\{flat\}\}\)^\{\-1\}\(u\)\\forall u\\in\(0,1\)\.IfUa,Ub​∼i\.i\.d\.​Unif​\(0,1\)U\_\{a\},U\_\{b\}\\overset\{\\mathrm\{i\.i\.d\.\}\}\{\\sim\}\\mathrm\{Unif\}\(0,1\)are independent ofZZ, andQ~jm=\(FQm\)−1​\(Uj\)\\widetilde\{Q\}\_\{j\}^\{m\}=\(F\_\{Q\}^\{m\}\)^\{\-1\}\(U\_\{j\}\)forj∈\{a,b\}j\\in\\\{a,b\\\}, then\(Q~am,Q~bm\)\(\\widetilde\{Q\}\_\{a\}^\{m\},\\widetilde\{Q\}\_\{b\}^\{m\}\)has the same law as\(Qam,Qbm\)\(Q\_\{a\}^\{m\},Q\_\{b\}^\{m\}\), whileQ~jtour≥Q~jflat\\widetilde\{Q\}\_\{j\}^\{\\mathrm\{tour\}\}\\geq\\widetilde\{Q\}\_\{j\}^\{\\mathrm\{flat\}\}almost surely forj=a,bj=a,b\. By coordinatewise monotonicity ofΓ\\Gamma,

Γ​\(Q~atour,Q~btour,Z\)≥Γ​\(Q~aflat,Q~bflat,Z\)\\displaystyle\\Gamma\(\\widetilde\{Q\}\_\{a\}^\{\\mathrm\{tour\}\},\\widetilde\{Q\}\_\{b\}^\{\\mathrm\{tour\}\},Z\)\\geq\\Gamma\(\\widetilde\{Q\}\_\{a\}^\{\\mathrm\{flat\}\},\\widetilde\{Q\}\_\{b\}^\{\\mathrm\{flat\}\},Z\)almost surely under the same coupling, which is FOSD:Ytour⪰FOSDYflatY^\{\\mathrm\{tour\}\}\\succeq\_\{\\mathrm\{FOSD\}\}Y^\{\\mathrm\{flat\}\}\. ∎

The reviewer\-independence assumption \(Assumption[3](https://arxiv.org/html/2606.05104#Thmassumption3)\) is necessary for the joint comparison: marginal FOSD alone does not suffice to compare the joint outputYmY^\{m\}\.

###### Corollary 1\(Conservative bonus calibration\)\.

AssumeΔ​pmin\>0\\Delta p\_\{\\min\}\>0, and define the generalized inverseG−1​\(u\)≜inf\{x∈ℝ:G​\(x\)≥u\}G^\{\-1\}\(u\)\\triangleq\\inf\\\{x\\in\\mathbb\{R\}:G\(x\)\\geq u\\\}\. The following sufficient conditions hold:

1. \(a\)Ifκi≡Δ​C\>0\\kappa\_\{i\}\\equiv\\Delta C\>0, thenB\>Δ​CΔ​pminB\>\\dfrac\{\\Delta C\}\{\\Delta p\_\{\\min\}\}makes high effort strictly dominant for every reviewer type\.
2. \(b\)For any target effort rateπ⋆∈\[0,1\)\\pi^\{\\star\}\\in\[0,1\),B≥\[G−1​\(π⋆\)\]\+Δ​pminB\\geq\\dfrac\{\[G^\{\-1\}\(\\pi^\{\\star\}\)\]\_\{\+\}\}\{\\Delta p\_\{\\min\}\}ensuresπtour≥π⋆\\pi^\{\\mathrm\{tour\}\}\\geq\\pi^\{\\star\}, where\[x\]\+≜max⁡\{x,0\}\[x\]\_\{\+\}\\triangleq\\max\\\{x,0\\\}\.
3. \(c\)If first moments exist andΔQ≜𝔼​\[QH\]−𝔼​\[QL\]\>0\\Delta\_\{Q\}\\triangleq\\mathbb\{E\}\[Q^\{H\}\]\-\\mathbb\{E\}\[Q^\{L\}\]\>0, then a feasible target mean\-quality gainη∈\[0,\(1−πflat\)​ΔQ\]\\eta\\in\\big\[0,\(1\-\\pi^\{\\mathrm\{flat\}\}\)\\Delta\_\{Q\}\\big\]is guaranteed by B≥\[G−1​\(πflat\+η/ΔQ\)\]\+Δ​pmin\.B\\geq\\dfrac\{\\big\[G^\{\-1\}\\\!\\big\(\\pi^\{\\mathrm\{flat\}\}\+\\eta/\\Delta\_\{Q\}\\big\)\\big\]\_\{\+\}\}\{\\Delta p\_\{\\min\}\}\.

##### Bonus expenditure accounting\.

For budget accounting, defineαu​v≜Pr⁡\(max⁡\{Sau,Sbv\}≥τ\)\\alpha\_\{uv\}\\triangleq\\Pr\\\!\\big\(\\max\\\{S\_\{a\}^\{u\},S\_\{b\}^\{v\}\\\}\\geq\\tau\\big\)for\(u,v\)∈\{L,H\}2\(u,v\)\\in\\\{L,H\\\}^\{2\}\. Under continuous scores, exactly one reviewer receives the bonus whenevermax⁡\{Sau,Sbv\}≥τ\\max\\\{S\_\{a\}^\{u\},S\_\{b\}^\{v\}\\\}\\geq\\tau\. Thus if the population high\-effort rate isπ\\pi, the expected variable bonus expenditure per task is

𝒞bonus​\(B,π\)\\displaystyle\\mathcal\{C\}\_\{\\mathrm\{bonus\}\}\(B,\\pi\)=B​\[π2​αH​H\+2​π​\(1−π\)​αH​L\+\(1−π\)2​αL​L\]≤B\.\\displaystyle=B\\Big\[\\pi^\{2\}\\alpha\_\{HH\}\+2\\pi\(1\-\\pi\)\\alpha\_\{HL\}\+\(1\-\\pi\)^\{2\}\\alpha\_\{LL\}\\Big\]\\leq B\.

##### Reduced\-form whitelist extension\.

Letziz\_\{i\}denote reviewerii’s current whitelist state,Ωi​\(zi\)≥0\\Omega\_\{i\}\(z\_\{i\}\)\\geq 0the discounted continuation value of remaining eligible for future tasks, andqi​\(ei,e−i;zi\)∈\[0,1\]q\_\{i\}\(e\_\{i\},e\_\{\-i\};z\_\{i\}\)\\in\[0,1\]the probability that the current task causes revieweriito lose whitelist status\. Holding fixed the whitelist rule, define

Uitour\+wl​\(ei;e−i,zi\)=w0\+B​pi​\(ei,e−i\)−ψi​\(ei\)\+\(1−qi​\(ei,e−i;zi\)\)​Ωi​\(zi\),\{\\begin\{multlined\}U\_\{i\}^\{\\mathrm\{tour\+wl\}\}\(e\_\{i\};e\_\{\-i\},z\_\{i\}\)=w\_\{0\}\+B\\,p\_\{i\}\(e\_\{i\},e\_\{\-i\}\)\-\\psi\_\{i\}\(e\_\{i\}\)\+\\big\(1\-q\_\{i\}\(e\_\{i\},e\_\{\-i\};z\_\{i\}\)\\big\)\\Omega\_\{i\}\(z\_\{i\}\),\\end\{multlined\}U\_\{i\}^\{\\mathrm\{tour\+wl\}\}\(e\_\{i\};e\_\{\-i\},z\_\{i\}\)=w\_\{0\}\+B\\,p\_\{i\}\(e\_\{i\},e\_\{\-i\}\)\-\\psi\_\{i\}\(e\_\{i\}\)\+\\big\(1\-q\_\{i\}\(e\_\{i\},e\_\{\-i\};z\_\{i\}\)\\big\)\\Omega\_\{i\}\(z\_\{i\}\),\}Uiflat\+wl​\(ei;e−i,zi\)=w0−ψi​\(ei\)\+\(1−qi​\(ei,e−i;zi\)\)​Ωi​\(zi\)\.\{\\begin\{multlined\}U\_\{i\}^\{\\mathrm\{flat\+wl\}\}\(e\_\{i\};e\_\{\-i\},z\_\{i\}\)=w\_\{0\}\-\\psi\_\{i\}\(e\_\{i\}\)\+\\big\(1\-q\_\{i\}\(e\_\{i\},e\_\{\-i\};z\_\{i\}\)\\big\)\\Omega\_\{i\}\(z\_\{i\}\)\.\\end\{multlined\}U\_\{i\}^\{\\mathrm\{flat\+wl\}\}\(e\_\{i\};e\_\{\-i\},z\_\{i\}\)=w\_\{0\}\-\\psi\_\{i\}\(e\_\{i\}\)\+\\big\(1\-q\_\{i\}\(e\_\{i\},e\_\{\-i\};z\_\{i\}\)\\big\)\\Omega\_\{i\}\(z\_\{i\}\)\.\}LetΔ​qi​\(e−i,zi\)≜qi​\(L,e−i;zi\)−qi​\(H,e−i;zi\)\\Delta q\_\{i\}\(e\_\{\-i\},z\_\{i\}\)\\triangleq q\_\{i\}\(L,e\_\{\-i\};z\_\{i\}\)\-q\_\{i\}\(H,e\_\{\-i\};z\_\{i\}\), and assumeΔ​qi​\(e−i,zi\)≥0\\Delta q\_\{i\}\(e\_\{\-i\},z\_\{i\}\)\\geq 0for alle−i,zie\_\{\-i\},z\_\{i\}\. Then

Δ​Uitour\+wl​\(e−i,zi\)\\displaystyle\\Delta U\_\{i\}^\{\\mathrm\{tour\+wl\}\}\(e\_\{\-i\},z\_\{i\}\)=Δ​Uiflat\+wl​\(e−i,zi\)\+B​Δ​pi​\(e−i\),Δ​Uiflat\+wl​\(e−i,zi\)=−κi\+Δ​qi​\(e−i,zi\)​Ωi​\(zi\)\.\\displaystyle=\\Delta U\_\{i\}^\{\\mathrm\{flat\+wl\}\}\(e\_\{\-i\},z\_\{i\}\)\+B\\,\\Delta p\_\{i\}\(e\_\{\-i\}\),\\Delta U\_\{i\}^\{\\mathrm\{flat\+wl\}\}\(e\_\{\-i\},z\_\{i\}\)=\-\\kappa\_\{i\}\+\\Delta q\_\{i\}\(e\_\{\-i\},z\_\{i\}\)\\Omega\_\{i\}\(z\_\{i\}\)\.Equivalently,

Δ​Uitour\+wl​\(e−i,zi\)=B​Δ​pi​\(e−i\)−κi\+Δ​qi​\(e−i,zi\)​Ωi​\(zi\)≥Δ​Uitour​\(e−i\)\.\\displaystyle\\Delta U\_\{i\}^\{\\mathrm\{tour\+wl\}\}\(e\_\{\-i\},z\_\{i\}\)=B\\,\\Delta p\_\{i\}\(e\_\{\-i\}\)\-\\kappa\_\{i\}\+\\Delta q\_\{i\}\(e\_\{\-i\},z\_\{i\}\)\\Omega\_\{i\}\(z\_\{i\}\)\\geq\\Delta U\_\{i\}^\{\\mathrm\{tour\}\}\(e\_\{\-i\}\)\.DefineΛi≜infe−i,ziΔ​qi​\(e−i,zi\)​Ωi​\(zi\)\\Lambda\_\{i\}\\triangleq\\inf\_\{e\_\{\-i\},z\_\{i\}\}\\Delta q\_\{i\}\(e\_\{\-i\},z\_\{i\}\)\\Omega\_\{i\}\(z\_\{i\}\), andκ~i≜κi−Λi\\widetilde\{\\kappa\}\_\{i\}\\triangleq\\kappa\_\{i\}\-\\Lambda\_\{i\}\. Then

Δ​Uitour\+wl​\(e−i,zi\)≥B​Δ​pmin−κ~i,∀e−i,zi\.\\displaystyle\\Delta U\_\{i\}^\{\\mathrm\{tour\+wl\}\}\(e\_\{\-i\},z\_\{i\}\)\\geq B\\,\\Delta p\_\{\\min\}\-\\widetilde\{\\kappa\}\_\{i\},\\qquad\\forall e\_\{\-i\},z\_\{i\}\.Thus all one\-shot sufficient conditions in Corollary[1](https://arxiv.org/html/2606.05104#Thmcorollary1)apply*pointwise*after replacingκi\\kappa\_\{i\}byκ~i\\widetilde\{\\kappa\}\_\{i\}\. Ifκi≡Δ​C\\kappa\_\{i\}\\equiv\\Delta CandΛi≥Λmin\\Lambda\_\{i\}\\geq\\Lambda\_\{\\min\}for allii, thenB\>\(Δ​C−Λmin\)\+Δ​pminB\>\\dfrac\{\(\\Delta C\-\\Lambda\_\{\\min\}\)\_\{\+\}\}\{\\Delta p\_\{\\min\}\}is a conservative sufficient condition for high effort to be strictly optimal at every state\.

##### Takeaway\.

Relative to flat payment, the noisy tournament changes private incentives throughB​Δ​pi​\(e−i\)B\\,\\Delta p\_\{i\}\(e\_\{\-i\}\), the bonus\-weighted increase in the probability of clearing the quality bar and outperforming the matched reviewer\. When this return is large enough relative to reviewer\-specific effort cost, the set of high\-effort types expands, inducing a first\-order stochastic right\-shift in per\-review latent quality\. Under reviewer independence and monotone adjudication, the same right\-shift carries over to released\-data quality\. Whitelist access further strengthens this comparison pointwise by lowering the effective net cost of effort through continuation value\.

## Appendix DTaxonomy of Disciplines

![[Uncaptioned image]](https://arxiv.org/html/2606.05104v2/x3.png)

Figure 4:Taxonomy of Disciplines
Table 8:Disciplines and Subfields StatisticsDisciplineFieldSubfieldCountAgronomyAnimal HusbandryAnimal Nutrition and Feed Science8Animal Rearing and Breeding7AquacultureAquaculture3Crop ScienceCrop Science7ForestryForest Cultivation and Genetic Breeding3Landscape Plants and Ornamental Horticulture4Veterinary MedicineVeterinary Medicine8EconomicsApplied EconomicsEconomic Statistics3Finance5Industrial Economics3International Trade3Labor Economics5Public Finance3Quantitative Economics3Theoretical EconomicsEconomic History1Political Economy3Western Economics5EducationEducationEducational Technology and Principles4Preschool Education6Special Education3Theory of Curriculum and Instruction2Physical EducationPhysical Education and Training6Sports Science and Medicine3PsychologyPsychology7EngineeringAeronautical and Astronautical Science and TechnologyAeronautical and Astronautical Science and Technology3Agricultural EngineeringAgricultural Environment and Soil\-Water Engineering6Agricultural Mechanization Engineering3ArchitectureArchitectural Design and Theory1Urban Planning and Design1Architectural History4Chemical Engineering and TechnologyChemical Transport Engineering4Elements of Chemical Reaction Engineering8Fluid Flow and Heat Transfer in Chemical Engineering4Mass Transport and Separation Process in Chemical Engineering3Civil EngineeringBridge and Tunnel Engineering2Geotechnical Engineering3Structural Engineering4Urban Infrastructure Engineering3Computer Science and TechnologyAdvanced Programming Languages5Computer Architecture3Computer Networks2Computer Software and Theory3Data Structures2Databases4Formal Languages4Operating Systems3Pattern Recognition3Principles of Computer Organization2Control Science and EngineeringControl Theory and Control Engineering10Guidance, Navigation and Control2Operations Research and Cybernetics2Electrical EngineeringElectrical Theory and New Technologies3High Voltage and Insulation Technology2Power Electronics and Electrical Drives8Power Systems and Automation3Electronic Science and TechnologyCircuits and Systems4Electromagnetic Field and Microwave Technology3Microelectronics and Solid\-State Electronics3Environmental Science and EngineeringEnvironmental Engineering4Environmental Science2Environmental and Resource Protection4Food Science and EngineeringFood Biochemistry3Food Processing and Storage Engineering3Forestry EngineeringForest Engineering4Wood Science and Technology4Geological Resources and Geological EngineeringGeological Resources and Geological Engineering5Hydraulic EngineeringHydraulics and Hydrology5Water conservancy and Hydropower Engineering2Information and Communication EngineeringAntenna and Radio Communication5Communication Principles3Communication and Information Systems3Optical Fiber Communication3Signal and Information Processing3Instrument Science and TechnologyInstrument Science and Technology8Materials Science and EngineeringMaterials Physics and Chemistry3Materials Processing Engineering5Mechanical EngineeringManufacturing Automation7Mechatronic Engineering3MechanicsRigid Body Mechanics3Theoretical Fluid Mechanics5Theoretical Mechanics3Metallurgical EngineeringIron and Steel Metallurgy3Non\-ferrous Metallurgy3Physical Chemistry of Metallurgical Process3Principles of Metallurgy3Mining EngineeringMineral Processing Engineering3Naval Architecture and Ocean EngineeringMarine Engineering3Ship Mechanics and Design Principles3Nuclear Science and TechnologyNuclear Energy and Reactor Technology2Radiation Protection and Nuclear Technology Applications3Optical EngineeringApplied Optics3Laser Technology4Optoelectronic Technology3Theoretical Optics5Petroleum and Natural Gas EngineeringOil and Gas Field Development and Storage & Transportation Engineering4Poromechanics and Reservoir Physics2Power Engineering and Engineering ThermophysicsEngineering Fluid Mechanics2Engineering Thermophysics3Fluid Machinery and Engineering2Heat Transfer5Internal Combustion Engineering3Power Machinery and Engineering2Refrigeration and Cryogenic Engineering3Thermal Energy Engineering3Surveying and Mapping Science and TechnologyCartography and Geographic Information Engineering2Geodesy and Surveying Engineering3Textile Science and EngineeringTextile Chemistry and Dyeing Engineering3Textile Materials Science3Transportation EngineeringRoad and Railway Engineering3Traffic Information Engineering and Control3Transportation Planning and Management3Vehicle Operation Engineering4Weapon Science and TechnologyMilitary Chemistry and Pyrotechnics1Weapon Systems Science and Engineering3HistoryHistoryArchaeology and Museology9Historical Geography2World History2LawLawCivil and Commercial Law5Constitutional and Administrative Law6Contract Law13Criminal Law5International Law14Law and Social Governance5Legal Theory and Legal History6Procedural Law7Political SciencePolitical Science4Literature and ArtsArt StudiesBroadcasting and Television Art3Dance Studies3Design Arts4Drama and Opera Studies6Film Studies3Fine Arts3Journalism and CommunicationCommunication and Broadcasting3History and Theory of Journalism and Media Management5Journalism and News Practice3Language and LiteratureClassical Asian Literature3French Language and Literature4Linguistics and Applied Linguistics2Literary History3Literary Theory3Modern and Contemporary Literature1Philology and Bibliography3Russian Language and Literature3MusicologyComposition3Harmony1Instrumentation and Performance1Music History, Education, and Technology2Musical Forms and Analysis3ManagementBusiness AdministrationBusiness and Accounting Management6Library, Information and Archival ManagementInformation Management Science5Library and Archival Science9Management Science and EngineeringManagement Science and Engineering1Public AdministrationEducation Economics, Management and Social Security3Land Resource Management and Administrative Management4Social Medicine and Health Management1MedicineBasic MedicineForensic Medicine3Human Anatomy and Histology\-Embryology3Immunology3Pathogen Biology1Pathology and Pathophysiology2Clinical MedicineAnesthesiology2Clinical Laboratory Diagnostics8Dermatology and Venereology1Emergency Medicine5Geriatric Medicine3Imaging and Nuclear Medicine3Internal Medicine4Neurology1Nursing and Rehabilitation Medicine7Obstetrics and Gynecology2Oncology1Ophthalmology1Otorhinolaryngology3Pediatrics1Psychiatry and Mental Health2Surgery1PharmacyMedicinal Chemistry3Microbiology and Biochemical Pharmacy3Pharmaceutical Analysis3Pharmaceutics2Pharmacology1Public Health and Preventive MedicineEpidemiology and Health Statistics1Health Toxicology and Environmental Health1StomatologyBasic Stomatology3Clinical Stomatology2Traditional MedicineTraditional Health Preservation3Traditional Medicine Theory2PhilosophyPhilosophyEthics2Logic1Philosophical Aesthetics1Philosophy of Science and Technology7Religious Studies2ScienceAstronomyAstronomical Observation and Technology2Astrophysics4Cosmology1Solar System Science5Stellar and Interstellar Evolution4Atmospheric ScienceAtmospheric Physics and Atmospheric Environment2Meteorology2BiologyBiochemistry and Molecular Biology3Biophysics2Botany3Cell Biology2Ecology3Genetics3Microbiology3Physiology3Zoology3ChemistryAnalytical Chemistry3Electrochemistry3Inorganic Chemistry2Organic Chemistry2Physical Chemistry3Polymer Chemistry and Physics4Radiochemistry6GeographyHuman Geography3Physical Geography5GeologyGeochemistry8Mineralogy, Petrology, and Economic Geology3Paleontology and Stratigraphy2Principles of Seismic Exploration3Structural Geology3GeophysicsSolid Earth Geophysics2MathematicsAdvanced Algebra4Combinatorial Mathematics2Computational Mathematics3Cryptography8Discrete Mathematics2Functions of Complex Variables4Functions of Real Variables1Fundamental Mathematics3Fuzzy Mathematics8Geometry and Topology3Graph Theory2Group Theory3Mathematical Analysis3Number Theory3Numerical Analysis3Ordinary Differential Equations3Polynomials and Series Expansions1Probability and Statistics2Special Number Theory3Stochastic Processes3OceanographyHydrogeology4Marine Biology2Marine Chemistry3Underwater Acoustics2Physical OceanographyPhysical Oceanography3PhysicsAcoustics2Atomic and Molecular Physics3Electrodynamics3Fluid Physics3Particle and Nuclear Physics1Polymer Physics3Quantum Mechanics3Semiconductor Physics3Subatomic and Atomic Physics1Thermodynamics2Thermodynamics and Statistical Physics2SociologySociologyDemography and Anthropology10Social and Folklore Studies9
## Appendix EData Collection Details

### E\.1Annotator Recruitment Test

1. 1\.Academic Background - •University \(Full Name in English\): - •Highest Academic Degree \(Bachelor’s / Master’s / Doctoral / Currently Enrolled\): - •Major Name \(Official Title\): - •Intended Disciplinary Fields for Question Design \(Limit: 1–3\):
2. 2\.Authoritative Domain Knowledge Please list3of the most authoritative publications, conferences or monographs \(titles included\) in your professional field\. We intend to evaluate your professional insight through your understanding of the authoritative sources in the industry\. Format requirement: Name \+ Type \(Journal/Conference/Monograph\) \+ Rationale
3. 3\.Understanding of Disciplinary Competency Assessment If you were asked to design3–6 multiple\-choice questionsto assess the advanced literacy and cognitive abilities of LLMs in your chosen field, what content would each question target? Please briefly elaborate on the question content and explain the rationale for designing each question\. Please ensure that the overall design logic of this set of questions is clearly demonstrated: specifically, how these questions, covering different dimensions within the discipline, can be integrated to outline a comprehensive cognitive framework of the subject\. Format requirement: Content \+ Rationale
4. 4\.Stress Tolerance and Work Attitude 1. \(a\)Attitude Towards Review and Revision \(Single Choice\) If your questions are returned due to incomplete citations or non\-standard formatting, and you are required to supplement supporting evidence item by item, what would your response be? - •A\. The rules are reasonable and necessary; I will supplement the required information accordingly - •B\. I feel a bit annoyed, but I will revise the questions as required - •C\. I think the requirements are overly detailed and somewhat confusing - •D\. I am not willing to continue with this process 2. \(b\)Tendency of Responsibility Attribution \(Single Choice\) If your questions fail to pass the review multiple times, what do you think is the primary cause? - •A\. My understanding of the standards and requirements is not thorough enough - •B\. The questions themselves still have room for improvement - •C\. The review standards are overly subjective - •D\. Bad luck
5. 5\.Rule Comprehension Verification 1. \(a\)Understanding of Question Types \(Multiple Choice\) Which of the following questions donotmeet the requirements for high\-quality disciplinary assessment questions? - •A\. Questions testing the reasoning and application of core disciplinary theories in specific scenarios - •B\. Questions requiring the model to memorize specific numerical values or factual information - •C\. Combined logical reasoning questions based on multiple statements - •D\. Static judgment questions involving only extremely niche or rare subjects - •E\. Questions developed based on individual scholars’ viewpoints that lack widespread consensus 2. \(b\)Understanding of Question Design \(Single Choice\) What is the main source of difficulty in high\-quality assessment questions? - •A\. Numerous calculation steps and complex numerical values - •B\. Obscure knowledge points and high memory difficulty - •C\. Interactions between multiple restrictive conditions - •D\. Long question length and large information volume
6. 6\.Workflow and Standard Confirmation \(Multiple Choice\) - •Develop all questions in English \(including punctuation marks\) throughout the entire process - •Each option \(whether correct or incorrect\) must be supported by evidence or citations - •Questions may undergo multiple rounds of review and revision - •Questions must demonstrate disciplinary representativeness, rather than being esoteric, biased, or tricky

### E\.2Annotation Manual

#### E\.2\.1Project Background

Current LLMs have demonstrated strong capabilities in programming, mathematics, and reasoning, but their performance still varies widely across different disciplines\. To enable researchers and practitioners in various disciplinary fields to select the most appropriate large AI model for their work, domain experts are needed to design questions that evaluate models’ disciplinary knowledge and competence in their respective fields\. The questions you provide will serve as critical data for evaluating LLMs\.

- •Your role:Question Design Expert
- •Your task: Create high\-difficulty multiple\-choice questions that at least three LLMs will answer incorrectly, and provide explanations and source materials\. - –Question: Consists of a stem and options\. Questions must berepresentative and uniqueto the discipline, and able to assess high\-level knowledge and literacy in the field\. Questions may be original, adapted, or directly excerpted;Full AI\-generated questions are strictly prohibited\. - –Explanation: Clearly explain why each option is correct or incorrect\. - –Source Materials: The origin of the question, which may be a book or high\-quality academic paper\.
- •Discipline List: Please check whether the list includes your specialized discipline\. Ensuring the accuracy and rationality of the questions is our top priority\.

#### E\.2\.2Core Criteria for Question Design

What We Require:

1. 1\.Disciplinary Representativeness:Instances must act as highly representative probes\. If a discipline were to be evaluated using only 3 to 5 questions, the submitted instance must be fundamental and comprehensive enough to be one of them\.
2. 2\.High\-Order Knowledge Application: - •Questions must construct a logically complete closed system where the solution is the inevitable product of general deductive reasoning applied to domain\-specific axioms\. - •Difficulty should stem from the complex coupling of constraints and multi\-hop reasoning, not merely computational heavy\-lifting\. - •For instance, physics questions should demand the integration of axioms and skills to solve a specific state transition, rather than asking for the formula of momentum conservation\.
3. 3\.Static Knowledge Breadth:For memory\-intensive disciplines \(e\.g\., Education\), questions should maximize knowledge coverage \(e\.g\., utilizing10\+10\+statements to encompass major pedagogical theories\)\.

What We Reject:

1. 1\.Narrow or Trivial Memorization:Questions testing pure rote memory devoid of core disciplinary literacy \(e\.g\., "What is the 100th digit ofπ\\pi?"\) or focusing on hyper\-niche, obscure sub\-entities\.
2. 2\.Idiosyncratic or Tricky Trivia:Questions universally recognized as flawed or unreasonable even in human examinations\.
3. 3\.Weak Epistemological Consensus:Hypotheses proposed by individual scholars that lack widespread academic consensus or are highly volatile \(e\.g\., highly debated legal interpretations\)\.
4. 4\.Non\-Disciplinary Failure Modes:Instances where LLMs fail due to semantic traps, ambiguous phrasing, or floating\-point calculation errors rather than a deficit in disciplinary literacy\.

#### E\.2\.3Standardized Annotation Workflow

The question authoring process is strictly compartmentalized into six components\. The specifications for each component are detailed in Table[10](https://arxiv.org/html/2606.05104#A5.T10)\.

Table 10:Standardized Annotation Workflow and Specifications\.
#### E\.2\.4Accepted and Rejected Cases

Table[11](https://arxiv.org/html/2606.05104#A5.T11)illustrates the dichotomy between accepted and rejected submissions based on our core criteria\.

The accepted case provided here is for reference only in terms of content and format\. Please note that question types are not limited to calculation problems\.

This question is selected as an excellent example for the following reasons:

1. 1\.It targets marine chemistry and assesses advanced reasoning and computational abilities, rather than only testing static knowledge\.
2. 2\.Even though it is a calculation question, the analysis of incorrect options is briefly summarized instead of being copied and pasted repetitively\.

Table 11:Accepted and Rejected CasesTable 12:Accepted and Rejected Cases \(continued\)

### E\.3Review Manual

#### E\.3\.1Core Philosophy and General Workflow

The primary objective of the review process is to ensure that each curated instance serves as a highly representative probe for evaluating LLMs on specific fine\-grained disciplines\. Reviewers must adopt azero\-tolerance policytoward ambiguous, peripheral, or substandard data\.

- •Comprehensive Feedback:Unless an instance requires a complete topic overhaul, reviewers must identify and aggregate all logical, factual, and formatting issues into a single comprehensive feedback report before returning it for revision\.
- •AI\-Generation Rejection:Submissions exhibiting clear patterns of unedited, large\-scale AI generation must be categorically rejected\.
- •Escalation Mechanism:For borderline cases or interdisciplinary disputes, reviewers are required to escalate the instance to the Core Committee for final arbitration\.

#### E\.3\.2Disciplinary Alignment and Factuality

- •Disciplinary Matching:Reviewers must first verify whether the question strictly aligns with the declared subfield\. Mismatched instances must be returned with a mandate to reclassify or rewrite\.
- •Epistemological Rigor \(Especially in Humanities & Social Sciences\):Questions must test established scientific truths, foundational theories, or widely recognized academic consensus\. Volatile opinions, transient scholarly debates, or highly subjective hypotheses without proper contextualization must be rejected\.
- •Memory Error Rejection:If an instance causes a model to fail solely due to a superficial memory error \(e\.g\., misremembering an obscure date, a peripheral name, or a raw statistic\) rather than a deficit in high\-order deductive reasoning or conceptual understanding, the instance must be rejected\. Questions must evaluate the mastery of causal mechanisms, not data retrieval\.

#### E\.3\.3Source Verification and Material Authenticity

Every claim must be fully traceable to authoritative human knowledge sources\.

- •Authoritative Sources:Referenced literature must be identifiable via valid URLs, DOIs, or standard APA citations\. Acceptable sources include high quality peer\-reviewed journals \(e\.g\., SCI/SSCI Q1, CSSCI\), authoritative monographs, or official databases\. Preprints \(e\.g\., arXiv\) are only acceptable if they demonstrate high citation counts or widespread community validation\.
- •Material Consistency:Any accompanying materials \(e\.g\., figures, charts\) must be intrinsically necessary to solve the question and perfectly correspond to the cited source\.

#### E\.3\.4Option Rigor and Structural Integrity

To prevent LLMs from exploiting logical loopholes, the 10\-option structure must adhere to strict constraints:

- •Pseudo\-Multi\-Choice Constraints:For combination\-based questions, the question stem must present a minimum of6 distinct foundational statements\.
- •No Proper Subset Relationships:To avoid logical leakage, options must not have proper subset relationships\. For example, if Option A is\{1,2\}\\\{1,2\\\}and Option B is\{1,2,3\}\\\{1,2,3\\\}, proving B correct implicitly validates A, introducing ambiguity\. Options must be mutually exclusive in their truth values\.
- •Explanation Completeness:The explanations provided for the options must comprehensively address why the correct option is uniquely valid and why all distractors are flawed\. For deterministic questions \(e\.g\., mathematical calculations\), overlapping explanations across options are permissible provided the core proof is sound\.

### E\.4Formatting and Rendering

Reviewers must ensure that all text, mathematical formulas, and symbolic logic within the Question, Options, and Explanations strictly conform to standard LaTeX syntax and render flawlessly in Markdown environments\. Formatting failures constitute immediate grounds for rejection\.

### E\.5Appeal Handling and Quality Enforcement

- •Annotator Appeals:Recognizing the domain expertise of annotators, the pipeline supports a formal appeal process\. Reviewers are required to objectively re\-evaluate contested instances based on newly provided academic evidence\.
- •Accountability:Reviewers who consistently exhibit superficial auditing \("lazy consensus"\), approve logically flawed instances, or violate the double\-blind protocols will be permanently removed from the reviewer pool and forfeit their compensation\.

### E\.6LLM\-based Filtering framework

`Feature Extraction Prompt Failure Pattern Analysis Prompt Consensus Voting Prompt`

`E\.7 Agentic Workflow Verification Diagnosis Agent Prompt Refinement Agent Prompt Appendix F Experiment Details F\.1 List of Models Evaluated Table 13: Overview of Evaluated Models on KINA Benchmark\. Provider / Family Specific Models OpenAI GPT\-5\.4 \[21\], GPT\-5\.2 \[20\] Google Gemini\-3\.1\-Pro\-Preview , Gemini\-3\-Flash\-Preview \[6\] Anthropic Claude\-Opus\-4\.6 \[1\], Claude\-Sonnet\-4\.6 \[2\] ByteDance Doubao\-Seed\-2\.0\-Pro\-260215, Doubao\-Seed\-2\.0\-Lite\-260215 \[24\] Tongyi \(Qwen\) Qwen3\.5\-397B\-A17B, Qwen3\.5\-122B\-A10B\-FP8, Qwen3\.5\-35B\-A3B\-FP8, Qwen3\.5\-27B \[28\] Qwen3\-Max\-250923 \[27\], Qwen3\-Next\-80B\-A3B\-Instruct, Qwen3\-Next\-80B\-A3B\-Thinking\-FP8 \[33\] Qwen3\-235B\-A22B, Qwen3\-235B\-A22B\-Thinking\-2507, Qwen3\-30B\-A3B, Qwen3\-30B\-A3B\-Thinking\-2507, Qwen3\-32B, Qwen3\-14B, Qwen3\-8B, Qwen3\-4B, Qwen3\-4B\-Thinking\-2507, Qwen3\-1\.7B, Qwen3\-0\.6B \[33\] Qwen2\.5\-72B\-Instruct \[26\], Qwen2\-72B\-Instruct \[34\] Meta \(Llama\) Llama\-4\-Maverick\-17B\-Instruct\-FP8 \[3\], Llama\-3\.1\-405B\-Instruct, Llama\-3\-70B\-Instruct \[10\] DeepSeek DeepSeek\-v3\.2\-Thinking \[15\] Moonshot \(Kimi\) Kimi\-k2\.5 \[25\] xAI Grok\-4\.1\-Fast\-Reasoning \[32\] Zhipu AI GLM\-5 \[35\] MiniMax Minimax\-M2\.5 \[17\] StepFun Step\-3\.5\-Flash \[12\] Mistral Mixtral\-8x7B\-Instruct\-v0\.1 \[13\] F\.2 Evaluation Prompt Evaluation Prompt F\.3 Result Details Table 14: Evaluation Stability of KINA\. Model Avg@4 Std Model Avg@4 Std Gemini\-3\.1\-Pro\-Preview 53\.17 0\.72 Claude\-Opus\-4\.6 49\.92 0\.32 GPT\-5\.4 48\.55 0\.56 Doubao\-Seed\-2\.0\-Pro\-260215 44\.99 0\.80 Gemini\-3\-Flash\-Preview 43\.91 0\.71 Qwen3\.5\-397B\-A17B 42\.99 0\.78 Doubao\-Seed\-2\.0\-Lite\-260215 41\.49 0\.61 Kimi\-K2\.5 40\.24 0\.66 GPT\-5\.2 39\.52 0\.99 Qwen3\.5\-27B 39\.35 0\.55 Qwen3\.5\-122B\-A10B 38\.88 0\.90 DeepSeek\-V3\.2\-Thinking 38\.01 0\.57 Qwen3\-Max\-2025\-09\-23 35\.90 0\.44 GLM\-5 35\.85 0\.55 Qwen3\.5\-35B\-A3B 35\.43 0\.69 Grok\-4\.1\-Fast\-Reasoning 33\.73 0\.64 Qwen3\-235B\-A22B\-Thinking\-2507 32\.15 0\.52 Llama\-4\-Maverick\-17B\-Instruct 31\.62 0\.59 Minimax\-M2\.5 30\.28 0\.46 Qwen3\.5\-9B 30\.09 0\.91 Qwen3\-Next\-80B\-A3B\-Instruct 30\.09 0\.62 Claude\-Sonnet\-4\.6 30\.01 0\.56 Llama\-3\.1\-405B\-Instruct 29\.59 0\.62 Qwen3\-235B\-A22B 29\.37 0\.83 Step\-3\.5\-Flash 29\.12 0\.52 Qwen3\.5\-4B 28\.50 0\.72 Qwen3\-Next\-80B\-A3B\-Thinking 28\.28 1\.33 Qwen3\-32B 27\.41 0\.38 Qwen3\-30B\-A3B\-Thinking\-2507 27\.03 0\.42 Meta\-Llama\-3\-70B\-Instruct 26\.61 0\.39 Qwen3\-14B 26\.00 0\.64 Qwen2\-72B\-Instruct 24\.92 1\.24 Qwen3\-4B\-Thinking\-2507 24\.83 0\.37 Qwen3\-30B\-A3B 24\.28 0\.45 Qwen2\.5\-72B\-Instruct 23\.03 1\.13 Qwen3\-8B 22\.27 1\.62 Qwen3\-4B 21\.50 0\.81 Qwen3\.5\-2B 20\.52 0\.46 Qwen3\-1\.7B 18\.33 1\.16 Mixtral\-8x7B\-Instruct\-v0\.1 17\.83 0\.55 Qwen3\.5\-0\.8B 16\.66 0\.80 Qwen3\-0\.6B 14\.49 0\.78 Figure 5: Subject\-Level Score Distribution Across Top\-10 Models Note\. Top\-10 models are ranked by overall discipline score\. Each ridge shows the smoothed density of one model’s subject\-level scores \(0–100\)\. Right\-shifted and narrower ridgelines indicate higher and more stable performance across subjects\. \(a\) Top\-10 Models: Discipline\-Level Weighted Score Distribution \(b\) Top\-10 Models: Field\-Level Weighted Score Distribution \(a\) Top\-10 Models: Subfield\-Level Weighted Score Distribution Figure 7: Top\-10 Model Performance Distributions Across Three Aggregation Levels \(continued\) Note\. Each violin shows the distribution of weighted model scores, with weights proportional to sample size\. Panels differ only in aggregation level: discipline, field, and subfield\. Wider sections indicate higher score density, and the central marker indicates the median\. Figure 8: Inference Cost Distribution of Qwen3 Dense Models\. Figure 9: Inference Cost Distribution of Qwen3 Moe Models\. Figure 10: Inference Cost Distribution of Qwen3\.5 Dense Models\. Figure 11: Inference Cost Distribution of Qwen3\.5 Moe Models\. Figure 12: UpSet plot visualizing intersections of high\-performing data points \(average score ≥\\geq 35\) among the KINA Top\-10 models\. The matrix at the bottom denotes set inclusion for distinct intersection combinations, paired with horizontal bars on the left showing total set sizes per discipline\. The vertical bar chart at the top specifies the sizes of these intersections\. Appendix G Datasheet for KINA We follow the structure of Gebru et al\. \[8\]\. Motivation For what purpose was the dataset created? KINA was created to evaluate frontier large language models on disciplinary\-representative knowledge tasks\. Existing benchmarks either prioritize scale \(SuperGPQA\), extreme difficulty \(HLE\), or narrow expert depth \(GPQA\), but none combine wide disciplinary coverage with an explicit representativeness criterion and an incentive\-aligned annotation pipeline\. Who created the dataset and on behalf of which entity? The dataset was created by the KINA team\. Who funded the creation? 2077AI & M\-A\-P & UTokyo & CMU Composition What do the instances represent? Each instance is a pseudo\-multiple\-choice question with 1010 combinatorial options\. An instance includes the stem, 1010 options, an option\-level explanation with cited sources, and the originating reference material\. How many instances are there in total? 899899 items, distributed across 1212 top\-level disciplines, 7070 fields, and 261261 fine\-grained subfields\. Does the dataset contain all possible instances or is it a sample? A sample\. The selection process is described in §3\.1 and §4\. Does the dataset identify any subpopulations? The instances are organized by discipline, field, and subfield following the CIP\-aligned taxonomy \(Appendix D\)\. Does the dataset contain data that might be considered confidential? No\. All source materials are drawn from publicly published academic literature\. Does the dataset contain data that might be offensive, insulting, threatening, or anxiety\-inducing? No\. Collection process How was the data collected? Three\-stage pipeline: rule\-based screening, double\-blind expert review under a bonus\-on\-bar tournament mechanism, and three\-judge LLM consensus \(§4\)\. Who was involved in the data collection process and how were they compensated? Annotators were graduate students from top\-tier global universities and senior industry experts, recruited via a two\-round examination \(Appendix E\.1\)\. Compensation followed the bonus\-on\-bar tournament described in §3\.2, calibrated per discipline\. Over what timeframe was the data collected? 2025\.10 to 2025\.12\. Was an ethical review process conducted? Yes\. KINA was reviewed in accordance with the applicable ethical requirements\. The research does not involve interventions on human participants or the collection of sensitive personal information\. All data used in the study were obtained and processed in a manner consistent with relevant privacy, consent, and data protection requirements\. Preprocessing, cleaning, labeling Was any preprocessing/cleaning of the data done? Yes\. Stage 1 enforces uniqueness \(cosine similarity <0\.8<0\.8\), formatting, and a 3\-of\-5 flagship\-LLM\-failure difficulty filter\. Stage 3 LLM\-as\-judge applied multidimensional feature scoring\. An agentic refinement loop modified ∼\\sim13\\frac\{1\}\{3\} of the candidate pool to remove residual boundary defects\. Uses Has the dataset been used for any tasks already? Yes: zero\-shot evaluation of 4242 frontier LLMs \(§5\), including web\-search\-augmented evaluation and parameter scaling analysis\. What other tasks could the dataset be used for? Calibration of LLM confidence, retrieval\-augmented generation evaluation, prompt\-engineering benchmarking, multilingual transfer studies \(after translation\), and meta\-evaluation of LLM\-as\-judge frameworks\. Are there tasks for which the dataset should not be used? KINA should not be used as a primary signal for clinical, legal, or high\-stakes deployment decisions\. The benchmark is designed to differentiate relative model capability under controlled conditions, not to certify absolute correctness on real\-world tasks\. Distribution Will the dataset be distributed to third parties? Yes\. The dataset is released under CC\-BY\-4\.0 and the evaluation code under MIT license at https://huggingface\.co/datasets/2077AIDataFoundation/KINA\. To mitigate contamination, the test split is distributed as an encrypted archive with a canary string; the development split is distributed openly\. Will there be a leaderboard? Will the dataset be updated? KINA will be refreshed \(versioned as KINA\-vx\.yx\.y\) when the strongest evaluated model exceeds 70%70\\% overall accuracy, to prevent saturation\. Maintenance Who is supporting/hosting/maintaining the dataset? The KINA team commits to maintaining the dataset and leaderboard for at least 3636 months from initial release\. Issue tracking and contribution guidelines are at the project repository\. If others want to contribute to the dataset, is there a mechanism for them to do so? Yes\. We accept community contributions of subfield\-specific items via a standardized submission form, subject to the same three\-stage verification pipeline described in §4\.`

Similar Articles

BAGEL: Benchmarking Animal Knowledge Expertise in Language Models

arXiv cs.CL

BAGEL is a new benchmark for evaluating animal-related knowledge in large language models, constructed from diverse scientific sources and covering taxonomy, morphology, habitat, behavior, and species interactions through closed-book question-answer pairs. The benchmark enables fine-grained analysis across taxonomic groups and knowledge categories, providing insights into model strengths and failure modes for biodiversity applications.