NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment
Summary
NovGauge is a human-anchored benchmark for diagnosing LLMs' capability in paper novelty assessment across task, problem, and method dimensions. Evaluation of 18 LLMs reveals high hallucination rates and logical mismatches, indicating current models are unreliable for this task.
View Cached Full Text
Cached at: 09/12/26, 08:24 AM
# NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs’ Capability in Paper Novelty Assessment
Source: [https://arxiv.org/html/2609.11234](https://arxiv.org/html/2609.11234)
Guoqiang ZhangAffiliation:College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, ChinaEmail:[zitago121@gmail\.com](mailto:)Kexin TanAffiliation:College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, ChinaAffiliation:Graduate School of Engineering, The University of Tokyo, Tokyo, JapanMing ZhangAffiliation:College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, ChinaWenqing JingAffiliation:College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, ChinaZhonghan YueAffiliation:College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, ChinaJiayi ChenAffiliation:College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, ChinaShiqiang WuAffiliation:College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, ChinaShaofan LiuAffiliation:College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, ChinaYue ZhangAffiliation:Atom Infinite Pte\. Ltd\., SingaporeYuankai YingAffiliation:College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, ChinaYang ShiAffiliation:AtomInnoLabTao GuiAffiliation:College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, ChinaQi ZhangAffiliation:College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, ChinaXuanjing HuangAffiliation:College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China
###### Abstract
Large language models \(LLMs\) are increasingly used in peer review at major AI conferences, yet novelty remains a persistent weak point\. Existing benchmarks assess novelty as a single holistic score, making it difficult to diagnose which dimension a model misjudges or whether its evidence is faithful\. We presentNovGauge, a human\-anchored benchmark for fine\-grained novelty assessment diagnosis\. The benchmark contains 619 paper pairs and 50 multi\-paper sets, drawn from two expert sources: ICLR reviewer overlap claims and survey co\-citations\. Instances are independently labeled along three dimensions:*task*,*problem*, and*method*, capturing application goals, technical challenges, and solution approaches\. We propose a cascading diagnostic pipeline that verifies per\-dimension correctness, evidence grounding, and logical support\. Evaluation of 18 LLMs shows hallucination rates ranging from 0% to 39% across dimensions, and among non\-hallucinated correct\-positive judgments, over 70% cite evidence fails to logically support the stated reason\. The best\-performing model, GPT\-5\.5, achieves 43–72% Verified F1 across dimensions, while most models retain less than half of their raw F1 after faithfulness verification\. These results suggest that current LLMs remain far from reliable scientific novelty assessment, particularly when correctness is conditioned on faithful evidence grounding\.
## 1Introduction
Figure 1:One label, three dimensions\.An ICLR 2026 submission compared against flagged prior work\.NovGaugeevaluates each dimension independently, revealing that this pair shares its method while targeting a distinct task\.Figure 2:Overview of theNovGaugeframework\.Two human\-anchored construction pipelines produce pairs and multi\-paper sets with task, problem, and method labels\. Model outputs pass through a cascading evaluation funnel that checks correctness, evidence hallucination, and reasoning mismatch, yielding a per\-dimension diagnosis\.Large language models already shape peer review at scale through two channels\. Reviewers increasingly rely on general\-purpose LLMs to draft reviews, with audits estimating that up to 20% of text at major AI venues is now LLM\-generated\([Liang et al\., 2024a](https://arxiv.org/html/2609.11234#bib.bib20);[Shen and Wang, 2026](https://arxiv.org/html/2609.11234#bib.bib22)\)\. Meanwhile, dedicated automated review systems continue to emerge\([Zhou et al\., 2024](https://arxiv.org/html/2609.11234#bib.bib1);[D’Arcy et al\., 2024](https://arxiv.org/html/2609.11234#bib.bib2);[Zhu et al\., 2025](https://arxiv.org/html/2609.11234#bib.bib3)\), with some already approaching human\-level agreement on acceptance decisions\([Weng et al\., 2025](https://arxiv.org/html/2609.11234#bib.bib4)\)\. However, the reliability of LLM\-generated reviews remains questionable[Wu et al\. \(2026a\)](https://arxiv.org/html/2609.11234#bib.bib8);[Shin et al\. \(2025b\)](https://arxiv.org/html/2609.11234#bib.bib33);[Du et al\. \(2024\)](https://arxiv.org/html/2609.11234#bib.bib30)\. Recent evaluations show that LLM reviews correlate weakly with human judgments and can be trivially gamed\([Baumann et al\., 2026](https://arxiv.org/html/2609.11234#bib.bib27)\)\. This concern is especially important for novelty assessment: large\-scale analysis of ICLR 2024–2025 identifies lack of novelty as one of the most frequently cited weaknesses in rejected papers\([Kargaran et al\., 2025](https://arxiv.org/html/2609.11234#bib.bib28)\)\.
Despite its importance, novelty remains a particularly challenging criterion for LLM reviewers\. Studies show that LLMs systematically underweight novelty in their reviews\([Shin et al\., 2025a](https://arxiv.org/html/2609.11234#bib.bib5)\)and that their novelty ratings even correlate negatively with a paper’s eventual impact\([Jiang, 2026](https://arxiv.org/html/2609.11234#bib.bib23)\)\. These results confirm that LLM novelty assessment is unreliable, but a single accuracy number on a holistic label reveals only that a model underperforms, not which aspect of novelty it misjudges\. This leaves the operational review question unanswered: given a submission and a flagged prior work, which novelty dimension overlaps, and is the cited evidence faithful?
Reliable diagnosis demands two capabilities\. The first is*per\-dimension*evaluation\. Novelty is not monolithic: scientometric studies decompose it into research question and method\([Luo et al\., 2022](https://arxiv.org/html/2609.11234#bib.bib21);[Shen et al\., 2026](https://arxiv.org/html/2609.11234#bib.bib24)\), and a paper may overlap with prior work on one dimension while remaining novel on others \(Figure[1](https://arxiv.org/html/2609.11234#S1.F1)\)\. Current benchmarks, however, either let dimensions vary across instances\([Moussa et al\., 2025](https://arxiv.org/html/2609.11234#bib.bib12)\)or treat novelty as one undifferentiated score among several review criteria\([Qiao et al\., 2026](https://arxiv.org/html/2609.11234#bib.bib26);[Schopf and Färber, 2026](https://arxiv.org/html/2609.11234#bib.bib11);[Lin et al\., 2025](https://arxiv.org/html/2609.11234#bib.bib17);[Wu et al\., 2026b](https://arxiv.org/html/2609.11234#bib.bib10)\)\. The second is*faithfulness*verification\. Even when a model’s judgment is correct, the reasoning behind it may rest on fabricated evidence or logical gaps\. No existing benchmark systematically diagnoses whether failures stem from judgment errors, fabricated evidence, or logical gaps\([Schopf and Färber, 2026](https://arxiv.org/html/2609.11234#bib.bib11)\)\.
We introduceNovGauge, a benchmark designed to enable fine\-grained diagnosis of novelty judgment\. Every instance is independently labeled for task, problem, and method, providing the fixed per\-dimension ground truth\. Labels are anchored in two complementary human sources, ICLR reviewer overlap claims and annotator\-verified survey co\-citations\. Beyond pairwise comparison, the benchmark further includes 50 multi\-paper grouping sets\. Rather than stopping at binary correctness, the evaluation protocol applies a*cascading*diagnostic pipeline that checks correctness, evidence hallucination, and reasoning mismatch\. Only predictions that survive all three stages count towardVerified F1, ensuring the model reached*the right answer for the right reason*\.
Our contributions are:
- •Abenchmark of 619 pairs and 50 multi\-paper setswith fixed three\-dimension ground truth, anchored in ICLR reviewer statements and annotator\-verified survey co\-citations\.
- •Acascading evaluation protocolthat diagnoses novelty judgment per dimension and audits evidence faithfulness through hallucination and mismatch checks\.
- •Asystematic evaluation of 18 LLMsrevealing that even correct judgments are frequently backed by fabricated or logically disconnected evidence, with failure patterns varying sharply across dimensions\.
Figure[2](https://arxiv.org/html/2609.11234#S1.F2)summarizes the benchmark construction and evaluation flow\.
## 2Related Work
### 2\.1Automated Paper Review
LLM\-based peer review has advanced from single\-model critique generation\([Zhou et al\., 2024](https://arxiv.org/html/2609.11234#bib.bib1)\)to multi\-agent architectures\([D’Arcy et al\., 2024](https://arxiv.org/html/2609.11234#bib.bib2);[Yu et al\., 2024](https://arxiv.org/html/2609.11234#bib.bib29);[Bougie and Watanabe, 2024](https://arxiv.org/html/2609.11234#bib.bib32)\)and fine\-tuned models trained on structured reasoning chains\([Zhu et al\., 2025](https://arxiv.org/html/2609.11234#bib.bib3);[Weng et al\., 2025](https://arxiv.org/html/2609.11234#bib.bib4);[Taechoyotin and Acuna, 2025](https://arxiv.org/html/2609.11234#bib.bib31)\)\. However, systematic analyses reveal that LLMs over\-attend to technical validity while underweighting novelty\([Shin et al\., 2025a](https://arxiv.org/html/2609.11234#bib.bib5)\), assign inflated scores to LLM\-generated papers\([Li et al\., 2025](https://arxiv.org/html/2609.11234#bib.bib6)\), and align with human judgments only on lower\-quality submissions\([Liang et al\., 2024b](https://arxiv.org/html/2609.11234#bib.bib7)\)\.[Li et al\. \(2026\)](https://arxiv.org/html/2609.11234#bib.bib9)further show that aligning critique focus with human experts is a prerequisite for reliable scoring\. These findings motivate benchmarks that diagnose novelty judgment specifically\.
### 2\.2Novelty and Similarity Benchmarks
Table[1](https://arxiv.org/html/2609.11234#S2.T1)situates our work within a progression from document\-level similarity to fine\-grained novelty assessment\. SciDocs\([Cohan et al\., 2020](https://arxiv.org/html/2609.11234#bib.bib15)\)and SciRepEval\([Singh et al\., 2023](https://arxiv.org/html/2609.11234#bib.bib16)\)evaluate retrieval relevance from citation\-graph labels, measuring similarity holistically without distinguishing which dimensions overlap\.
Recent novelty benchmarks move beyond retrieval to assess whether a paper or idea advances the state of the art\. NovBench\([Wu et al\., 2026b](https://arxiv.org/html/2609.11234#bib.bib10)\)frames novelty as text generation from a single paper’s introductory claim, while literature\-grounded systems assess manuscript novelty against existing work\([Mostafa et al\., 2026](https://arxiv.org/html/2609.11234#bib.bib13);[Zhang et al\., 2026a](https://arxiv.org/html/2609.11234#bib.bib25)\)\. Idea\-level benchmarks evaluate research ideas against retrieved related works: RINoBench\([Schopf and Färber, 2026](https://arxiv.org/html/2609.11234#bib.bib11)\)and InnoEval\([Qiao et al\., 2026](https://arxiv.org/html/2609.11234#bib.bib26)\)assign novelty scores, while ScholarEval\([Moussa et al\., 2025](https://arxiv.org/html/2609.11234#bib.bib12)\)and the Idea Novelty Checker\([Shahid et al\., 2025](https://arxiv.org/html/2609.11234#bib.bib14)\)dynamically identify contribution dimensions for each idea\. Pairwise approaches take designated paper pairs as input but target review quality assessment\([Zhang et al\., 2026b](https://arxiv.org/html/2609.11234#bib.bib18);[Zheng et al\., 2026](https://arxiv.org/html/2609.11234#bib.bib19)\)or relative novelty ranking\([Lin et al\., 2025](https://arxiv.org/html/2609.11234#bib.bib17)\)rather than overlap detection on fixed dimensions\.
BenchmarkTaskPairwiseGroupingExpert\#Dim\.ScaleSciDocs\([Cohan et al\., 2020](https://arxiv.org/html/2609.11234#bib.bib15)\)Retrieval✓——125K papersSciRepEval\([Singh et al\., 2023](https://arxiv.org/html/2609.11234#bib.bib16)\)Retrieval✓——160K papersNovBench\([Wu et al\., 2026b](https://arxiv.org/html/2609.11234#bib.bib10)\)Generation——✓11\.7K pairsRINoBench\([Schopf and Färber, 2026](https://arxiv.org/html/2609.11234#bib.bib11)\)Scoring——✓11\.4K ideasScholarEval\([Moussa et al\., 2025](https://arxiv.org/html/2609.11234#bib.bib12)\)Scoring \+ RAG——✓var\.117 ideasIdea Novelty Checker\([Shahid et al\., 2025](https://arxiv.org/html/2609.11234#bib.bib14)\)Scoring \+ RAG——✓167 ideasInnoEval\([Qiao et al\., 2026](https://arxiv.org/html/2609.11234#bib.bib26)\)Scoring——✓1217 ideasSchNovel\([Lin et al\., 2025](https://arxiv.org/html/2609.11234#bib.bib17)\)Ranking✓——115K pairsOurs \(pairwise\)Classification✓—✓3619 pairsOurs \(grouping\)Grouping—✓✓350 setsTable 1:Comparison with existing novelty and similarity benchmarks\.\#Dim\.denotes the number of orthogonal dimensions along which novelty overlap is independently judged\. These dimensions include task, problem, and method, rather than general evaluation rubrics such as relevance or clarity\. “var\.” denotes dimensions that vary across instances\.NovGaugeaddresses these gaps by providing per\-instance ground truth under a fixed three\-dimension taxonomy, enabling systematic diagnosis of where and why LLM novelty judgments fail\.
## 3Benchmark Construction
SourceSizeMain Dim\.PairwiseICLR review positives236method and problemSurvey co\-citation positives227taskSurvey cross\-group negatives156method and taskMulti\-paper groupingSurvey co\-citation sets50task, problem, methodTable 2:Final benchmark composition\. “Main Dim\.” denotes the dominant dimension in each subset\.### 3\.1Design Principles
NovGaugeaddresses the limitations of existing benchmarks through three core design principles\.
#### Principle 1: Per\-dimension decomposition\.
Novelty is not monolithic\. We decompose it into three orthogonal dimensions and evaluate each one independently\.
- •Task: the application domain \(e\.g\.,*image classification*vs\.*machine translation*\)\.
- •Problem: the core research challenge \(e\.g\.,*distribution shift*vs\.*catastrophic forgetting*\)\.
- •Method: the technical approach \(e\.g\.,*contrastive learning*vs\.*supervised fine\-tuning*\)\.
This follows the*what*,*why*, and*how*of a research contribution and extends prior taxonomies that separate research question and method\([Luo et al\., 2022](https://arxiv.org/html/2609.11234#bib.bib21);[Shen et al\., 2026](https://arxiv.org/html/2609.11234#bib.bib24)\)\. We treat task as a separate axis because papers may share a problem or method while targeting different domains\. Other contribution types are represented within these dimensions rather than treated as separate axes\. Figure[1](https://arxiv.org/html/2609.11234#S1.F1)illustrates partial overlap across dimensions\. For this reason, every instance carries fixed per\-dimension ground truth, enabling us to identify which dimension a model misjudges\.
#### Principle 2: Expert\-sourced labels\.
Every label is anchored in explicit expert evidence\. Positive labels originate from two complementary sources: ICLR reviewer overlap claims from public reviews, and survey co\-citation groupings verified by three annotators\. Negatives are constructed through cross\-group sampling, where pairs are drawn from papers a survey author placed into*distinct*similarity groups within the same sentence\.
#### Principle 3: Evidence\-grounded evaluation\.
Correct judgments are insufficient if the reasoning behind them is unreliable\. We therefore require models to ground every positive judgment in verbatim evidence spans and evaluate outputs through the three\-stage cascade described in Section[3\.4](https://arxiv.org/html/2609.11234#S3.SS4)\.
### 3\.2Data Construction
NovGaugeis constructed from two complementary sources \(Table[2](https://arxiv.org/html/2609.11234#S3.T2)\), both verified by three PhD\-level annotators\. ICLR\-grounded pairs predominantly capture method and problem overlap, while survey\-derived pairs are dominated by task\-level similarity\. The ICLR source contributes positives only\. The survey source supplies both positives and dimension\-specific negatives\.
Source 1: ICLR Review\-Grounded Pairs\.We process 42,682 ICLR 2023–2026 OpenReview forums through a two\-stage LLM pipeline that extracts and verifies novelty\-overlap claims from official reviews and author rebuttals\. Three PhD\-level annotators adjudicate every retained candidate, verifying prior\-work identity, dimension labels, and reviewer evidence\. The released test set retains only high\-confidence pairs with convergent evidence from multiple reviewers and author concession, yielding236 pairs\. Detailed annotation procedures are in Appendix[D](https://arxiv.org/html/2609.11234#A4)\.
Source 2: Survey Co\-citation Groups\.We extract similarity groupings from 459 CS arXiv surveys \(2023–2024\)\. A rule\-based extractor identifies 1,130 co\-citation sentences with explicit similarity cues, and an LLM proposes per\-dimension groupings\. The same annotators verify each grouping, yielding118 verified sentences\. We enumerate within\-group pairs as positives \(227 pairs\) and cross\-group pairs as negatives \(156 pairs\)\. The survey source also contributes85 multi\-paper grouping instancesderived from 50 multi\-paper sets \(Section[3\.3](https://arxiv.org/html/2609.11234#S3.SS3)\)\. Detailed pipeline is in Appendix[E](https://arxiv.org/html/2609.11234#A5)\.
Final composition and quality control\.Table[2](https://arxiv.org/html/2609.11234#S3.T2)summarizes the four subsets that make up the benchmark\. At the paper\-pair level, the pairwise subset contains 619 records, of which 463 are labeled positive and 156 are labeled negative\. Because each available dimension is evaluated independently, these records yield 875 dimension\-specific evaluation instances, of which 649 carry positive labels and 226 carry negative labels\. To verify label consistency, we audited a random sample of 30 items stratified by dimension\. The same annotators re\-examined the labels independently, yielding 93\.3% agreement with the original annotations\. Table[5](https://arxiv.org/html/2609.11234#A4.T5)reports the per\-dimension breakdown of this audit and the main causes of disagreement\.
### 3\.3Evaluation Formulation
We define two complementary evaluations, both scored independently per dimension\. The first tests pairwise comparison, where a model decides whether a submission and a specific prior work overlap on the target dimension\. The second requires synthesizing evidence across multiple documents, a more realistic setting that mirrors how reviewers survey related work\.
#### Pairwise Novelty Judgment\.
Input: a pair of papers\(A,B\)\(A,B\)and a target dimensiond∈\{task,problem,method\}d\\in\\\{\\text\{task\},\\text\{problem\},\\text\{method\}\\\}, with each paper rendered at one of three content granularities \(abstract, abstract \+ introduction, or full paper\)\. A separate model call, with a dimension\-specific prompt, is issued for each available labeled dimension\. A paper\-pair record with labels on multiple dimensions therefore contributes one dimension\-specific evaluation instance per label rather than one holistic instance\.
Output: a structured JSON record judging similarity on*dd*and grounding it in verbatim evidence:
> \{is\_similar: bool, evidence\_a: <verbatim span from A\>, evidence\_b: <verbatim span from B\>, reason: <≤\\leq2\-sentence justification\> \}
Predictions are evaluated against the ground\-truth labels in Table[2](https://arxiv.org/html/2609.11234#S3.T2)\.
#### Multi\-paper Similarity Grouping\.
Given a*set*of related papers, the model must identify which subsets share a given dimension\.
Input: a set ofNNpapers \(3≤N≤103\\leq N\\leq 10\), together with a target dimensiondd\. TheNNpapers are not all mutually similar\. Only certain subsets share the dimension\.
Output: the model must partition theNNpapers into similarity subgroups for dimensiondd:
> \{groups: \[ \{ paper\_indices: \[…\], evidence: \{…\}, reason: …\}, …\] \}
Each output subgroup is a maximal set of papers judged mutually similar ondd\. Each input set is evaluated on every dimension for which a ground\-truth label exists, yielding 85 grouping instances across the 50 input sets \(44 task, 13 problem, 28 method\)\.
### 3\.4Cascading Evaluation Protocol
A model may produce a correct label with fabricate supporting evidence, or cite real evidence that does not logically support its conclusion\. To separate grounded reasoning from surface\-level correctness, we evaluate outputs along three cascading stages \(Figure[2](https://arxiv.org/html/2609.11234#S1.F2), right panel\), each filtering the outputs of the previous stage\.
Stage 1: Correctness\.For pairwise judgment, we report per\-dimension Accuracy and F1 for classifying whether a pair is similar\. For multi\-paper grouping, each of the 85 \(set, dimension\) instances yields a predicted partition, scored against the ground\-truth partition by converting both into their sets of within\-set paper pairs and computing pairwise Precision, Recall, and F1 \(TP=\|pred pairs∩GT pairs\|\\text\{TP\}=\|\\text\{pred pairs\}\\cap\\text\{GT pairs\}\|\)\. These per\-instance F1 scores are then macro\-averaged within each dimension\. For both evaluations, every reported metric is the mean acrossn=3n\{=\}3inference samples\.
Stage 2: Hallucination Rate\.Every positive output must cite verbatim evidence from the source papers\. For pairwise predictions, this means one span per paper\. For grouping predictions, one span per member in each predicted subgroup\. Each cited span is verified against the source paper’s text via substring matching\. A span is classified*hallucinated*if verification fails\. The Hallucination Rate is the fraction of positive outputs \(pairs or subgroups\) in which at least one cited span is hallucinated\. Details are in Appendix[F](https://arxiv.org/html/2609.11234#A6)\.
Stage 3: Mismatch Rate\.Among the*non\-hallucinated*positive outputs, an external LLM judge decides whether the stated reason is logically entailed by the cited evidence\. The judge model and prompt are detailed in Section[4](https://arxiv.org/html/2609.11234#S4)and Appendix[P](https://arxiv.org/html/2609.11234#A16)\.
Verified F1\.The three stages compose into a single summary metric:Verified F1\. Only true positives that pass all three stages count in the numerator, while hallucinated or mismatched predictions remain in the denominator\. For multi\-paper grouping, subgroups are first classified by their hallucination and mismatch status, then converted to pairs for F1 computation\. The gap between raw F1 and Verified F1 quantifies the*faithfulness cliff*, the fraction of seemingly correct judgments undermined by fabricated or unsupported evidence\. Formal definitions are in Appendix[C](https://arxiv.org/html/2609.11234#A3)\.
## 4Experimental Setup
We evaluateNovGaugethrough three experiments\. The first tests pairwise novelty judgment, the main task\. The second varies how much paper content is shown to the model\. The third tests whether models can group several papers that overlap along the same novelty dimension\. All experiments use the cascading protocol in §[3\.4](https://arxiv.org/html/2609.11234#S3.SS4)\.
### 4\.1Pairwise Judgment
Pairwise judgment asks a model to decide whether two papers overlap in task, problem, or method\. We evaluate 18 LLMs spanning frontier systems, mid\-scale open\-weight models, and specialised peer\-review models\. Each model produces a binary label, a short reason, and evidence spans from the papers\. We report raw F1, Hallucination Rate, Mismatch Rate, and Verified F1\. The full model list and system settings appear in Appendix[B](https://arxiv.org/html/2609.11234#A2)\.
### 4\.2Granularity Ablation
This experiment tests whether more paper context improves novelty judgment\. We run the same pairwise task under three input settings: abstract only, abstract plus introduction, and full paper\. The corresponding per\-paper content caps are 1k, 4k, and 16k tokens\. The full\-paper setting is the default for the main pairwise results\.
### 4\.3Multi\-Paper Grouping
Multi\-paper grouping asks a model to recover clusters of papers that share the same task, problem, or method overlap\. This setting is harder than pairwise judgment because the model must compare multiple papers and provide collective evidence for each predicted group\. We evaluate the 12 frontier models on 85 grouping instances, using full paper inputs\.
Table 3:Pairwise novelty judgment results\(full\-paper content granularity\)\.Haldenotes Hallucination Rate,Misdenotes Mismatch Rate, andVF1denotes Verified F1\. Definitions appear in §[3\.4](https://arxiv.org/html/2609.11234#S3.SS4)\. All metrics reported as percentages\.Boldmarks the best value per column when fewer than half the models tie\.‡GPT\-family judged by Gemini\-3\.1\-Pro\. Standard deviations appear in Appendix[H](https://arxiv.org/html/2609.11234#A8)\.
### 4\.4Implementation Details
Each input and dimension is decoded three times at temperature 0\.6, and we report mean scores with standard deviations in Appendix[H](https://arxiv.org/html/2609.11234#A8)\. All models use their maximum reasoning setting when available\. The mismatch stage uses an LLM judge at temperature 0\.2 \(prompt in Appendix[P](https://arxiv.org/html/2609.11234#A16)\)\. GPT\-family models are judged by Gemini\-3\.1\-Pro, while all other models are judged by GPT\-5\.2\. A human validation study on 100 non\-hallucinated positive outputs finds 87\.0% agreement with the judge, with Cohen’sκ=0\.74\\kappa=0\.74\(details in Appendix[K](https://arxiv.org/html/2609.11234#A11)\)\.
## 5Experimental Results
### 5\.1Main Results
Table[3](https://arxiv.org/html/2609.11234#S4.T3)reports pairwise novelty judgment results for the 18 evaluated models across task, problem, and method\. Raw F1 is often high, but the cascading evaluation exposes a large gap between correct labels and faithful reasoning \(Figure[3](https://arxiv.org/html/2609.11234#S5.F3)a\)\. The average model retains only 26% of its raw F1 after faithfulness filtering\.
#### Faithfulness filtering exposes unsupported reasoning\.
Only 5 of 54 model\-dimension cells retain at least half of their raw F1 after faithfulness filtering, as shown in Figure[3](https://arxiv.org/html/2609.11234#S5.F3)b\. The main source of degradation is mismatch\. The mismatch stage flags 72\.5% of non\-hallucinated correct\-positive predictions as unsupported because models identify similarity at the contribution level but cite evidence at the implementation level\. For example, Gemini 3\.1 Flash\-Lite reaches 89\.2% raw F1 on problem but drops to 22\.9% Verified F1\. GPT\-5\.5 is the clear exception, with zero detected hallucinations and mismatch rates below 46% across all dimensions\. Notably, CycleReviewer\-8B and DeepReviewer\-14B reach at most 15\.9% Verified F1, suggesting that review\-oriented fine\-tuning does not necessarily transfer to fine\-grained, evidence\-grounded novelty judgment\.
Figure 3:Panel \(a\) shows the faithfulness cascade for Claude Opus 4\.7 on the method dimension, where Raw F1 drops to low Verified F1 after hallucination and mismatch checks\. Panel \(b\) shows the distribution of VF1/F1 ratios across the 54 model\-dimension cells of Table[3](https://arxiv.org/html/2609.11234#S4.T3)\.
#### Method\-level novelty is the hardest to ground\.
The method dimension is the least faithful dimension for nearly all models\. Hallucination on method often exceeds task and problem by a wide margin\. For GLM\-5\.1, the method hallucination rate is 27\.2%, compared with 0\.0% on task\. Even GPT\-5\.5, the strongest model overall, retains 84% of its raw F1 on task but only 55% on method, confirming that method\-level novelty is the most fragile setting for LLM\-based review support\.
#### Incorrect positive predictions fabricate more evidence\.
Figure[4](https://arxiv.org/html/2609.11234#S5.F4)compares faithfulness on correct and incorrect positive predictions\. Hallucination Rate rises from 6\.4% on true positives to 24\.9% on false positives\. Mismatch Rate also increases, from 72\.5% to 85\.1%\. Only 11% of false\-positive predictions survive the full cascade, compared with 26% of true positives\. This suggests that false positives are not merely label errors, they often combine an unsound conclusion with weaker evidence selection\.
Figure 4:Faithfulness on correct versus incorrect positive predictions\. Panel mean over 18 models\.
#### LLMs are conservative similarity judges\.
Models achieve 90–97% negative accuracy but only 47–59% positive accuracy, as shown in Figure[5](https://arxiv.org/html/2609.11234#S5.F5)\. The gap also holds within the survey source, which contains both labels: positive accuracy is 44–59%, compared with 90–97% for negatives \(Table[16](https://arxiv.org/html/2609.11234#A8.T16)\)\. Among the 61 pairs missed by all 18 LLMs, every case is a false negative\. In other words, LLMs are far more likely to miss genuine similarity than to fabricate spurious overlap, particularly when two papers differ in framing or implementation despite sharing a core contribution\.
Figure 5:Positive accuracy \(correctly identifying similar pairs\) versus negative accuracy \(correctly rejecting not\-similar pairs\) across three dimensions, averaged over 18 models\.Table 4:Multi\-paper similarity grouping results\(full content granularity, 85 instances\)\. Metrics follow the same cascade as Table[3](https://arxiv.org/html/2609.11234#S4.T3)\. All metrics reported as percentages\.‡GPT\-family judged by Gemini\-3\.1\-Pro\.Boldmarks the best value per column when fewer than half the models tie\.
### 5\.2Granularity Results
Figure 6:Verified F1 across three content granularities for five representative models\.We rerun pairwise judgment at three content granularities to test whether richer paper context improves both accuracy and faithfulness\. Figure[6](https://arxiv.org/html/2609.11234#S5.F6)shows five representative models\. Full results appear in Appendix[I](https://arxiv.org/html/2609.11234#A9)\.
#### More context improves raw F1 but reduces Verified F1\.
Raw F1 increases for 13 of 18 models, with a mean gain of 3\.9 points\. But verified F1 decreases for 16 of 18 models, with a mean loss of 5\.0 points\. GPT\-5\.5 and GPT\-5\.4\-mini are the only models that maintain or improve Verified F1 with longer content\. The pattern suggests that richer context helps models resolve ambiguous overlap labels, but makes it harder to keep the rationale tightly grounded\.
#### Longer inputs invite mismatch, not mainly hallucination\.
The degradation comes from unsupported reasoning rather than fabricated spans\. Mismatch Rate rises from 63\.6% on abstracts to 72\.6% on full papers, while Hallucination Rate remains low at 0\.5% on abstracts and 6\.3% on full papers\. Longer papers give models more plausible evidence to choose from, but much of that evidence is not logically tied to the stated novelty judgment\.
### 5\.3Grouping Results
Table[4](https://arxiv.org/html/2609.11234#S5.T4)reports results for the 12 frontier models on 85 multi\-paper grouping instances\. The task uses the same cascade as pairwise judgment, but each predicted group must cite evidence from its member papers to justify the shared\-dimension judgment\.
#### Grouping amplifies the faithfulness cliff\.
Panel\-mean Verified F1 drops to 17%, 8%, and 7% for task, problem, and method\. These scores retain only 38–62% of the corresponding pairwise values\. This indicates that models lose more faithful reasoning ability when the unit of judgment shifts from one paper pair to a multi\-paper set\. GPT\-5\.5 again leads all three dimensions, with Verified F1 of 41\.7%, 25\.2%, and 16\.7%\.
#### Shorter evidence spans reduce hallucination but do not solve mismatch\.
Hallucination rates are lower in grouping than in pairwise evaluation, with panel means of 0\.0%, 3\.0%, and 5\.6%\. Grouping tends to produce shorter evidence spans than pairwise judgment, averaging 34 words compared with 53 words in the pairwise setting\. This leaves less room for fabrication\. The main failure is still mismatch, with panel means of 75%, 87%, and 86%\. Appendix[J](https://arxiv.org/html/2609.11234#A10)examines the resulting rank drift between pairwise and grouping leaderboards\.
## 6Conclusion
We introducedNovGauge, a benchmark that evaluates LLM novelty judgment with per\-dimension labels and cascading faithfulness checks\. Our evaluation of 18 LLMs reveals that accuracy is a misleading metric\. Only 5 of 54 model\-dimension pairs retain more than half of their raw F1 after faithfulness filtering\. Multi\-paper grouping amplifies this faithfulness cliff, where even GPT\-5\.5 degrades substantially\. This bottleneck is most severe for method\-level novelty, worsens with longer paper context, and often arises from a gap between contribution\-level reasoning and implementation\-level evidence\. Across settings, LLMs also exhibit a conservative similarity bias, tending to reject genuine overlap more often than they invent spurious similarity\. Together, these findings show that reliable novelty assessment requires more faithful links between verdicts, reasons, and evidence\.
## Limitations
- •Domain scope\.All data derives from ICLR and CS arXiv surveys\. Cross\-disciplinary generalization is untested\.
- •LLM judge dependency\.The cascading metric relies on LLM judges \(GPT\-5\.2 and Gemini\-3\.1\-Pro\)\. Researchers without access can substitute another capable model at some cost in comparability\.
- •Per\-dimension negative imbalance\.Negative samples are uneven across dimensions, with the problem dimension having only 33 negatives, yielding wider variance in F1 estimates\.
- •Reviewer LLM contamination\.A fraction of ICLR 2025–2026 reviews may have been drafted with LLM assistance, though the underlying novelty judgments still reflect human expertise\.
- •Retrieval scope\.NovGaugeevaluates judgment after candidate papers are retrieved\. It can serve as the judgment stage of a retrieval pipeline\. Evaluation of the full retrieval process is left to future work\.
## Ethics Statement
All data derives from publicly available OpenReview submissions and arXiv papers\. No personally identifiable information beyond author names on public papers is included\. The benchmark is intended for research on evaluation methodology and does not endorse automated replacement of human peer review\.
## References
- Baumannet al\.\(2026\)J\. Baumann, J\. Pei, S\. Koyejo, and D\. HovyStop automating peer review without rigorous evaluation\.arXiv preprint arXiv:2605\.03202\.Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p1.1)\.
- Bougie and Watanabe \(2024\)N\. Bougie and N\. WatanabeGenerative adversarial reviews: when llms become the critic\.arXiv preprint arXiv:2412\.10415\.Cited by:[§2\.1](https://arxiv.org/html/2609.11234#S2.SS1.p1.1)\.
- Cohanet al\.\(2020\)A\. Cohan, S\. Feldman, I\. Beltagy, D\. Downey, and D\. WeldSPECTER: document\-level representation learning using citation\-informed transformers\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 2270–2282\.External Links:[Link](https://aclanthology.org/2020.acl-main.207/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.207)Cited by:[§2\.2](https://arxiv.org/html/2609.11234#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2609.11234#S2.T1.2.2.1)\.
- Duet al\.\(2024\)J\. Du, Y\. Wang, W\. Zhao, Z\. Deng, S\. Liu, R\. Lou, H\. P\. Zou, P\. N\. Venkit, N\. Zhang, M\. Srinath,et al\.LLMs assist nlp researchers: critique paper \(meta\-\) reviewing\.InProceedings of the 2024 conference on empirical methods in natural language processing,pp\. 5081–5099\.Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p1.1)\.
- D’Arcyet al\.\(2024\)M\. D’Arcy, T\. Hope, L\. Birnbaum, and D\. DowneyMARG: multi\-agent review generation for scientific papers\.External Links:2401\.04259,[Link](https://arxiv.org/abs/2401.04259)Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.11234#S2.SS1.p1.1)\.
- Jiang \(2026\)B\. JiangHindSight: evaluating llm\-generated research ideas via future impact\.arXiv preprint arXiv:2603\.15164\.Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p2.1)\.
- Kargaranet al\.\(2025\)A\. H\. Kargaran, N\. Nikeghbal, J\. Yang, and N\. OusidhoumInsights from the iclr peer review and rebuttal process\.arXiv preprint arXiv:2511\.15462\.Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p1.1)\.
- Liet al\.\(2026\)B\. Li, H\. Ma, Y\. Wang, J\. Yang, Y\. Zheng, X\. Chen, X\. Huang, and X\. QiuBeyond rating: a comprehensive evaluation and benchmark for ai reviews\.External Links:2604\.19502,[Link](https://arxiv.org/abs/2604.19502)Cited by:[§2\.1](https://arxiv.org/html/2609.11234#S2.SS1.p1.1)\.
- Liet al\.\(2025\)R\. Li, J\. Gu, P\. Kung, H\. Xia, J\. liu, X\. Kong, Z\. Sui, and N\. PengLLM\-reval: can we trust llm reviewers yet?\.External Links:2510\.12367,[Link](https://arxiv.org/abs/2510.12367)Cited by:[§2\.1](https://arxiv.org/html/2609.11234#S2.SS1.p1.1)\.
- Lianget al\.\(2024a\)W\. Liang, Z\. Izzo, Y\. Zhang, H\. Lepp, H\. Cao, X\. Zhao, L\. Chen, H\. Ye, S\. Liu, Z\. Huang, D\. Mcfarland, and J\. Y\. ZouMonitoring AI\-modified content at scale: a case study on the impact of ChatGPT on AI conference peer reviews\.InProceedings of the 41st International Conference on Machine Learning,R\. Salakhutdinov, Z\. Kolter, K\. Heller, A\. Weller, N\. Oliver, J\. Scarlett, and F\. Berkenkamp \(Eds\.\),Proceedings of Machine Learning Research, Vol\.235,pp\. 29575–29620\.External Links:[Link](https://proceedings.mlr.press/v235/liang24b.html)Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p1.1)\.
- Lianget al\.\(2024b\)W\. Liang, Y\. Zhang, H\. Cao, B\. Wang, D\. Y\. Ding, X\. Yang, K\. Vodrahalli, S\. He, D\. S\. Smith, Y\. Yin, D\. A\. McFarland, and J\. ZouCan large language models provide useful feedback on research papers? a large\-scale empirical analysis\.NEJM AI1\(8\),pp\. AIoa2400196\.External Links:[Document](https://dx.doi.org/10.1056/AIoa2400196),[Link](https://ai.nejm.org/doi/full/10.1056/AIoa2400196),https://ai\.nejm\.org/doi/pdf/10\.1056/AIoa2400196Cited by:[§2\.1](https://arxiv.org/html/2609.11234#S2.SS1.p1.1)\.
- Linet al\.\(2025\)E\. Lin, Z\. Peng, and Y\. FangEvaluating and enhancing large language models for novelty assessment in scholarly publications\.InProceedings of the 1st Workshop on AI and Scientific Discovery: Directions and Opportunities,pp\. 46–57\.Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.11234#S2.SS2.p2.1),[Table 1](https://arxiv.org/html/2609.11234#S2.T1.2.9.1)\.
- Luoet al\.\(2022\)Z\. Luo, W\. Lu, J\. He, and Y\. WangCombination of research questions and methods: a new measurement of scientific novelty\.Journal of Informetrics16\(2\),pp\. 101282\.External Links:ISSN 1751\-1577,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.joi.2022.101282),[Link](https://www.sciencedirect.com/science/article/pii/S1751157722000347)Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p3.1),[§3\.1](https://arxiv.org/html/2609.11234#S3.SS1.SSS0.Px1.p1.2)\.
- Mostafaet al\.\(2026\)A\. Mostafa, T\. H\. Nguyen, and Z\. AhmadiWhat is novel? a knowledge\-driven framework for bias\-aware literature originality evaluation\.arXiv preprint arXiv:2602\.06054\.Cited by:[§2\.2](https://arxiv.org/html/2609.11234#S2.SS2.p2.1)\.
- Moussaet al\.\(2025\)H\. N\. Moussa, P\. Q\. Da Silva, D\. Adu\-Ampratwum, A\. East, Z\. Lu, N\. Puccetti, M\. Xue, H\. Sun, B\. P\. Majumder, and S\. KumarScholarEval: research idea evaluation grounded in literature\.arXiv preprint arXiv:2510\.16234\.Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.11234#S2.SS2.p2.1),[Table 1](https://arxiv.org/html/2609.11234#S2.T1.2.6.1)\.
- Qiaoet al\.\(2026\)S\. Qiao, Y\. Wei, X\. Wang, B\. Wu, B\. Xue, N\. Zhang, H\. A\. Rahmani, Y\. Wang, Q\. Zhang, K\. Ding,et al\.InnoEval: on research idea evaluation as a knowledge\-grounded, multi\-perspective reasoning problem\.arXiv preprint arXiv:2602\.14367\.Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.11234#S2.SS2.p2.1),[Table 1](https://arxiv.org/html/2609.11234#S2.T1.2.8.1)\.
- Schopf and Färber \(2026\)T\. Schopf and M\. FärberIs this idea novel? An automated benchmark for judgment of research ideas\.Note:Accepted to LREC 2026External Links:2603\.10303,[Link](https://arxiv.org/abs/2603.10303)Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.11234#S2.SS2.p2.1),[Table 1](https://arxiv.org/html/2609.11234#S2.T1.2.5.1)\.
- Shahidet al\.\(2025\)S\. Shahid, M\. Radensky, R\. Fok, P\. Siangliulue, D\. S\. Weld, and T\. HopeLiterature\-grounded novelty assessment of scientific ideas\.InProceedings of the Fifth Workshop on Scholarly Document Processing \(SDP 2025\),pp\. 96–113\.Cited by:[§2\.2](https://arxiv.org/html/2609.11234#S2.SS2.p2.1),[Table 1](https://arxiv.org/html/2609.11234#S2.T1.2.7.1)\.
- Shen and Wang \(2026\)S\. Shen and K\. WangDetecting ai\-generated content in academic peer reviews\.arXiv preprint arXiv:2602\.00319\.Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p1.1)\.
- Shenet al\.\(2026\)Y\. Shen, M\. Liu, D\. Zhou, and L\. HuangNavigating ideation space: decomposed conceptual representations for positioning scientific ideas\.arXiv preprint arXiv:2601\.08901\.Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p3.1),[§3\.1](https://arxiv.org/html/2609.11234#S3.SS1.SSS0.Px1.p1.2)\.
- Shinet al\.\(2025a\)H\. Shin, J\. Tang, Y\. Lee, N\. Kim, H\. Lim, J\. Y\. Cho, H\. Hong, M\. Lee, and J\. KimMind the blind spots: a focus\-level evaluation framework for LLM reviews\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 35630–35656\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.1805/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.1805),ISBN 979\-8\-89176\-332\-6Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.11234#S2.SS1.p1.1)\.
- Shinet al\.\(2025b\)H\. Shin, J\. Tang, Y\. Lee, N\. Kim, H\. Lim, J\. Y\. Cho, H\. Hong, M\. Lee, and J\. KimMind the blind spots: a focus\-level evaluation framework for llm reviews\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 35618–35644\.Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p1.1)\.
- Singhet al\.\(2023\)A\. Singh, M\. D’Arcy, A\. Cohan, D\. Downey, and S\. FeldmanSciRepEval: a multi\-format benchmark for scientific document representations\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 5548–5566\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.338/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.338)Cited by:[§2\.2](https://arxiv.org/html/2609.11234#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2609.11234#S2.T1.2.3.1)\.
- Taechoyotin and Acuna \(2025\)P\. Taechoyotin and D\. AcunaRemor: automated peer review generation with llm reasoning and multi\-objective reinforcement learning\.arXiv preprint arXiv:2505\.11718\.Cited by:[§2\.1](https://arxiv.org/html/2609.11234#S2.SS1.p1.1)\.
- Wenget al\.\(2025\)Y\. Weng, M\. Zhu, G\. Bao, H\. Zhang, J\. Wang, Y\. Zhang, and L\. YangCycleResearcher: improving automated research via automated review\.InProceedings of the Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=bjcsVLoHYs)Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.11234#S2.SS1.p1.1)\.
- Wuet al\.\(2026a\)S\. Wu, O\. Jiang, Y\. Zhao, T\. Hu, Y\. Ma, K\. Zhang, M\. Patwardhan, and A\. CohanCan ai be a good peer reviewer? a survey of peer review process, evaluation, and the future\.External Links:2604\.27924,[Link](https://arxiv.org/abs/2604.27924)Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p1.1)\.
- Wuet al\.\(2026b\)W\. Wu, Y\. Zhao, Y\. Wang, S\. Li, J\. Shao, Y\. Long, and C\. ZhangNovBench: evaluating large language models on academic paper novelty assessment\.arXiv preprint arXiv:2604\.11543\.Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.11234#S2.SS2.p2.1),[Table 1](https://arxiv.org/html/2609.11234#S2.T1.2.4.1)\.
- Yuet al\.\(2024\)J\. Yu, Z\. Ding, J\. Tan, K\. Luo, Z\. Weng, C\. Gong, L\. Zeng, R\. Cui, C\. Han, Q\. Sun,et al\.Automated peer reviewing in paper sea: standardization, evaluation, and analysis\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 10164–10184\.Cited by:[§2\.1](https://arxiv.org/html/2609.11234#S2.SS1.p1.1)\.
- Zhanget al\.\(2026a\)M\. Zhang, K\. Tan, Y\. Huang, Y\. Shen, C\. Ma, L\. Ju, X\. Zhang, Y\. Wang, W\. Jing, J\. Deng,et al\.Opennovelty: an llm\-powered agentic system for verifiable scholarly novelty assessment\.arXiv preprint arXiv:2601\.01576\.Cited by:[§2\.2](https://arxiv.org/html/2609.11234#S2.SS2.p2.1)\.
- Zhanget al\.\(2026b\)Y\. Zhang, H\. Zhang, W\. Ji, T\. Hua, N\. Haber, H\. Cao, and W\. LiangFrom replication to redesign: exploring pairwise comparisons for llm\-based peer review\.Advances in Neural Information Processing Systems38,pp\. 11778–11805\.Cited by:[§2\.2](https://arxiv.org/html/2609.11234#S2.SS2.p2.1)\.
- Zhenget al\.\(2026\)P\. Zheng, J\. Yao, J\. Zheng, C\. Gu, G\. He, J\. Liu, Y\. Huang, T\. Guo, and W\. LuFrom isolated scoring to collaborative ranking: a comparison\-native framework for llm\-based paper evaluation\.arXiv preprint arXiv:2603\.17588\.Cited by:[§2\.2](https://arxiv.org/html/2609.11234#S2.SS2.p2.1)\.
- Zhouet al\.\(2024\)R\. Zhou, L\. Chen, and K\. YuIs LLM a reliable reviewer? a comprehensive evaluation of LLM on automatic paper reviewing tasks\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\),N\. Calzolari, M\. Kan, V\. Hoste, A\. Lenci, S\. Sakti, and N\. Xue \(Eds\.\),Torino, Italia,pp\. 9340–9351\.External Links:[Link](https://aclanthology.org/2024.lrec-main.816/)Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.11234#S2.SS1.p1.1)\.
- Zhuet al\.\(2025\)M\. Zhu, Y\. Weng, L\. Yang, and Y\. ZhangDeepReview: improving LLM\-based paper review with human\-like deep thinking process\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 29330–29355\.External Links:[Link](https://aclanthology.org/2025.acl-long.1420/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1420),ISBN 979\-8\-89176\-251\-0Cited by:[§1](https://arxiv.org/html/2609.11234#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.11234#S2.SS1.p1.1)\.
## Appendix AUse of AI Assistance
Generative AI tools were used to assist with the preparation of this manuscript, including language polishing, organization of draft material, LaTeX editing, and consistency checks\. AI tools were not used as authors and did not determine the research questions, benchmark design, experimental results, or scientific claims\. The authors reviewed, edited, and verified all AI\-assisted output and take full responsibility for the content of the paper\.
## Appendix BModel and Inference Details
### B\.1Model Coverage
We evaluate 18 LLMs in three groups\. The frontier group contains 12 systems: GPT\-5\.5, GPT\-5\.4\-mini, Claude Opus 4\.7, Claude Sonnet 4\.6, Gemini 3\.1 Pro, Gemini 3\.1 Flash\-Lite, Kimi\-K2\.6, GLM\-5\.1, Qwen3\.6\-Plus, DeepSeek\-V4\-Pro, DeepSeek\-V4\-Flash, and MiniMax\-M2\.7\. The mid\-scale open\-weight group contains Qwen3\-8B, Qwen3\-32B, Qwen3\.5\-27B, and Gemma4\-31B\-IT\. The specialised peer\-review group contains CycleReviewer\-8B, DeepReviewer\-7B, DeepReviewer\-14B, and LLaMA\-OpenReviewer\-8B\.
DeepReviewer\-7B and LLaMA\-OpenReviewer\-8B failed to produce valid structured outputs and are excluded from quantitative tables\. Appendix[L](https://arxiv.org/html/2609.11234#A12)reports the failure cases\.
### B\.2Inference Settings
All models are run with their maximum reasoning setting when the API or local runtime exposes one\. Pairwise judgment is evaluated for all 18 models\. Multi\-paper grouping is restricted to the 12 frontier models because the task requires substantially longer context windows\.
Papers are rendered at three granularities: abstract only, abstract plus introduction, and full paper\. The corresponding token caps are 1k, 4k, and 16k per paper\. Each input and dimension is decoded three times at temperature 0\.6\. Per\-family prompt and configuration details appear in Appendix[P](https://arxiv.org/html/2609.11234#A16)\.
## Appendix CMetric Definitions
This section provides formal definitions for the three cascading metrics introduced in Section[3\.4](https://arxiv.org/html/2609.11234#S3.SS4)\. All metrics are computed per dimension \(task, problem, method\) and averaged across independent samples\.
We first define the shared notation\. For a given model and dimension, let:
- •TP= true positives \(pred=similar, label=similar\),
- •FP= false positives \(pred=similar, label=dissimilar\),
- •FN= false negatives \(pred=dissimilar, label=similar\)\.
Among the true positives, the cascading evaluation partitionsTPinto three mutually exclusive subsets based on Stages 2 and 3:
- •TPh\\text\{TP\}\_\{h\}= hallucinated \(Stage 2 failed: evidence not grounded in paper text\),
- •TPm\\text\{TP\}\_\{m\}= mismatched \(Stage 2 passed but Stage 3 failed: reason not supported by evidence\),
- •TPv\\text\{TP\}\_\{v\}= verified \(both Stage 2 and Stage 3 passed\)\.
By construction,TP=TPv\+TPh\+TPm\\text\{TP\}=\\text\{TP\}\_\{v\}\+\\text\{TP\}\_\{h\}\+\\text\{TP\}\_\{m\}\.
### C\.1Hallucination Rate
Hallucination Rate measures the fraction of true\-positive predictions whose cited evidence is not grounded in the source paper text\. An evidence span is considered hallucinated if it fails both exact substring match and 4\-gram soft match \(ROUGE\-L recall≥0\.75\\geq 0\.75\) against the paper content\.
Hallucination Rate=\|TPh\|\|TP\|\\text\{Hallucination Rate\}=\\frac\{\|\\text\{TP\}\_\{h\}\|\}\{\|\\text\{TP\}\|\}
### C\.2Mismatch Rate
Among the non\-hallucinated true positives \(TPv\+TPm\\text\{TP\}\_\{v\}\+\\text\{TP\}\_\{m\}\), Mismatch Rate measures the fraction whose stated reason is judged unsupported by the cited evidence\. An external LLM judge makes this determination, using the prompt in Appendix[P](https://arxiv.org/html/2609.11234#A16)\.
Mismatch Rate=\|TPm\|\|TPv\|\+\|TPm\|\\text\{Mismatch Rate\}=\\frac\{\|\\text\{TP\}\_\{m\}\|\}\{\|\\text\{TP\}\_\{v\}\|\+\|\\text\{TP\}\_\{m\}\|\}
Tables[3](https://arxiv.org/html/2609.11234#S4.T3)and[4](https://arxiv.org/html/2609.11234#S5.T4)report Hallucination Rate and Mismatch Rate computed over true positives only\. False\-positive predictions \(pred=similar, label=dissimilar\) undergo the same two\-stage verification, their substantially higher hallucination and mismatch rates are reported separately in Figure[4](https://arxiv.org/html/2609.11234#S5.F4)\.
### C\.3Verified F1
Verified F1 measures end\-to\-end performance: a true positive counts toward the numerator only if it passes both Stage 2 and Stage 3\.
Verified Precision=TPvTPv\+TPh\+TPm\+FP\\displaystyle=\\frac\{\\text\{TP\}\_\{v\}\}\{\\text\{TP\}\_\{v\}\+\\text\{TP\}\_\{h\}\+\\text\{TP\}\_\{m\}\+\\text\{FP\}\}\(1\)Verified Recall=TPvTPv\+TPh\+TPm\+FN\\displaystyle=\\frac\{\\text\{TP\}\_\{v\}\}\{\\text\{TP\}\_\{v\}\+\\text\{TP\}\_\{h\}\+\\text\{TP\}\_\{m\}\+\\text\{FN\}\}\(2\)Verified F1 is the harmonic mean:
Verified F1=2⋅VP⋅VRVP\+VR\\text\{Verified F1\}=\\frac\{2\\cdot\\text\{VP\}\\cdot\\text\{VR\}\}\{\\text\{VP\}\+\\text\{VR\}\}where VP and VR denote Verified Precision and Verified Recall\.
Note that the denominator of Verified Precision equalsTP\+FP\\text\{TP\}\+\\text\{FP\}, identical to standard Precision\. The only difference is that the numerator counts only the verified subsetTPv\\text\{TP\}\_\{v\}rather than allTP\.
Granularity for multi\-paper grouping\.For the grouping evaluation, hallucination and mismatch checks operate at the*subgroup*level\. Each predicted subgroup is classified as verified, hallucinated, or mismatched based on its collective evidence and reasoning\. However, F1 computation operates at the*pair*level to remain comparable with pairwise evaluation\. The procedure is as follows\.
1. 1\.Each predicted subgroupGGis classified as verified, hallucinated, or mismatched\.
2. 2\.The subgroup is then converted to its set of within\-group paper pairs:pairs\(G\)=\{\(i,j\):i,j∈G,i<j\}\\text\{pairs\}\(G\)=\\\{\(i,j\):i,j\\in G,i<j\\\}\.
3. 3\.All pairs from verified subgroups contribute toTPv\\text\{TP\}\_\{v\}when they match ground\-truth pairs and toFPotherwise\. Similarly, pairs from hallucinated or mismatched subgroups contribute toTPh\\text\{TP\}\_\{h\}orTPm\\text\{TP\}\_\{m\}when correct and toFPotherwise\.
This design ensures that a single hallucinated evidence span invalidates all pairs within that subgroup, reflecting the higher reliability standard for multi\-document reasoning, while maintaining pair\-level F1 comparability with the pairwise evaluation\.
## Appendix DICLR Construction Pipeline Details
This appendix provides the complete construction pipeline for the ICLR review\-grounded subset, including extraction rules, verification procedures, tiering criteria, and quality controls\.
### D\.1Candidate Extraction Rules
Novelty\-overlap claim definition\.A candidate novelty\-overlap claim is a statement in an official review or author rebuttal that satisfies both conditions below\.
- •A reviewer asserts that the submission’s task, problem, or method is already covered by a specific prior work
- •The authors explicitly concede such a claim in their rebuttal\.
The extraction step uses GPT\-5\.5\. We prompt it to operate at high recall and capture any statement that could plausibly indicate overlap, even if the reviewer’s phrasing is indirect or the concern is later refuted by the authors\.
Extraction unit\.Each candidate is represented as a pair consisting of a submission and a cited prior work\. The fields are listed below\.
- •Thesubmissionfield contains the ICLR forum ID and year\.
- •Thecited\_prior\_workfield contains the paper title, authors, and venue as mentioned in the review or rebuttal\.
- •Thereviewer\_quotefield contains the verbatim text from the review that contains the overlap claim\.
- •Theauthor\_rebuttal\_quotefield is optional\. It contains the verbatim text from the rebuttal where the authors acknowledge the concern\.
Raw extraction yields tens of thousands of noisy candidates per year, including false positives where reviewers mention prior work for comparison without claiming overlap\. GPT\-5\.5 also proposes preliminary dimension labels, and a coding agent using GPT\-5\.5 separately collects paper metadata\. The verification stage revisits the labels\.
### D\.2LLM Verification and Dimension Correction
Verification criteria\.Claude Opus 4\.6 then verifies each candidate using the following criteria\.
1. 1\.The claim is genuinely a novelty concern, not merely a citation for background or comparison\.
2. 2\.The overlap is specific to the cited prior work, not a general statement about the field\.
3. 3\.The labels for task, problem, and method are correctly assigned based on the reviewer’s statement\.
Dimension correction\.The verification LLM may correct the initial dimension labels if the reviewer’s statement indicates a different type of overlap than initially extracted\. For example, a reviewer statement “This method is similar to \[Prior Work\]” is labeled asmethod, while “This paper addresses the same problem as \[Prior Work\]” is labeled asproblem\.
### D\.3Prior\-Work Identity Resolution
Multi\-source resolution chain\.Each candidate’s cited prior work is resolved to a stable identity through a cascading lookup across the following sources\.
1. 1\.arXiv\.Title and author matching against arXiv metadata\.
2. 2\.Semantic Scholar\.API lookup by title, with fuzzy matching for minor variations\.
3. 3\.OpenReview\.Direct forum ID if the prior work was also submitted to ICLR\.
4. 4\.Publisher pages\.DOI resolution for conference or journal papers\.
5. 5\.DBLP\.Bibliographic lookup for venue\-published papers\.
6. 6\.Manual web search\.Human fallback for ambiguous cases\.
Resolution confidence\.Each resolved prior work receives one of the following confidence levels\.
- •High\.Exact match on title and authors, or a confirmed DOI or arXiv ID\.
- •Medium\.Fuzzy title match with author overlap, or confirmation from a single source\.
- •Low\.Ambiguous match or multiple plausible candidates\.
The final test set retains only pairs withhighresolution confidence\.
### D\.4Author Concession Verification
Structured signal\.Author concession is recorded only when the author*explicitly admits the specific concern*raised by the reviewer\. Generic statements like “We thank the reviewer for the suggestion” or “We will cite this work” do not count as concession\.
Verification protocol\.The human adjudicator verifies the following conditions\.
1. 1\.The rebuttal quote directly responds to the reviewer’s overlap claim\.
2. 2\.The rebuttal acknowledges the overlap\. For example, it may state “We agree that our method shares similarities with \[Prior Work\]\.”
3. 3\.The rebuttal refers to the same prior work the reviewer named\.
### D\.5Tier Assignment Rules
Tier\-1 criteria\.A pair is assignedTier\-1only when both conditions hold\.
- •Multiple reviewers independently flag the same prior work\. At least two distinct reviewers must mention the same paper\.
- •The authors explicitly concede the concern in their rebuttal\.
Tier\-2 criteria\.A pair is assignedTier\-2when either condition holds\.
- •Only one reviewer flags the prior work\.
- •Multiple reviewers flag it but the authors do not concede\.
Pairs with neither signal are dropped from the benchmark draft\. A pair has neither signal when only one reviewer flags the prior work and the authors do not concede\.
### D\.6Human Adjudication Protocol
Partition and audit procedure\.Three PhD\-level annotators adjudicate every candidate pair\. All three have research experience in NLP or ML\. They follow the protocol below\.
1. 1\.Partition\.Each pair is assigned to one primary annotator\.
2. 2\.Label\.The primary annotator labels the pair, verifying prior\-work identity, dimension labels, and tier assignment\.
3. 3\.Audit\.The other two annotators independently audit the primary annotator’s decision\.
4. 4\.Resolve\.Disagreements are resolved through discussion among all three annotators\.
Adjudication checklist\.For each pair, the primary annotator verifies the following conditions\.
- •The cited prior work is correctly resolved to a stable identity\.
- •The reviewer quotes anchor to that specific paper\.
- •The labels for task, problem, and method are supported by the raw forum evidence\.
- •The tier assignment follows the Tier\-1 and Tier\-2 rules and reflects the number of reviewers and any author concession\.
Adjudication outcomes\.Adjudication has one of the following outcomes\.
- •Accept\.The pair is retained with the verified labels\.
- •Correct\.The pair is retained, but its dimension labels or tier are corrected\.
- •Reject\.The pair is excluded when the prior work is misidentified or the claim is not genuinely about novelty\.
Consistency audit by dimension\.For the consistency audit in Section[3\.2](https://arxiv.org/html/2609.11234#S3.SS2), we sampled 30 items with 10 from each dimension\. Three annotators independently reviewed each item, giving 90 reviews in total\. Table[5](https://arxiv.org/html/2609.11234#A4.T5)reports 93\.3% overall agreement\. Disagreements mainly involved the boundary between task and problem and differences in method granularity\. All were resolved through discussion\.
Table 5:Label consistency audit by dimension\. Three annotators independently reviewed 30 sampled items, with 10 from each dimension, giving 90 reviews in total\. Agreement is computed against the original annotations\.
### D\.7Final Selection Criteria
The released test set retains pairs that meet all of the following conditions\.
- •tier == "Tier\-1"
- •human\_adjudicated == true
- •adjudication\_source == "human"\. Agent\-adjudicated entries are excluded\.
- •overlap\.is\_valid == true
- •prior\_resolved == true
- •prior\_resolution\_confidence == "high"
- •reviewer\_evidence\.quotesis non\-empty
- •Both submission and prior work have parsed Markdown content
This strict filtering ensures that every pair in the test set is grounded in convergent human evidence from multiple reviewers and an author concession, and is anchored to a correctly identified prior work\.
### D\.8Year Distribution and Temporal Skew
Table 6:ICLR dimension labels by year\. Pairs with multi\-dimension labels appear in more than one row\. The total pair count is the unique pair count per year\.The pairs concentrate in ICLR 2026\. It contributes 189 of 236 pairs, or 80%\. This concentration reflects two*construction*factors rather than a change in reviewer behavior\.
1. 1\.ICLR 2026 contributes by far the most raw OpenReview forums to the extraction pipeline\.
2. 2\.The benchmark’s human\-review effort prioritized a test\-heavy 2026 split, while 2025 adjudication was directed mainly toward a training split that is not released with this benchmark\.
The year distribution should therefore not be read as an estimate of the true per\-year novelty\-overlap rate in ICLR submissions\.
Evaluation by year\.Because ICLR contributes positive pairs only, F1 and Verified F1 are not comparable across ICLR years\. We therefore report metrics based on positive pairs only\. The method dimension has 169 pairs from 2026 and 28 pairs from earlier years\. The problem dimension has 60 pairs from 2026 and 41 pairs from earlier years\. We omit the task dimension because it has only 3 pairs from 2026 and 1 pair from earlier years\. Table[7](https://arxiv.org/html/2609.11234#A4.T7)averages the metrics over input settings and evaluated models\. The values are similar across years\.
Table 7:ICLR metrics for positive pairs, stratified by year\. Each row reports one dimension and year group\. Values are percentages averaged over input settings and models\. The task dimension is omitted because it has only 3 pairs from 2026 and 1 pair from earlier years\.We also compared the full ICLR model ranking with the ranking on data from earlier years\. Table[8](https://arxiv.org/html/2609.11234#A4.T8)shows Spearman correlations of 0\.87–0\.98 across input settings and metrics\. These values indicate stable rankings after removing 2026\.
Table 8:Spearman correlation between the full ICLR ranking and the pre\-2026 ranking\. Values are reported for each input setting and metric\.
### D\.9Construction Funnel Summary
Table 9:ICLR construction funnel from raw forums to final test set\.The 236 released pairs are a deliberately selected, content\-available benchmark subset, not an exhaustive set of all novelty\-overlap concerns in ICLR\. This funnel reflects high\-precision filtering for evaluation rather than prevalence estimation\.
## Appendix ESurvey Construction Pipeline Details
This appendix provides the complete construction pipeline for the survey co\-citation subset, including survey filtering rules, sentence extraction patterns, annotation protocols, and cross\-group sampling constraints\.
### E\.1Survey Collection Rules
Precision\-oriented filter\.A paper qualifies as a candidate survey when one of the following conditions holds\.
- •Its title containssurvey,review, ortutorial\.
- •Its abstract opens with an explicit survey or review declaration\. For example, it may begin “This survey reviews…” or “We present a comprehensive review of…”\.
False positive exclusion\.The filter excludes papers whose title contains any of the following phrases\.
- •peer review, as in “Peer Review Analysis”
- •code review, as in “Automated Code Review”
- •literature reviewin a non\-survey context, as in “A New Method with Literature Review”
This precision\-first design prioritizes high\-quality survey papers over exhaustive coverage\.
### E\.2Sentence Extraction Rules
Co\-citation requirement\.A sentence qualifies as a candidate if it cites two or more papers\. Citations are identified in standard forms such as\[Author et al\., Year\],\(Author et al\., Year\),\\cite\{\.\.\.\}, and numbered references such as\[1, 2, 3\]\.
Similarity cues\.The sentence must contain at least one of the following explicit similarity cues\.
- •similar, orshare the same task, problem, method, or approach
- •address the same problem,follow the same approach
- •both papers tackle,fall into the same, orbelong to the same
- •group of methods,family of approaches
Anti\-patterns\.The extractor excludes sentences containing the following technical uses of “similarity”\.
- •cosine similarity,semantic similarity score,similarity metric
- •similarity function,similarity measure,similarity threshold
- •Jaccard similarity,edit distance similarity
This ensures that only sentences asserting paper\-level similarity are kept, not sentences discussing similarity as a technical concept\.
### E\.3LLM Pre\-screening Protocol
The survey pre\-screening model is Gemini 3\.1 Pro, the only LLM in this pipeline\. It labels sentences selected by the rule\-based extractor, and human annotators confirm or revise each grouping\.
Input format\.For each candidate sentence, the LLM receives the following information\.
- •The co\-citation sentence \(verbatim\)
- •The list of cited papers \(titles and authors\)
- •The surrounding context \(previous and next sentences\)
Output format\.The LLM proposes a three\-dimensional similarity grouping in the following JSON format\.
```
{
"task": [[p1, p2], [p3]],
"problem": [[p1, p3]],
"method": [[p2, p3]]
}
```
Each dimension contains zero or more similarity groups\. Papers not grouped with any other paper are omitted\.
### E\.4Human Verification Protocol
Two\-outcome protocol\.For each candidate sentence, the assigned annotator performs the following steps\.
1. 1\.Reads the co\-citation sentence, the grouping proposed by the LLM, and the cited papers’ full text alongside one another\.
2. 2\.Decides whether toacceptthe LLM labels orrevisethem\.
Explicit revision\.If any dimension label is judged incorrect or incomplete, the annotator writes an explicit correction in thehuman\_revisefield\. The correction specifies the following\.
- •Which papers are similar on which dimension
- •Which papers are*not*similar when the LLM incorrectly grouped them
Implicit approval\.If no correction is needed, the annotator records a non\-emptyhuman\_judgmententry such as “Verified”\. This signals that all three LLM labels are accepted\.
Audit\.The other two annotators independently audit each decision\. Disagreements are resolved through discussion\.
Final label merge rule\.The labels for each verified sentence are merged as follows\.
- •Ifhuman\_reviseis non\-empty, its per\-dimension labels replace the LLM output in full\.
- •Otherwise, thellm\_judgmentlabels are adopted as\-is\.
### E\.5Cross\-Group Sampling Constraints
Dimension\-specific sampling\.For each dimensiond∈\{task,problem,method\}d\\in\\\{\\text\{task\},\\text\{problem\},\\text\{method\}\\\}, we sample negative pairs only from sentences where the annotators verified at least one similarity group in dimensiondd\.
Sampling procedure\.For each verified sentence withk≥2k\\geq 2similarity groups in dimensiondd, we do the following\.
1. 1\.Enumerate all pairs\(pi,pj\)\(p\_\{i\},p\_\{j\}\)wherepip\_\{i\}is in groupg1g\_\{1\}andpjp\_\{j\}is in groupg2≠g1g\_\{2\}\\neq g\_\{1\}\.
2. 2\.Label each such pair asnot\_similarin dimensiondd\.
Conflict handling\.If a cross\-group pair\(pi,pj\)\(p\_\{i\},p\_\{j\}\)appears as a positive in dimensionddelsewhere in the survey corpus, we exclude it from the negative set\. This occurs when the same two papers are grouped together in another sentence\.
Negatives in multiple dimensions\.A single pair may carry negative labels in multiple dimensions if the survey author placed the two papers into distinct groups across multiple dimensions\. 56 of the 156 negative pairs carry labels in more than one dimension\.
### E\.6Per\-Survey Contribution Distribution
Table 10:Per\-survey contribution to the 118 verified sentences\. The top\-10 contributing surveys account for 54% of all sentences\.The per\-survey skew reflects both the precision\-first design of the extractor, which favors surveys with explicit similarity statements, and the natural variation in how surveys structure their related\-work prose\. We treat the resulting concentration as a known limitation\.
## Appendix FEvidence Verification Technical Details
This appendix provides the complete technical specification for the hallucination check in Stage 2 of the cascading evaluation protocol\.
### F\.1Normalization Rules
Before exact substring matching, both the cited evidence span and the source paper text undergo the following normalizations:
Ligature unification\.Unicode ligatures are expanded to their component characters:
- •fi\(U\+FB01\)→\\rightarrowfi
- •fl\(U\+FB02\)→\\rightarrowfl
- •ff\(U\+FB00\)→\\rightarrowff
- •ffi\(U\+FB03\)→\\rightarrowffi
- •ffl\(U\+FB04\)→\\rightarrowffl
Typographic quote and dash normalization\.Typographic punctuation is normalized to ASCII equivalents:
- •"\(U\+201C\),"\(U\+201D\)→\\rightarrow"
- •’\(U\+2018\),’\(U\+2019\)→\\rightarrow’
- •–\(U\+2013, en\-dash\),—\(U\+2014, em\-dash\)→\\rightarrow\-
Soft\-hyphen removal\.Soft hyphens \(U\+00AD\) are removed entirely\.
PDF line\-wrap hyphenation handling\.Hyphens at line breaks are conditionally removed:
- •If a word ends with\-\\nand the next line starts with a lowercase letter, the hyphen and newline are removed\.
- •Otherwise, the hyphen is retained\.
Whitespace collapse\.All sequences of whitespace \(spaces, tabs, newlines\) are collapsed to a single space\.
### F\.2Character\-Level 4\-Gram Recall
For spans that fail exact substring matching, we compute character\-level 4\-gram recall as follows:
4\-gram extraction\.For a stringssof lengthnn, the set of character 4\-grams is:
4grams\(s\)=\{s\[i:i\+4\]∣0≤i≤n−4\}\\text\{4grams\}\(s\)=\\\{s\[i:i\{\+\}4\]\\mid 0\\leq i\\leq n\{\-\}4\\\}
Recall computation\.For a cited evidence spaneeand source paper textpp:
4gram\-recall\(e,p\)=\|4grams\(e\)∩4grams\(p\)\|\|4grams\(e\)\|\\text\{4gram\-recall\}\(e,p\)=\\frac\{\|\\text\{4grams\}\(e\)\\cap\\text\{4grams\}\(p\)\|\}\{\|\\text\{4grams\}\(e\)\|\}
Soft match threshold\.A sentence is admitted as a soft match if its 4\-gram recall is at least 0\.75\.
Rationale\.This soft match catches PDF parsing artifacts \(e\.g\.,‘‘per forms’’matching‘‘performs’’\) and benign rewordings \(e\.g\., punctuation changes, minor typos\) without admitting unrelated content\. The 0\.75 threshold was chosen empirically to balance precision and recall on a held\-out validation set of 50 manually labeled evidence spans\.
Table 11:Per\-model raw F1 and Verified F1 with cross\-sample standard deviations for pairwise judgment, full content granularity\. Values are mean±\\pmstd across then=3n\{=\}3inference samples \(same samples used to compute Table[3](https://arxiv.org/html/2609.11234#S4.T3)\)\.‡GPT\-family judged by Gemini\-3\.1\-Pro\.
## Appendix GDataset Schema and Samples
The benchmark comprises three dataset types corresponding to the two evaluation tasks described in §[3\.3](https://arxiv.org/html/2609.11234#S3.SS3)\.
### G\.1Pairwise Positive Pairs
Confirmed positive pairs from ICLR submissions citing prior work\. Each pair includes the submitted paper, the cited prior work, and human\-annotated overlap dimensions\.
Key fields:
- •submitted\_paper,prior\_work: paper metadata \(title, abstract, etc\.\)
- •similar\_dimensions:\["task"\],\["problem"\],\["method"\], or combinations
- •overlap: human\-written overlap summary
### G\.2Cross\-Group Negative Pairs
Negative pairs from papers in different groups within survey sentences\. Used to construct hard negatives where papers share context but differ on specific dimensions\.
Key fields:
- •paper\_a,paper\_b: paper metadata
- •negative\_dimension: the dimension on which they differ \("task","problem", or"method"\)
- •labels: per\-dimension binary labels \(0 = dissimilar\)
Table 12:Content granularity ablation\.Metrics averaged over the task, problem, and method dimensions\.F1measures binary prediction accuracy\.Haldenotes Hallucination Rate,Misdenotes Mismatch Rate, andVF1denotes Verified F1\. All metrics are percentages\. The columns are*abs*for abstract only,*a\+i*for abstract and intro, and*full*for the full paper\.‡GPT\-family models are judged by Gemini\-3\.1\-Pro\.Boldindicates the best value for each content granularity and metric\.
### G\.3Multi\-paper Grouping Tasks
Sets of 3–4 papers from survey sentences with ground\-truth groupings\. Models must partition papers by shared task, problem, or method\.
Key fields:
- •papers: list of paper metadata
- •ground\_truth:\{task\_groups, problem\_groups, method\_groups\}, each a list of paper\-index lists \(e\.g\.,\[\[1,2\], \[3\]\]means papers 1,2 form one group, paper 3 is alone\)
- •eval\_dims: dimensions to evaluate on this instance
Example structure:
\{
"papers":\[
\{"title":"\.\.\.","abstract":"\.\.\."\},
\{"title":"\.\.\.","abstract":"\.\.\."\},
\{"title":"\.\.\.","abstract":"\.\.\."\}
\],
"ground\_truth":\{
"task\_groups":\[\[0,1,2\]\],
"method\_groups":\[\[0,1\],\[2\]\]
\},
"eval\_dims":\["task","method"\]
\}
### G\.4Paper Type and Domain Coverage
We classify 677 unique papers by contribution type and scientific domain using their titles and abstracts\. Tables[13](https://arxiv.org/html/2609.11234#A7.T13)and[14](https://arxiv.org/html/2609.11234#A7.T14)report the overall and source\-specific distributions\. Method papers account for 73\.3%, while the remaining papers include analysis, benchmarks, theory, and surveys\. Computer science accounts for 86\.7% of the papers\. The survey source is more diverse than ICLR\. Computer science accounts for 66\.7% of the survey papers and 96\.5% of the ICLR papers\.
Among 619 pairs, 19\.2% span different paper types and 10\.8% span different scientific domains\. The corresponding rates are 19\.5% and 7\.2% for ICLR, and 19\.1% and 13\.1% for the survey source\.
Table 13:Contribution\-type distribution over 677 unique papers\. The\#column gives the count\. The other columns are percentages overall and within each source\.Table 14:Domain distribution over 677 unique papers\. The\#column gives the count\. The other columns are percentages overall and within each source\.
## Appendix HDetailed Per\-Sample Results
This section provides per\-sample variability estimates for the main pairwise evaluation results\. Table[11](https://arxiv.org/html/2609.11234#A6.T11)reports raw F1 and Verified F1 with cross\-sample standard deviations for thefullcontent granularity, the default setting in Table[3](https://arxiv.org/html/2609.11234#S4.T3)\. Each standard deviation is computed across three independent inference runs at temperature 0\.6 on the same test set, reflecting output variability due to sampling randomness\.
Hallucination Rate and Mismatch Rate standard deviations are omitted for space and are available in the released per\-sample raw counts\. The reported standard deviations are typically small, at 0\.5–3\.0 percentage points, confirming that the benchmark’s conclusions are robust to sample variation\.
Dimension coverage and accuracy by source\.Table[15](https://arxiv.org/html/2609.11234#A8.T15)gives per\-dimension label counts, and Table[16](https://arxiv.org/html/2609.11234#A8.T16)separates accuracy by source\. Since ICLR contributes positives only, we compare both labels within the survey source\. Positive accuracy is 44–59% versus 90–97% for negatives\. ICLR positive accuracy is similar on the shared dimensions\.
Table 15:Dimension\-label counts\. Counts are over dimension\-specific evaluation instances rather than paper\-pair records\. The 619 paper\-pair records yield 875 pairwise evaluation instances, including 649 positive and 226 negative labels\. The benchmark contains 463 positive paper\-pair records, including 236 from ICLR and 227 from the survey source, and 156 negative paper\-pair records\. All negative pairs come from the survey source\. Fifty\-six negative paper\-pair records have multiple labels\.Table 16:Per\-sample accuracy by source, averaged over evaluated models and reported as percentages\. ICLR contributes positive pairs only\. Its task dimension is omitted because only four ICLR task pairs are available\.
## Appendix IContent Granularity Ablation \(Full Results\)
This section provides the complete content granularity ablation results for all evaluated models across three content modes:abstract only\(abs\),abstract \+ introduction\(a\+i\), andfull paper\(full\)\.
Table[12](https://arxiv.org/html/2609.11234#A7.T12)reports macro\-averaged metrics over the three dimensions \(task, problem, method\)\. The main paper \(Table[3](https://arxiv.org/html/2609.11234#S4.T3)\) uses the full\-paper setting as default\. Figure[6](https://arxiv.org/html/2609.11234#S5.F6)in the main text highlights five representative models\.
Figure 7:Rank drift between pairwise and grouping tasks for the 12 Frontier models\. Top\-4 is stable, shown with teal lines, but the mid\-pack reorders substantially, shown with grey lines\. GLM\-5\.1 drops 6 ranks and DeepSeek\-V4\-Pro jumps 4 ranks\. Red lines highlight major drifts with\|Δ\|≥4\|\\Delta\|\\geq 4\.#### Key observations\.
- •F1 generally improves with more content: Most models show higher prediction accuracy when given full papers compared to abstracts alone\.
- •Hallucination Rate remains low: Across all granularities, most frontier models maintain hallucination rates below 2%, indicating that evidence extraction is generally reliable\.
- •Mismatch Rate increases substantially with full\-paper content: Models struggle more with nuanced reasoning when given complete papers, with mismatch rates rising from 38–76% \(abstract\) to 24–90% \(full paper\)\. This suggests that while models can extract correct evidence from longer texts, they have difficulty maintaining reasoning quality\.
## Appendix JRank Drift: Pairwise vs Grouping
The pairwise and grouping tasks described in §[4\.1](https://arxiv.org/html/2609.11234#S4.SS1)and §[4\.3](https://arxiv.org/html/2609.11234#S4.SS3)share the same 12 Frontier models and the same cascading metric, so their Verified F1 numbers are directly comparable\. Figure[7](https://arxiv.org/html/2609.11234#A9.F7)shows each model’s rank on both evaluations\.
The top of the leaderboard is stable: GPT\-5\.5, GPT\-5\.4\-mini, Gemini 3\.1 Pro and Flash\-Lite hold the top four positions on both evaluations\. The middle of the leaderboard is not: GLM\-5\.1 drops from 5th to 11th \(−6\-6ranks\), Claude Opus 4\.7 collapses to last place \(VF1 = 2\.6%\), while two DeepSeek\-V4 variants jump four ranks in the opposite direction\. Multi\-paper grouping separates models that sustain grounded reasoning across documents from those that rely on pairwise shortcuts\.
## Appendix KEvidence Judge Validation
Table 17:Human validation of the LLM\-based mismatch judge on 100 stratified non\-hallucinated positive outputs\. The judge agrees with human labels in 87\.0% of cases, with Cohen’sκ=0\.74\\kappa=0\.74\.Table 18:Verified F1 of GPT\-family models under the GPT\-5\.2 self\-judge and the third\-party Gemini\-3\.1\-Pro judge\. The larger method gap indicates self\-preference in the default judge\.### K\.1Human Validation of Mismatch Judgments
To validate the reliability of the LLM mismatch judge, we manually annotated a stratified sample of 100 non\-hallucinated positive outputs\. The sample was balanced across pairwise and grouping tasks, the task, problem, and method dimensions, and decisions that the judge supported or did not support\. Human annotators judged whether the cited evidence logically supported the model’s stated reason\.
Table[17](https://arxiv.org/html/2609.11234#A11.T17)shows that the LLM judge agrees with human labels in 87\.0% of cases \(95% Wilson CI: 79\.0–92\.2\), with Cohen’sκ=0\.74\\kappa=0\.74\. For identifying unsupported reasoning, the judge achieves 78\.8% precision, 95\.3% recall, and 86\.3% F1\. Only 2 of 48 judge\-supported cases were marked unsupported by humans, indicating that cases passing the judge are rarely false passes under human assessment\. The judge is therefore reliable for detecting unsupported reasoning, while tending to err conservatively by marking some human\-supported cases as unsupported\.
### K\.2Evidence of GPT Self\-Preference
To validate the fairness of our judge routing strategy, we conducted a controlled experiment with GPT\-5\.2 as the default judge and Gemini\-3\.1\-Pro as a third\-party judge\. We applied both judges to GPT\-5\.5 and GPT\-5\.4\-mini\.
We reran the reasoning check for evidence and reason mismatches with both judges and report the resulting Verified F1 in Table[18](https://arxiv.org/html/2609.11234#A11.T18)\. Hallucination Rate measures verbatim evidence and does not depend on the judge, so it is unchanged across the two runs\.
The results show that on task and problem dimensions, the two judges agree within±2\\pm 2percentage points Verified F1\. However, on the method dimension, where evidence is most nuanced, the GPT\-5\.2 self\-judge produces Verified F1 that is 7\.3 to 20\.6 percentage points higher than the third\-party Gemini\-3\.1\-Pro judge\. This one\-sided inflation is what we mean by*self\-preference bias*\.
This finding motivates the per\-model judge routing used throughout the paper\. GPT\-family models are judged by Gemini\-3\.1\-Pro, as shown in Tables[3](https://arxiv.org/html/2609.11234#S4.T3)and[4](https://arxiv.org/html/2609.11234#S5.T4)\. All other models are judged by GPT\-5\.2\.
### K\.3Consistency Across Judges
To test judge dependence, we reran the mismatch judgment with GPT\-5\.2, Gemini\-3\.1\-Pro, and DeepSeek\-V4\-Pro on Claude Sonnet 4\.6 outputs under all three input settings\. Each setting contains 1128–1237 records\. Pairwise agreement is 79–83% withκ=0\.52\\kappa=0\.52–0\.620\.62\. Agreement among all three judges is 66–76%, with the lowest value for method, as shown in Tablesand[19](https://arxiv.org/html/2609.11234#A11.T19)\.
On the 60 contested records, GPT\-5\.2 reaches 80% accuracy, compared with 55% for majority vote, 43% for Gemini\-3\.1\-Pro, and 32% for DeepSeek\-V4\-Pro as shown in Table[20](https://arxiv.org/html/2609.11234#A11.T20)\. Majority voting performs worse because the two lenient judges make correlated errors\. These accuracies on contested cases are not comparable to the 87\.0% balanced human\-validation result above, which uses a different sample\.
Table 19:Percentage of records with the same verdict from all three judges, by dimension and input setting\.Table 20:Accuracy against human labels on the 60 contested records where the judges disagree\.
## Appendix LFailed Specialised Models
Two specialised peer\-review models could not be evaluated under our structured\-output protocol:
DeepReviewer\-7B\.This model disregarded the JSON output format specified in the prompt and instead emitted full review\-template prose in natural language\. None of its outputs could be parsed as the required JSON structure containingis\_similar,evidence\_a,evidence\_b, andreasonfields\. Manual inspection of 50 outputs confirmed that the model had lost the ability to follow structured\-output instructions after task\-specific supervised fine\-tuning on peer\-review generation\.
LLaMA\-OpenReviewer\-8B\.This model produced syntactically valid JSON but exhibited severe degeneration in its predictions\. In its completed positive\-pair run, covering 649 dimension\-specific instances, it predictedis\_similar=falsefor 647 of 649 instances and copied the prompt’s JSON schema verbatim into the evidence fields \(e\.g\.,evidence\_a: "copy\-paste the exact span of text from Paper A\.\.\."\)\. Its available negative\-pair run contains only 60 records and is therefore incomplete\. The model appeared to have memorized the prompt template during fine\-tuning but lost the semantic understanding required to instantiate it with actual paper content\.
Both failures suggest that task\-specific SFT for review generation may erode the general instruction\-following and structured\-output discipline needed for cascading evaluation protocols\.
## Appendix MCase Study: Right Label, Wrong Reason
The cliff of §[5\.1](https://arxiv.org/html/2609.11234#S5.SS1.SSS0.Px1)is most legible at the level of a single pair\. We illustrate with a method\-dimension case from Claude Sonnet 4\.6 \(full transcript in Appendix[N](https://arxiv.org/html/2609.11234#A14)\)\.
Diagnosis\.On the pair*IndicSuperTokenizer*↔\\leftrightarrow*SuperBPE*, the model predictsmethod = similar\(correct\) and cites verbatim evidence from both papers \(Stage 2 passes\)\. However, Evidence A states only a high\-level contribution \(“two\-stage subword–superword learning”\) while Evidence B provides implementation specifics \(whitespace enforcement, lifted restriction, bridging word boundaries\)\. The model’s reason silently transfers Paper\-B specifics onto Paper A without textual support, and the cascade correctly flags this asmismatched\.
The final label happens to be correct because IndicSuperTokenizer does adopt the SuperBPE curriculum, but the cascade reveals that the model reached the right answer through unreliable reasoning\. This is the same contribution\-implementation gap identified in §[5\.1](https://arxiv.org/html/2609.11234#S5.SS1.SSS0.Px1), observed in 72\.5% of mismatch cases\. An F1\-only metric would miss this failure entirely\.
## Appendix NIllustrative Examples
This section provides two contrasting examples to illustrate the cascading verification pipeline:Case Ademonstrates a verified prediction where all three stages pass, whileCase Bshows a mismatched prediction that fails Stage 3 \(reasoning verification\) despite passing Stages 1 and 2\.
Case A: Verified Prediction \(GPT\-5\.5\)Paper pair: Paper A:*Forget Unlearning: Towards True Data\-Deletion in Machine Learning* Paper B:*Adaptive Machine Unlearning*Dimension:methodGold label:similarPrediction:similar \(✓ correct\)Model output \(sample 1 of 3\):Evidence A\.“We also provide a reduction theorem that if a mechanism is differentially private with respect to the remaining records and satisfies a data\-deletion guarantee under non\-adaptive edit requests, then it also satisfies a data\-deletion guarantee under adaptive requests\. Based on this reduction, we redefine the problem of data\-deletion as designing a mechanism that \(1\.\) satisfies a data\-deletion guarantee against non\-adaptive deletion requests, \(2\.\) is differentially private for remaining records, and \(3\.\) has the same utility guarantee as retraining under identical differential privacy constraints\.”Evidence B\.“We give a generic reduction\. We show that if:\(1\) A data deletion algorithmRAR\_\{A\}for a learning algorithmAAhas deletion guarantees for oblivious sequences of deletion requests \(as those from past work do\), and\(2\) Information about the internal randomness ofRAR\_\{A\}is revealed only in a manner that satisfies differential privacy, then\(A,RA\)\(A,R\_\{A\}\)also satisfies data deletion guarantees against an adaptive sequence of deletion requests, that can depend in arbitrary ways on the information that the model provider has made public\.”Reason\.“Both propose a reduction where a deletion/unlearning guarantee for non\-adaptive or oblivious request sequences, together with differential privacy, yields a guarantee for adaptive deletion/update requests\.”Cascading verification:Stage 1 \(Correctness\):✓\\checkmarkPass— prediction matches gold label \(method\-similar\)Stage 2 \(Hallucination\):✓\\checkmarkPass— both evidence strings are verbatim substrings of source papersStage 3 \(Reasoning\):✓\\checkmarkPass— judge critique: “Both evidence strings describe the same specific reduction mechanism where differential privacy combined with non\-adaptive \(or oblivious\) deletion guarantees implies adaptive deletion guarantees\.”Final verdict:Verified— This is the rare case where the model reaches*the right answer for the right reason*\. All three samples for this pair passed all verification stages\.
Case B: Mismatched Prediction \(Claude Sonnet 4\.6\)Paper pair: Paper A:*IndicSuperTokenizer: An Optimized Tokenizer for Indic Multilingual LLMs* Paper B:*SuperBPE: Space Travel for Language Models*\(Liu et al\., 2025\)Dimension:methodGold label:similarPrediction:similar \(✓ correct\)Model output \(sample 1 of 3\):Evidence A\.“our tokenizer combines linguistically grounded pre\-tokenization with a two\-stage subword–superword learning process \(Liu et al\., 2025b\), yielding a more compact and semantically faithful vocabulary\.”Evidence B\.“SuperBPE introduces a simple intervention in the pretokenization step, separating tokenizer training into two discrete phases, wherein the tokenizer \(1\) first learns subwords \(by using pretokenization to prevent merges across whitespace\) and then \(2\) learns superwords \(by lifting this restriction\)\.”Reason\.“Both papers share the same core mechanism: a two\-stage BPE curriculum that first learns subword tokens \(with whitespace pretokenization enforced\) and then learns superword tokens that bridge word boundaries \(by lifting that restriction\), producing a vocabulary of both subword and multi\-word tokens\.”Cascading verification:Stage 1 \(Correctness\):✓\\checkmarkPass— prediction matches gold label \(method\-similar\)Stage 2 \(Hallucination\):✓\\checkmarkPass— both evidence strings are verbatim substringsStage 3 \(Reasoning\):×\\timesFail— judge critique: “Evidence A mentions a ‘two\-stage subword–superword learning process’ but does not describe the specific mechanism of enforcing and then lifting whitespace restrictions, nor does it mention BPE, both of which the Reason attributes to Paper A\.”Final verdict:Mismatched— Evidence A only states a generic “two\-stage subword–superword learning process” and cites SuperBPE \(Liu et al\., 2025b\)\. Evidence B provides the specific BPE\-curriculum details: whitespace pretokenization enforced in phase 1, then lifted in phase 2 to enable cross\-word merges\. The Reason silently transfers all of Paper B’s specifics \(“whitespace pretokenization enforced,” “lifting that restriction,” “bridge word boundaries”\) onto Paper A without textual support\.Failure type:*Specificity transfer*— The model correctly identifies that Paper A cites Paper B, but assumes the implementation details carry over without verifying that Evidence A actually describes them\. Sample 2 of 3 passed Stage 3 because Evidence B was more detailed, but samples 1 and 3 failed\.
## Appendix OCascading Failure Case Studies
This appendix provides detailed case studies illustrating failures at each stage of the cascading verification pipeline:Stage 2 \(hallucination detection\)andStage 3 \(reasoning mismatch detection\)\. Each case includes the model’s output, the source text for verification, and the judge’s determination\.
### O\.1Stage 2 Failures: Hallucination
Case H1: Cross\-Paper Evidence Contamination \(DeepReviewer\-14B\)Paper pair: Paper A:*Beyond Pass@k: Breadth\-Depth Metrics for Reasoning Boundaries* Paper B:*Are Your LLMs Capable of Stable Reasoning?*Dimension:problemGold label:similarPrediction:similar \(✓ correct\)Content granularity:abstract onlyModel output:Evidence A\(verbatim\): “To assess the reasoning boundary \(the fraction of problems a model can solve\) researchers often report Pass@k at large sampling budgets\.”Evidence B\(hallucinated\): “To assess the reasoning boundary \(the fraction of problems a model can solve\) researchers often report Pass@k at large sampling budgets\.”Reason:“Both papers identify the limitation of using Pass@k at large sampling budgets as a metric for assessing the reasoning boundary of large language models\.”Cascading verification:Stage 1 \(Correctness\):✓\\checkmarkPass— prediction matches gold label \(similar\)Stage 2 \(Hallucination\):×\\timesFail— Evidence B is*character\-for\-character identical*to Evidence A \(cls\_b: hallucination\)Failure type:*Cross\-paper evidence contamination*— The model copied Paper A’s text verbatim in place of extracting evidence from Paper B\.
Case H2: Citation Misattribution as Evidence \(DeepSeek\-V4 Flash\)Paper pair: Paper A:*QWEN Technical Report* Paper B:*MobileLLM: Optimizing Sub\-billion Parameter Language Models for On\-Device Use Cases*Dimension:taskGold label:not similarPrediction:similar \(×\\timesincorrect\)Content granularity:abstract onlyModel output:Evidence A\(hallucinated\): “We present Open Pre\-trained Transformers \(OPT\), a suite of decoder\-only pre\-trained transformers ranging from 125M to 175B parameters, which we aim to fully and responsibly share with interested researchers\.”Evidence B\(hallucinated\): “In this work, we introduce QWEN, the first installment of our large language model series\. QWEN is a comprehensive language model series that encompasses distinct models with varying parameter counts\.”Cascading verification:Stage 1 \(Correctness\):×\\timesFail— prediction does not match gold label \(predicted similar, actually not similar\)Stage 2 \(Hallucination\):×\\timesFail— Evidence A classified as hallucination \(cls\_a: hallucination\)\. Evidence B classified as hallucination \(cls\_b: hallucination\)\.Failure type:*Citation misattribution*— Evidence A describes OPT, which appears in Paper A’s reference section, not as the paper’s own contribution\.
Case H3: Fabricated Title as Evidence \(DeepSeek\-V4 Flash\)Paper pair: Paper A:*Certain Generalizations in the Analysis of Variance* Paper B:*A Generalization of Fisher’s z Test*Dimension:taskGold label:not similarPrediction:similar \(×\\timesincorrect\)Content granularity:abstract onlyModel output:Evidence A\(hallucinated\): “Some New Test Criteria in Multivariate Analysis”Evidence B\(verbatim\): “Wilks \(1932\) defined a generalized variance and found the appropriate A\-criteria for testing certain hypotheses concerning the means, variances, and covariances of k normal multivariate populations from which k independent samples have been drawn\.”Reason:“Both papers target the task of generalizing classical univariate statistical tests \(analysis of variance, Fisher’s z test\) to multivariate populations, focusing on variance components and hypothesis testing for multiple populations\.”Cascading verification:Stage 1 \(Correctness\):×\\timesFail— prediction does not match gold label \(predicted similar, actually not similar\)Stage 2 \(Hallucination\):×\\timesFail— Evidence A classified as hallucination \(cls\_a: hallucination\)\. The phrase “Some New Test Criteria in Multivariate Analysis” does not appear in Paper A\.Failure type:*Fabricated content*— Evidence A is a plausible\-sounding but non\-existent phrase\.
### O\.2Stage 3 Failures: Reasoning Mismatch
Case M1: Evidence Description Error \(minimax\-m2\.7\)Paper pair: Paper A:*Iteratively Refined Regularization for Nash Equilibria* Paper B:*Adaptively Perturbed Mirror Descent for Learning in Games*Dimension:methodGold label:similarPrediction:similar \(✓ correct\)Content mode:full paperModel output:Evidence A\(verbatim\): “Instead, we fix the regularization strength at a large value for robustness and achieve convergence by iteratively refining the reference policy\.”Evidence B\(verbatim\): “In response, we propose Adaptively Perturbed MD \(APMD\), which adjusts the magnitude of the perturbation by repeatedly updating the slingshot strategy at a predefined interval\.”Reason:“Both papers share the fundamental mechanism of keeping regularization/perturbation strength constant while iteratively updating a reference/slingshot strategy to progressively approach Nash equilibrium\. Paper A explicitly states this iterative refinement of reference policy with fixed regularization strength, and Paper B describes updating the slingshot strategy at intervals to adjust perturbation magnitude—both represent the same algorithmic backbone of constant\-strength regularization with iterative anchor updates in Mirror Descent\.”Cascading verification:Stage 1 \(Correctness\):✓\\checkmarkPassStage 2 \(Hallucination\):✓\\checkmarkPassStage 3 \(Reasoning\):×\\timesFail— supported: falseJudge critique:“Evidence A mentions fixed regularization strength and iterative refinement of a reference policy, but Evidence B explicitly says the method*adjusts*perturbation magnitude \(not keeping it constant\) and neither excerpt mentions Nash equilibrium, Mirror Descent, or an algorithmic backbone of constant\-strength regularization with anchor updates, so the stated methodological similarity is not supported by the provided text\.”Failure type:*Evidence description error*— The model misrepresents Evidence B by claiming it keeps perturbation strength constant, when it explicitly states the method “adjusts the magnitude of the perturbation\.”
Case M2: Reason Exceeds Evidence Support \(Claude Opus 4\.7\)Paper pair: Paper A:*Can Large Language Models Explore In\-Context?* Paper B:*Large Language Models as Optimizers*Dimension:taskGold label:similarPrediction:similar \(✓ correct\)Content mode:full paperModel output:Evidence A\(verbatim\): “We evaluate the ability of the current generation of large language models \(LLMs\) to help a decision\-making agent facing an exploration\-exploitation tradeoff\. We use LLMs to explore and exploit in silos in various \(contextual\) bandit tasks\.”Evidence B\(verbatim\): “We investigate the extent to which contemporary Large Language Models \(LLMs\) can engage in exploration, a core capability in reinforcement learning and decision making\. We focus on native performance of existing LLMs, without training interventions\. We deploy LLMs as agents in simple multi\-armed bandit environments, specifying the environment description and interaction history entirely in\-context, i\.e\., within the LLM prompt\.”Reason:“Both papers target the same task of evaluating LLMs’ ability to perform exploration/exploitation in \(contextual\) bandit settings via in\-context prompting, deploying LLMs as decision\-making agents in MAB environments without specialized training interventions\. A researcher would file both under the same sub\-topic of LLM\-based in\-context bandit learning\.”Cascading verification:Stage 1 \(Correctness\):✓\\checkmarkPassStage 2 \(Hallucination\):✓\\checkmarkPassStage 3 \(Reasoning\):×\\timesFail— supported: falseJudge critique:“While both excerpts clearly study evaluating LLMs on exploration/exploitation behavior in bandit tasks, Evidence A does not state that this is done via in\-context prompting or that it excludes training interventions, so the reason over\-attributes shared task details not supported by the provided text\.”Failure type:*Reason exceeds evidence support*— The model adds specific details \(“via in\-context prompting,” “without training interventions”\) that only appear in Evidence B, incorrectly attributing them to both papers\.
Case M3: Dimension Confusion \(GPT\-5\.4\-mini\)Paper pair: Paper A:*Action Robust Reinforcement Learning with Probabilistic Policy Execution Uncertainty* Paper B:*Robust Reinforcement Learning via Adversarial Training*Dimension:methodGold label:similarPrediction:similar \(✓ correct\)Content mode:abstract onlyModel output:Evidence A\(verbatim\): “we focus on action robust RL with the probabilistic policy execution uncertainty, in which, instead of always carrying out the action specified by the policy, the agent will take the action specified by the policy with probability 1\-ρ\\rhoand an alternative adversarial action with probabilityρ\\rho\. We establish the existence of an optimal policy on the action robust MDPs with probabilistic policy execution uncertainty and provide the action robust Bellman optimality equation for its solution\.”Evidence B\(verbatim\): “Specifically, we consider two scenarios in which the agent attempts to perform an actionaa, and \(i\) with probabilityα\\alpha, an alternative adversarial actiona¯\\bar\{a\}is taken, or \(ii\) an adversary adds a perturbation to the selected action in the case of continuous action space\.”Reason:“Both papers model action uncertainty by replacing the agent’s chosen action with an adversarial alternative with some probability\. They each build reinforcement\-learning methods on top of that same probabilistic action\-corruption mechanism\.”Cascading verification:Stage 1 \(Correctness\):✓\\checkmarkPassStage 2 \(Hallucination\):✓\\checkmarkPassStage 3 \(Reasoning\):×\\timesFail— supported: falseJudge critique:“The reason identifies a shared problem formulation \(probabilistic action corruption\) rather than a specific algorithmic mechanism or architectural design, and Evidence B lacks any description of the actual reinforcement learning method used to solve the scenario\.”Failure type:*Dimension confusion*— The reason describes a shared*problem*\(action uncertainty\), but the pair is labeled as method dimension\. The model confuses problem formulation with methodological similarity\.
## Appendix PPrompt Templates
This section provides the complete prompt templates used for all evaluation tasks in NoveltyBench\. We include these prompts to support reproducibility \(researchers can replicate our evaluation protocol exactly\), transparency \(prompt design directly affects model behavior and evaluation outcomes\), and extensibility \(these templates can be adapted to evaluate novelty in new domains or dimensions\)\.
All prompts follow a consistent structure: \(1\) aRoledefinition establishing the evaluator’s expertise; \(2\) aDimension Definitionproviding a precise, operational definition of the target dimension; \(3\)Decision Rulesspecifying when to mark pairs as similar, with edge\-case guidance; \(4\) anOutput Formatenforcing structured JSON with evidence extraction and reasoning; and \(5\)Inputwith paper content placeholders \(<TITLE\_A\>,<CONTENT\_A\>, etc\.\) filled at runtime\.
We provide seven templates: three for pairwise novelty judgment \(task, problem, method\), three for multi\-paper grouping \(task, problem, method\), and one for reasoning verification \(Stage 3 of the cascading evaluation\)\.
### P\.1Pairwise Novelty Judgment: Task Dimension
Pairwise Task Prompt[⬇](data:text/plain;base64,IyBSb2xlCllvdSBhcmUgYW4gZXhwZXJ0IHJlc2VhcmNoIGFuYWx5c3QuIFlvdXIgdGFzayBpcyB0byBqdWRnZSB3aGV0aGVyIFBhcGVyIEEgYW5kIFBhcGVyIEIgdGFyZ2V0IHRoZSBzYW1lICoqdGFzayoqLgoKIyBEaW1lbnNpb24gRGVmaW5pdGlvbgpgdGFza2A6IHRoZSBoaWdoLWxldmVsIHJlc2VhcmNoIG9iamVjdGl2ZSBvciBhcHBsaWNhdGlvbiBzY2VuYXJpbyB0aGF0IGEgcGFwZXIgZm9jdXNlcyBvbi4KLSBBIHRhc2sgaXMgc3BlY2lmaWMgZW5vdWdoIHRvIGRpc3Rpbmd1aXNoIHJlc2VhcmNoIGRpcmVjdGlvbnMgKGUuZy4sICJvcGVuLXNldCBvYmplY3QgZGV0ZWN0aW9uIiwgInRleHQtdG8tdmlkZW8gZ2VuZXJhdGlvbiIsICJvZmZsaW5lIHJlaW5mb3JjZW1lbnQgbGVhcm5pbmcgZnJvbSBodW1hbiBmZWVkYmFjayIpLCB5ZXQgYnJvYWQgZW5vdWdoIHRoYXQgbXVsdGlwbGUgcGFwZXJzIGNhbiB0YWNrbGUgaXQgd2l0aCBkaWZmZXJlbnQgbWV0aG9kcy4KLSBUaGUgdGFzayBkZXNjcmliZXMgKip3aGF0KiogdGhlIHBhcGVyIGlzIHRyeWluZyB0byBhY2NvbXBsaXNoIGluIHRoZSB3b3JsZCwgbm90ICoqaG93KiogKG1ldGhvZCkgb3IgKip3aHkgZXhpc3Rpbmcgc29sdXRpb25zIGZhaWwqKiAocHJvYmxlbSkuCi0gUGFwZXJzIGNhbiBzaGFyZSBhIHRhc2sgZXZlbiBpZiB0aGV5IHVzZSBlbnRpcmVseSBkaWZmZXJlbnQgbWV0aG9kczsgcGFwZXJzIGNhbiBzaGFyZSBhIG1ldGhvZCB3aGlsZSB0YXJnZXRpbmcgZGlmZmVyZW50IHRhc2tzLgoKIyBEZWNpc2lvbiBSdWxlcwotIE1hcmsgYGlzX3NpbWlsYXI6IHRydWVgIGlmIGJvdGggcGFwZXJzIHRhcmdldCB0aGUgc2FtZSBvciBjbG9zZWx5IHJlbGF0ZWQgc3BlY2lmaWMgdGFzayBzZXR0aW5nLgotICJDbG9zZWx5IHJlbGF0ZWQiIG1lYW5zIGEgcmVzZWFyY2hlciBmYW1pbGlhciB3aXRoIGJvdGggd291bGQgZmlsZSB0aGVtIHVuZGVyIHRoZSBzYW1lIHN1Yi10b3BpYyBoZWFkaW5nIGluIGEgbGl0ZXJhdHVyZSByZXZpZXcuCi0gRG8gTk9UIG1hcmsgc2ltaWxhciBtZXJlbHkgYmVjYXVzZSBib3RoIHBhcGVycyBiZWxvbmcgdG8gdGhlIHNhbWUgYnJvYWQgZmllbGQgKGUuZy4sIGJvdGggZG8gTkxQLCBDb21wdXRlciBWaXNpb24sIG9yIFJlaW5mb3JjZW1lbnQgTGVhcm5pbmcpLgotIERvIE5PVCBtYXJrIHNpbWlsYXIgYmFzZWQgb24gc2hhcmVkIG1ldGhvZHMgb3IgcHJvYmxlbXMgYWxvbmUg4oCUIGZvY3VzIGV4Y2x1c2l2ZWx5IG9uIHRoZSB0YXNrIHNldHRpbmcuCi0gVXNlIG9ubHkgdGhlIHByb3ZpZGVkIHBhcGVyIGNvbnRlbnQgYXMgZXZpZGVuY2U7IGRvIG5vdCBpbnZlbnQgaW5mb3JtYXRpb24uCgojIE91dHB1dCBGb3JtYXQKT3V0cHV0IHZhbGlkIEpTT04gb25seS4gTm8gbWFya2Rvd24sIG5vIHRleHQgb3V0c2lkZSB0aGUgSlNPTiBvYmplY3QuCgpXaGVuIGBpc19zaW1pbGFyYCBpcyB0cnVlOgp7CiAgImlzX3NpbWlsYXIiOiB0cnVlLAogICJldmlkZW5jZV9hIjogImNvcHktcGFzdGUgZXhhY3Qgc3BhbiBmcm9tIFBhcGVyIEEgaWRlbnRpZnlpbmcgaXRzIHRhc2sg4oCUIHZlcmJhdGltIHN1YnN0cmluZywgMeKAkzMgc2VudGVuY2VzIiwKICAiZXZpZGVuY2VfYiI6ICJjb3B5LXBhc3RlIGV4YWN0IHNwYW4gZnJvbSBQYXBlciBCIGlkZW50aWZ5aW5nIGl0cyB0YXNrIOKAlCBzYW1lIHJlcXVpcmVtZW50cyIsCiAgInJlYXNvbiI6ICIx4oCTMiBzZW50ZW5jZXMgd2h5IHRhc2tzIGFyZSBzYW1lL3JlbGF0ZWQg4oCUIGRlcml2ZWQgU1RSSUNUTFkgZnJvbSBldmlkZW5jZSBhYm92ZSIKfQoKV2hlbiBgaXNfc2ltaWxhcmAgaXMgZmFsc2U6CnsKICAiaXNfc2ltaWxhciI6IGZhbHNlLAogICJldmlkZW5jZV9hIjogImNvcHktcGFzdGUgZXhhY3Qgc3BhbiBmcm9tIFBhcGVyIEEgaWRlbnRpZnlpbmcgaXRzIHRhc2sg4oCUIHZlcmJhdGltIHN1YnN0cmluZyIsCiAgImV2aWRlbmNlX2IiOiAiY29weS1wYXN0ZSBleGFjdCBzcGFuIGZyb20gUGFwZXIgQiBpZGVudGlmeWluZyBpdHMgdGFzayDigJQgdmVyYmF0aW0gc3Vic3RyaW5nIgp9CgojIElucHV0CltQQVBFUiBBXQpUaXRsZTogPFRJVExFX0E+CkNvbnRlbnQgdHlwZTogPENPTlRFTlRfVFlQRV9BPgpDb250ZW50OiA8Q09OVEVOVF9BPgpbUEFQRVIgQSBFTkRdCgpbUEFQRVIgQl0KVGl0bGU6IDxUSVRMRV9CPgpDb250ZW50IHR5cGU6IDxDT05URU5UX1RZUEVfQj4KQ29udGVudDogPENPTlRFTlRfQj4KW1BBUEVSIEIgRU5EXQ==)\#RoleYouareanexpertresearchanalyst\.YourtaskistojudgewhetherPaperAandPaperBtargetthesame\*\*task\*\*\.\#DimensionDefinition‘task‘:thehigh\-levelresearchobjectiveorapplicationscenariothatapaperfocuseson\.\-Ataskisspecificenoughtodistinguishresearchdirections\(e\.g\.,"open\-setobjectdetection","text\-to\-videogeneration","offlinereinforcementlearningfromhumanfeedback"\),yetbroadenoughthatmultiplepaperscantackleitwithdifferentmethods\.\-Thetaskdescribes\*\*what\*\*thepaperistryingtoaccomplishintheworld,not\*\*how\*\*\(method\)or\*\*whyexistingsolutionsfail\*\*\(problem\)\.\-Paperscanshareataskeveniftheyuseentirelydifferentmethods;paperscanshareamethodwhiletargetingdifferenttasks\.\#DecisionRules\-Mark‘is\_similar:true‘ifbothpaperstargetthesameorcloselyrelatedspecifictasksetting\.\-"Closelyrelated"meansaresearcherfamiliarwithbothwouldfilethemunderthesamesub\-topicheadinginaliteraturereview\.\-DoNOTmarksimilarmerelybecausebothpapersbelongtothesamebroadfield\(e\.g\.,bothdoNLP,ComputerVision,orReinforcementLearning\)\.\-DoNOTmarksimilarbasedonsharedmethodsorproblemsalone\-\-\-focusexclusivelyonthetasksetting\.\-Useonlytheprovidedpapercontentasevidence;donotinventinformation\.\#OutputFormatOutputvalidJSONonly\.Nomarkdown,notextoutsidetheJSONobject\.When‘is\_similar‘istrue:\{"is\_similar":true,"evidence\_a":"copy\-pasteexactspanfromPaperAidentifyingitstask\-\-\-verbatimsubstring,1\-\-3sentences","evidence\_b":"copy\-pasteexactspanfromPaperBidentifyingitstask\-\-\-samerequirements","reason":"1\-\-2sentenceswhytasksaresame/related\-\-\-derivedSTRICTLYfromevidenceabove"\}When‘is\_similar‘isfalse:\{"is\_similar":false,"evidence\_a":"copy\-pasteexactspanfromPaperAidentifyingitstask\-\-\-verbatimsubstring","evidence\_b":"copy\-pasteexactspanfromPaperBidentifyingitstask\-\-\-verbatimsubstring"\}\#Input\[PAPERA\]Title:<TITLE\_A\>Contenttype:<CONTENT\_TYPE\_A\>Content:<CONTENT\_A\>\[PAPERAEND\]\[PAPERB\]Title:<TITLE\_B\>Contenttype:<CONTENT\_TYPE\_B\>Content:<CONTENT\_B\>\[PAPERBEND\]
### P\.2Pairwise Novelty Judgment: Problem Dimension
Pairwise Problem Prompt[⬇](data:text/plain;base64,IyBSb2xlCllvdSBhcmUgYW4gZXhwZXJ0IHJlc2VhcmNoIGFuYWx5c3QuIFlvdXIgdGFzayBpcyB0byBqdWRnZSB3aGV0aGVyIFBhcGVyIEEgYW5kIFBhcGVyIEIgYWRkcmVzcyB0aGUgc2FtZSAqKnByb2JsZW0qKi4KCiMgRGltZW5zaW9uIERlZmluaXRpb24KYHByb2JsZW1gOiB0aGUgc3BlY2lmaWMgdGVjaG5pY2FsIGNoYWxsZW5nZSwgYm90dGxlbmVjaywgbGltaXRhdGlvbiwgb3IgZ2FwIHRoYXQgdGhlIHBhcGVyIGlzIHRyeWluZyB0byBzb2x2ZSDigJQgdGhlIHByZWNpc2UgcmVhc29uIHdoeSBleGlzdGluZyBzb2x1dGlvbnMgYXJlIGluc3VmZmljaWVudC4KLSBBIHByb2JsZW0gZGVzY3JpYmVzICoqd2h5KiogY3VycmVudCBhcHByb2FjaGVzIGZhaWwgb3IgZmFsbCBzaG9ydCwgbm90IHdoYXQgdGhlIHBhcGVyIGRvZXMgKG1ldGhvZCkgb3Igd2hhdCBpdCBhcHBsaWVzIHRvICh0YXNrKS4KLSBQcm9ibGVtcyBjYW4gYmUgc2hhcmVkIGFjcm9zcyBlbnRpcmVseSBkaWZmZXJlbnQgdGFza3MgKGUuZy4sICJkaXN0cmlidXRpb24gc2hpZnQgYXQgdGVzdCB0aW1lIiBhZmZlY3RzIGJvdGggaW1hZ2UgY2xhc3NpZmljYXRpb24gYW5kIG1hY2hpbmUgdHJhbnNsYXRpb24pLgotIEEgYnJvYWQgZ29hbCBvciBhcHBsaWNhdGlvbiBjYXBhYmlsaXR5IChlLmcuLCAiaW1wcm92ZSBhY2N1cmFjeSIsICJzY2FsZSB0byBsYXJnZXIgZGF0YXNldHMiKSBkb2VzIE5PVCBjb25zdGl0dXRlIGEgcHJvYmxlbS4KLSBUaGUgcHJvYmxlbSBtdXN0IGlkZW50aWZ5IGEgY29uY3JldGUgdGVjaG5pY2FsIG9ic3RhY2xlLCBjb25zdHJhaW50LCBvciBnYXAgdGhhdCBtb3RpdmF0ZXMgdGhlIHNwZWNpZmljIGNvbnRyaWJ1dGlvbi4KCiMgRGVjaXNpb24gUnVsZXMKLSBNYXJrIGBpc19zaW1pbGFyOiB0cnVlYCBpZiBib3RoIHBhcGVycyBhZGRyZXNzIHRoZSBzYW1lIHNwZWNpZmljIHRlY2huaWNhbCBvYnN0YWNsZSBvciBib3R0bGVuZWNrLCBldmVuIGlmIHRoZXkgb3BlcmF0ZSBvbiBkaWZmZXJlbnQgdGFza3Mgb3IgZG9tYWlucy4KLSBEbyBOT1QgbWFyayBzaW1pbGFyIGlmIHRoZSBwYXBlcnMgb25seSBzaGFyZSB0aGUgc2FtZSB0YXNrIG9yIGFwcGxpY2F0aW9uIGRvbWFpbiB3aXRob3V0IGEgc2hhcmVkIGJvdHRsZW5lY2suCi0gRG8gTk9UIG1hcmsgc2ltaWxhciBpZiB0aGUgcHJvYmxlbXMgYXJlIGxvb3NlbHkgcmVsYXRlZCBhdCBhIHN1cmZhY2UgbGV2ZWwgYnV0IHRhcmdldCBjbGVhcmx5IGRpZmZlcmVudCBmYWlsdXJlIG1vZGVzIG9yIGNvbnN0cmFpbnRzLgotIFVzZSBvbmx5IHRoZSBwcm92aWRlZCBwYXBlciBjb250ZW50IGFzIGV2aWRlbmNlOyBkbyBub3QgaW52ZW50IGluZm9ybWF0aW9uLgoKIyBPdXRwdXQgRm9ybWF0Ck91dHB1dCB2YWxpZCBKU09OIG9ubHkuIE5vIG1hcmtkb3duLCBubyB0ZXh0IG91dHNpZGUgdGhlIEpTT04gb2JqZWN0LgoKV2hlbiBgaXNfc2ltaWxhcmAgaXMgdHJ1ZToKewogICJpc19zaW1pbGFyIjogdHJ1ZSwKICAiZXZpZGVuY2VfYSI6ICJjb3B5LXBhc3RlIGV4YWN0IHNwYW4gZnJvbSBQYXBlciBBIHN0YXRpbmcgaXRzIHRlY2huaWNhbCBwcm9ibGVtIOKAlCB2ZXJiYXRpbSwgc2VsZi1jb250YWluZWQsIDHigJMzIHNlbnRlbmNlcyIsCiAgImV2aWRlbmNlX2IiOiAiY29weS1wYXN0ZSBleGFjdCBzcGFuIGZyb20gUGFwZXIgQiBzdGF0aW5nIGl0cyB0ZWNobmljYWwgcHJvYmxlbSDigJQgc2FtZSByZXF1aXJlbWVudHMiLAogICJyZWFzb24iOiAiMeKAkzIgc2VudGVuY2VzIHdoeSBwcm9ibGVtcyBhcmUgc2FtZS9yZWxhdGVkIOKAlCBkZXJpdmVkIFNUUklDVExZIGZyb20gZXZpZGVuY2UsIG5vIG5ldyB0ZXJtcyIKfQoKV2hlbiBgaXNfc2ltaWxhcmAgaXMgZmFsc2U6CnsKICAiaXNfc2ltaWxhciI6IGZhbHNlLAogICJldmlkZW5jZV9hIjogImNvcHktcGFzdGUgZXhhY3Qgc3BhbiBmcm9tIFBhcGVyIEEgaWRlbnRpZnlpbmcgaXRzIHRlY2huaWNhbCBwcm9ibGVtIOKAlCB2ZXJiYXRpbSwgMeKAkzMgc2VudGVuY2VzIiwKICAiZXZpZGVuY2VfYiI6ICJjb3B5LXBhc3RlIGV4YWN0IHNwYW4gZnJvbSBQYXBlciBCIGlkZW50aWZ5aW5nIGl0cyB0ZWNobmljYWwgcHJvYmxlbSDigJQgdmVyYmF0aW0sIDHigJMzIHNlbnRlbmNlcyIKfQoKIyBJbnB1dApbUEFQRVIgQV0KVGl0bGU6IDxUSVRMRV9BPgpDb250ZW50IHR5cGU6IDxDT05URU5UX1RZUEVfQT4KQ29udGVudDogPENPTlRFTlRfQT4KW1BBUEVSIEEgRU5EXQoKW1BBUEVSIEJdClRpdGxlOiA8VElUTEVfQj4KQ29udGVudCB0eXBlOiA8Q09OVEVOVF9UWVBFX0I+CkNvbnRlbnQ6IDxDT05URU5UX0I+CltQQVBFUiBCIEVORF0=)\#RoleYouareanexpertresearchanalyst\.YourtaskistojudgewhetherPaperAandPaperBaddressthesame\*\*problem\*\*\.\#DimensionDefinition‘problem‘:thespecifictechnicalchallenge,bottleneck,limitation,orgapthatthepaperistryingtosolve\-\-\-theprecisereasonwhyexistingsolutionsareinsufficient\.\-Aproblemdescribes\*\*why\*\*currentapproachesfailorfallshort,notwhatthepaperdoes\(method\)orwhatitappliesto\(task\)\.\-Problemscanbesharedacrossentirelydifferenttasks\(e\.g\.,"distributionshiftattesttime"affectsbothimageclassificationandmachinetranslation\)\.\-Abroadgoalorapplicationcapability\(e\.g\.,"improveaccuracy","scaletolargerdatasets"\)doesNOTconstituteaproblem\.\-Theproblemmustidentifyaconcretetechnicalobstacle,constraint,orgapthatmotivatesthespecificcontribution\.\#DecisionRules\-Mark‘is\_similar:true‘ifbothpapersaddressthesamespecifictechnicalobstacleorbottleneck,eveniftheyoperateondifferenttasksordomains\.\-DoNOTmarksimilarifthepapersonlysharethesametaskorapplicationdomainwithoutasharedbottleneck\.\-DoNOTmarksimilariftheproblemsarelooselyrelatedatasurfacelevelbuttargetclearlydifferentfailuremodesorconstraints\.\-Useonlytheprovidedpapercontentasevidence;donotinventinformation\.\#OutputFormatOutputvalidJSONonly\.Nomarkdown,notextoutsidetheJSONobject\.When‘is\_similar‘istrue:\{"is\_similar":true,"evidence\_a":"copy\-pasteexactspanfromPaperAstatingitstechnicalproblem\-\-\-verbatim,self\-contained,1\-\-3sentences","evidence\_b":"copy\-pasteexactspanfromPaperBstatingitstechnicalproblem\-\-\-samerequirements","reason":"1\-\-2sentenceswhyproblemsaresame/related\-\-\-derivedSTRICTLYfromevidence,nonewterms"\}When‘is\_similar‘isfalse:\{"is\_similar":false,"evidence\_a":"copy\-pasteexactspanfromPaperAidentifyingitstechnicalproblem\-\-\-verbatim,1\-\-3sentences","evidence\_b":"copy\-pasteexactspanfromPaperBidentifyingitstechnicalproblem\-\-\-verbatim,1\-\-3sentences"\}\#Input\[PAPERA\]Title:<TITLE\_A\>Contenttype:<CONTENT\_TYPE\_A\>Content:<CONTENT\_A\>\[PAPERAEND\]\[PAPERB\]Title:<TITLE\_B\>Contenttype:<CONTENT\_TYPE\_B\>Content:<CONTENT\_B\>\[PAPERBEND\]
### P\.3Pairwise Novelty Judgment: Method Dimension
Pairwise Method Prompt[⬇](data:text/plain;base64,IyBSb2xlCllvdSBhcmUgYW4gZXhwZXJ0IHJlc2VhcmNoIGFuYWx5c3QuIFlvdXIgdGFzayBpcyB0byBqdWRnZSB3aGV0aGVyIFBhcGVyIEEgYW5kIFBhcGVyIEIgcHJvcG9zZSBzaW1pbGFyICoqbWV0aG9kcyoqLgoKIyBEaW1lbnNpb24gRGVmaW5pdGlvbgpgbWV0aG9kYDogdGhlIGNvcmUgdGVjaG5pY2FsIGFwcHJvYWNoLCBhbGdvcml0aG1pYyBpZGVhLCBtb2RlbCBhcmNoaXRlY3R1cmUsIG9yIHRyYWluaW5nL2luZmVyZW5jZSBzdHJhdGVneSB0aGF0IGEgcGFwZXIgY29udHJpYnV0ZXMuCi0gQSBtZXRob2QgZGVzY3JpYmVzICoqaG93KiogdGhlIHBhcGVyIHNvbHZlcyBpdHMgcHJvYmxlbSDigJQgdGhlIHNwZWNpZmljIHRlY2huaWNhbCBtZWNoYW5pc20sIG5vdCB0aGUgcHJvYmxlbSBpdHNlbGYgb3IgdGhlIHRhc2suCi0gVHdvIG1ldGhvZHMgYXJlIHNpbWlsYXIgaWYgdGhleSBzaGFyZSBhIGZ1bmRhbWVudGFsIGFsZ29yaXRobWljIG1lY2hhbmlzbSBvciBjb25jZXB0dWFsIGJhY2tib25lIGF0IGEgbWVhbmluZ2Z1bCBsZXZlbCBvZiBzcGVjaWZpY2l0eS4KLSBNZXRob2RzIGNhbiBiZSBzaGFyZWQgYWNyb3NzIGRpZmZlcmVudCB0YXNrcyAoZS5nLiwgY29udHJhc3RpdmUgbGVhcm5pbmcsIFJMSEYsIGtub3dsZWRnZSBkaXN0aWxsYXRpb24gZWFjaCBhcHBlYXIgYWNyb3NzIG1hbnkgYXBwbGljYXRpb24gYXJlYXMpLgotIEEgYnJvYWQgdGVjaG5pcXVlIGNhdGVnb3J5IChlLmcuLCAidXNlcyB0cmFuc2Zvcm1lcnMiLCAidXNlcyBkaWZmdXNpb24iLCAidXNlcyByZWluZm9yY2VtZW50IGxlYXJuaW5nIikgaXMgTk9UIHN1ZmZpY2llbnQg4oCUIHJlcXVpcmUgc2hhcmVkIHNwZWNpZmljIG1lY2hhbmlzbXMuCgojIERlY2lzaW9uIFJ1bGVzCi0gTWFyayBgaXNfc2ltaWxhcjogdHJ1ZWAgaWYgYm90aCBwYXBlcnMgc2hhcmUgYSBmdW5kYW1lbnRhbCB0ZWNobmljYWwgbWVjaGFuaXNtLCBjb3JlIGFsZ29yaXRobWljIHBpcGVsaW5lLCBvciBzcGVjaWZpYyBhcmNoaXRlY3R1cmFsIGRlc2lnbiBwcmluY2lwbGUuCi0gIlNoYXJlZCBtZWNoYW5pc20iIG1lYW5zOiBpZiB5b3UgZGVzY3JpYmVkIHRoZSBtZXRob2Qgb2Ygb25lIHBhcGVyIHRvIGFuIGV4cGVydCwgdGhleSB3b3VsZCByZWNvZ25pemUgdGhlIGNvcmUgaWRlYSBhcyB0aGUgc2FtZSBhcyB0aGUgb3RoZXIgcGFwZXIncyBtZXRob2QuCi0gRG8gTk9UIG1hcmsgc2ltaWxhciBiYXNlZCBvbjoKICAtIEJvdGggYmVsb25naW5nIHRvIHRoZSBzYW1lIGhpZ2gtbGV2ZWwgdGVjaG5pcXVlIGZhbWlseSB3aXRob3V0IHNoYXJpbmcgYSBzcGVjaWZpYyBtZWNoYW5pc20gKGUuZy4sICJib3RoIGFyZSBkaWZmdXNpb24gbW9kZWxzIiwgImJvdGggdXNlIGF0dGVudGlvbiIpLgogIC0gU2ltaWxhciBwcm9ibGVtIGZyYW1pbmcgb3IgdGFzayBzZXR0aW5nIHdpdGhvdXQgc2hhcmVkIHRlY2huaWNhbCBhcHByb2FjaC4KICAtIFNoYXJlZCBwcmVwcm9jZXNzaW5nLCBldmFsdWF0aW9uIHByb3RvY29sLCBvciBkYXRhc2V0LgogIC0gQm90aCBiZWluZyBmaW5lLXR1bmluZyBhcHByb2FjaGVzIHRvIExMTXMgd2l0aG91dCBzaGFyaW5nIHNwZWNpZmljIGZpbmUtdHVuaW5nIHN0cmF0ZWd5LgoKIyBPdXRwdXQgRm9ybWF0Ck91dHB1dCB2YWxpZCBKU09OIG9ubHkuIE5vIG1hcmtkb3duLCBubyB0ZXh0IG91dHNpZGUgdGhlIEpTT04gb2JqZWN0LgoKV2hlbiBgaXNfc2ltaWxhcmAgaXMgdHJ1ZToKewogICJpc19zaW1pbGFyIjogdHJ1ZSwKICAiZXZpZGVuY2VfYSI6ICJjb3B5LXBhc3RlIGV4YWN0IHNwYW4gZnJvbSBQYXBlciBBIGRlc2NyaWJpbmcgaXRzIGNvcmUgdGVjaG5pY2FsIG1lY2hhbmlzbSDigJQgdmVyYmF0aW0sIHNlbGYtY29udGFpbmVkLCAx4oCTMyBzZW50ZW5jZXMiLAogICJldmlkZW5jZV9iIjogImNvcHktcGFzdGUgZXhhY3Qgc3BhbiBmcm9tIFBhcGVyIEIgZGVzY3JpYmluZyBpdHMgY29yZSB0ZWNobmljYWwgbWVjaGFuaXNtIOKAlCBzYW1lIHJlcXVpcmVtZW50cyIsCiAgInJlYXNvbiI6ICIx4oCTMiBzZW50ZW5jZXMgc3RhdGluZyB0aGUgc3BlY2lmaWMgc2hhcmVkIG1lY2hhbmlzbSDigJQgZGVyaXZlZCBTVFJJQ1RMWSBmcm9tIGV2aWRlbmNlLCBubyBuZXcgdGVybXMiCn0KCldoZW4gYGlzX3NpbWlsYXJgIGlzIGZhbHNlOgp7CiAgImlzX3NpbWlsYXIiOiBmYWxzZSwKICAiZXZpZGVuY2VfYSI6ICJjb3B5LXBhc3RlIGV4YWN0IHNwYW4gZnJvbSBQYXBlciBBIGlkZW50aWZ5aW5nIGl0cyBjb3JlIHRlY2huaWNhbCBtZWNoYW5pc20g4oCUIHZlcmJhdGltLCAx4oCTMyBzZW50ZW5jZXMiLAogICJldmlkZW5jZV9iIjogImNvcHktcGFzdGUgZXhhY3Qgc3BhbiBmcm9tIFBhcGVyIEIgaWRlbnRpZnlpbmcgaXRzIGNvcmUgdGVjaG5pY2FsIG1lY2hhbmlzbSDigJQgdmVyYmF0aW0sIDHigJMzIHNlbnRlbmNlcyIKfQoKIyBJbnB1dApbUEFQRVIgQV0KVGl0bGU6IDxUSVRMRV9BPgpDb250ZW50IHR5cGU6IDxDT05URU5UX1RZUEVfQT4KQ29udGVudDogPENPTlRFTlRfQT4KW1BBUEVSIEEgRU5EXQoKW1BBUEVSIEJdClRpdGxlOiA8VElUTEVfQj4KQ29udGVudCB0eXBlOiA8Q09OVEVOVF9UWVBFX0I+CkNvbnRlbnQ6IDxDT05URU5UX0I+CltQQVBFUiBCIEVORF0=)\#RoleYouareanexpertresearchanalyst\.YourtaskistojudgewhetherPaperAandPaperBproposesimilar\*\*methods\*\*\.\#DimensionDefinition‘method‘:thecoretechnicalapproach,algorithmicidea,modelarchitecture,ortraining/inferencestrategythatapapercontributes\.\-Amethoddescribes\*\*how\*\*thepapersolvesitsproblem\-\-\-thespecifictechnicalmechanism,nottheproblemitselforthetask\.\-Twomethodsaresimilariftheyshareafundamentalalgorithmicmechanismorconceptualbackboneatameaningfullevelofspecificity\.\-Methodscanbesharedacrossdifferenttasks\(e\.g\.,contrastivelearning,RLHF,knowledgedistillationeachappearacrossmanyapplicationareas\)\.\-Abroadtechniquecategory\(e\.g\.,"usestransformers","usesdiffusion","usesreinforcementlearning"\)isNOTsufficient\-\-\-requiresharedspecificmechanisms\.\#DecisionRules\-Mark‘is\_similar:true‘ifbothpapersshareafundamentaltechnicalmechanism,corealgorithmicpipeline,orspecificarchitecturaldesignprinciple\.\-"Sharedmechanism"means:ifyoudescribedthemethodofonepapertoanexpert,theywouldrecognizethecoreideaasthesameastheotherpaper’smethod\.\-DoNOTmarksimilarbasedon:\-Bothbelongingtothesamehigh\-leveltechniquefamilywithoutsharingaspecificmechanism\(e\.g\.,"botharediffusionmodels","bothuseattention"\)\.\-Similarproblemframingortasksettingwithoutsharedtechnicalapproach\.\-Sharedpreprocessing,evaluationprotocol,ordataset\.\-Bothbeingfine\-tuningapproachestoLLMswithoutsharingspecificfine\-tuningstrategy\.\#OutputFormatOutputvalidJSONonly\.Nomarkdown,notextoutsidetheJSONobject\.When‘is\_similar‘istrue:\{"is\_similar":true,"evidence\_a":"copy\-pasteexactspanfromPaperAdescribingitscoretechnicalmechanism\-\-\-verbatim,self\-contained,1\-\-3sentences","evidence\_b":"copy\-pasteexactspanfromPaperBdescribingitscoretechnicalmechanism\-\-\-samerequirements","reason":"1\-\-2sentencesstatingthespecificsharedmechanism\-\-\-derivedSTRICTLYfromevidence,nonewterms"\}When‘is\_similar‘isfalse:\{"is\_similar":false,"evidence\_a":"copy\-pasteexactspanfromPaperAidentifyingitscoretechnicalmechanism\-\-\-verbatim,1\-\-3sentences","evidence\_b":"copy\-pasteexactspanfromPaperBidentifyingitscoretechnicalmechanism\-\-\-verbatim,1\-\-3sentences"\}\#Input\[PAPERA\]Title:<TITLE\_A\>Contenttype:<CONTENT\_TYPE\_A\>Content:<CONTENT\_A\>\[PAPERAEND\]\[PAPERB\]Title:<TITLE\_B\>Contenttype:<CONTENT\_TYPE\_B\>Content:<CONTENT\_B\>\[PAPERBEND\]
### P\.4Multi\-paper Grouping: Task Dimension
Grouping Task Prompt[⬇](data:text/plain;base64,IyBSb2xlCllvdSBhcmUgYW4gZXhwZXJ0IHJlc2VhcmNoIGFuYWx5c3QuIFlvdSBhcmUgZ2l2ZW4gPE5fUEFQRVJTPiByZXNlYXJjaCBwYXBlcnMuIFlvdXIgdGFzayBpcyB0byBpZGVudGlmeSB3aGljaCBwYXBlcnMgYWRkcmVzcyB0aGUgc2FtZSAqKnRhc2sqKiBhbmQgZ3JvdXAgdGhlbSBhY2NvcmRpbmdseS4KCiMgRGltZW5zaW9uIERlZmluaXRpb24KYHRhc2tgOiB0aGUgc3BlY2lmaWMgYXBwbGljYXRpb24gZG9tYWluLCBwcm9ibGVtIHNldHRpbmcsIG9yIHVzZSBjYXNlIHRoYXQgdGhlIHBhcGVyIHRhcmdldHMg4oCUIHRoZSBjb25jcmV0ZSBzY2VuYXJpbyB3aGVyZSB0aGUgbWV0aG9kIGlzIGFwcGxpZWQuCi0gQSB0YXNrIGRlc2NyaWJlcyAqKndoYXQqKiB0aGUgcGFwZXIgaXMgdHJ5aW5nIHRvIGFjY29tcGxpc2ggaW4gdGVybXMgb2YgaW5wdXQtb3V0cHV0IGJlaGF2aW9yIG9yIGFwcGxpY2F0aW9uIGNvbnRleHQsIG5vdCBob3cgaXQgZG9lcyBpdCAobWV0aG9kKSBvciB3aHkgZXhpc3Rpbmcgc29sdXRpb25zIGZhaWwgKHByb2JsZW0pLgotIFRhc2tzIGNhbiBiZSBzaGFyZWQgYWNyb3NzIGRpZmZlcmVudCBtZXRob2RzIChlLmcuLCAiaW1hZ2UgY2xhc3NpZmljYXRpb24iLCAibWFjaGluZSB0cmFuc2xhdGlvbiIsICJxdWVzdGlvbiBhbnN3ZXJpbmciIGVhY2ggaGF2ZSBtYW55IGRpZmZlcmVudCB0ZWNobmljYWwgYXBwcm9hY2hlcykuCi0gQSBicm9hZCByZXNlYXJjaCBhcmVhIChlLmcuLCAiY29tcHV0ZXIgdmlzaW9uIiwgIk5MUCIsICJyZWluZm9yY2VtZW50IGxlYXJuaW5nIikgaXMgTk9UIGEgdGFzayDigJQgcmVxdWlyZSBhIHNwZWNpZmljIGFwcGxpY2F0aW9uIHNldHRpbmcuCi0gVGhlIHRhc2sgbXVzdCBiZSBjb25jcmV0ZSBlbm91Z2ggdGhhdCBhIHByYWN0aXRpb25lciBjb3VsZCByZWNvZ25pemUgd2hldGhlciB0aGVpciB1c2UgY2FzZSBtYXRjaGVzIGl0LgoKIyBHcm91cGluZyBSdWxlcwotIFBsYWNlIHR3byBvciBtb3JlIHBhcGVycyBpbiB0aGUgc2FtZSBncm91cCBpZiB0aGV5IHRhcmdldCB0aGUgc2FtZSBvciBjbG9zZWx5IHJlbGF0ZWQgc3BlY2lmaWMgdGFzayBzZXR0aW5nLgotICJDbG9zZWx5IHJlbGF0ZWQiIG1lYW5zIGEgcmVzZWFyY2hlciBmYW1pbGlhciB3aXRoIGJvdGggd291bGQgZmlsZSB0aGVtIHVuZGVyIHRoZSBzYW1lIHN1Yi10b3BpYyBoZWFkaW5nIGluIGEgbGl0ZXJhdHVyZSByZXZpZXcuCi0gRG8gTk9UIGdyb3VwIHBhcGVycyBtZXJlbHkgYmVjYXVzZSB0aGV5IGJlbG9uZyB0byB0aGUgc2FtZSBicm9hZCBmaWVsZCAoZS5nLiwgYm90aCBkbyBOTFAsIENvbXB1dGVyIFZpc2lvbiwgb3IgUmVpbmZvcmNlbWVudCBMZWFybmluZykuCi0gRG8gTk9UIGdyb3VwIHBhcGVycyBiYXNlZCBvbiBzaGFyZWQgbWV0aG9kcyBvciBwcm9ibGVtcyBhbG9uZSDigJQgZm9jdXMgZXhjbHVzaXZlbHkgb24gdGhlIHRhc2sgc2V0dGluZy4KLSBBIHBhcGVyIGNhbiBiZWxvbmcgdG8gYXQgbW9zdCBvbmUgZ3JvdXAuIElmIGEgcGFwZXIgZG9lcyBub3Qgc2hhcmUgYSB0YXNrIHdpdGggYW55IG90aGVyIHBhcGVyLCBpdCBkb2VzIG5vdCBhcHBlYXIgaW4gYW55IGdyb3VwLgotIFRoZXJlIG1heSBiZSB6ZXJvLCBvbmUsIG9yIG11bHRpcGxlIGdyb3Vwcy4KLSBVc2Ugb25seSB0aGUgcHJvdmlkZWQgcGFwZXIgY29udGVudCBhcyBldmlkZW5jZTsgZG8gbm90IGludmVudCBpbmZvcm1hdGlvbi4KCiMgT3V0cHV0IEZvcm1hdApPdXRwdXQgdmFsaWQgSlNPTiBvbmx5LiBObyBtYXJrZG93biwgbm8gdGV4dCBvdXRzaWRlIHRoZSBKU09OIG9iamVjdC4KCldoZW4gZ3JvdXBzIGV4aXN0Ogp7CiAgImdyb3VwcyI6IFsKICAgIHsKICAgICAgInBhcGVyX2luZGljZXMiOiBbMSwgMl0sCiAgICAgICJldmlkZW5jZSI6IHsKICAgICAgICAiMSI6ICJjb3B5LXBhc3RlIGV4YWN0IHNwYW4gZnJvbSBQYXBlciAxIGlkZW50aWZ5aW5nIGl0cyB0YXNrIOKAlCB2ZXJiYXRpbSIsCiAgICAgICAgIjIiOiAiY29weS1wYXN0ZSBleGFjdCBzcGFuIGZyb20gUGFwZXIgMiBpZGVudGlmeWluZyBpdHMgdGFzayDigJQgdmVyYmF0aW0iCiAgICAgIH0sCiAgICAgICJyZWFzb24iOiAiMeKAkzIgc2VudGVuY2VzIHdoeSB0aGVzZSBwYXBlcnMgc2hhcmUgdGhlIHNhbWUgdGFzaywgcmVmZXJlbmNpbmcgZXZpZGVuY2UgYWJvdmUiCiAgICB9CiAgXQp9CgpXaGVuIG5vIGdyb3VwcyBleGlzdDoKeyJncm91cHMiOiBbXX0KCiMgSW5wdXQKPFBBUEVSUz4=)\#RoleYouareanexpertresearchanalyst\.Youaregiven<N\_PAPERS\>researchpapers\.Yourtaskistoidentifywhichpapersaddressthesame\*\*task\*\*andgroupthemaccordingly\.\#DimensionDefinition‘task‘:thespecificapplicationdomain,problemsetting,orusecasethatthepapertargets\-\-\-theconcretescenariowherethemethodisapplied\.\-Ataskdescribes\*\*what\*\*thepaperistryingtoaccomplishintermsofinput\-outputbehaviororapplicationcontext,nothowitdoesit\(method\)orwhyexistingsolutionsfail\(problem\)\.\-Taskscanbesharedacrossdifferentmethods\(e\.g\.,"imageclassification","machinetranslation","questionanswering"eachhavemanydifferenttechnicalapproaches\)\.\-Abroadresearcharea\(e\.g\.,"computervision","NLP","reinforcementlearning"\)isNOTatask\-\-\-requireaspecificapplicationsetting\.\-Thetaskmustbeconcreteenoughthatapractitionercouldrecognizewhethertheirusecasematchesit\.\#GroupingRules\-Placetwoormorepapersinthesamegroupiftheytargetthesameorcloselyrelatedspecifictasksetting\.\-"Closelyrelated"meansaresearcherfamiliarwithbothwouldfilethemunderthesamesub\-topicheadinginaliteraturereview\.\-DoNOTgrouppapersmerelybecausetheybelongtothesamebroadfield\(e\.g\.,bothdoNLP,ComputerVision,orReinforcementLearning\)\.\-DoNOTgrouppapersbasedonsharedmethodsorproblemsalone\-\-\-focusexclusivelyonthetasksetting\.\-Apapercanbelongtoatmostonegroup\.Ifapaperdoesnotshareataskwithanyotherpaper,itdoesnotappearinanygroup\.\-Theremaybezero,one,ormultiplegroups\.\-Useonlytheprovidedpapercontentasevidence;donotinventinformation\.\#OutputFormatOutputvalidJSONonly\.Nomarkdown,notextoutsidetheJSONobject\.Whengroupsexist:\{"groups":\[\{"paper\_indices":\[1,2\],"evidence":\{"1":"copy\-pasteexactspanfromPaper1identifyingitstask\-\-\-verbatim","2":"copy\-pasteexactspanfromPaper2identifyingitstask\-\-\-verbatim"\},"reason":"1\-\-2sentenceswhythesepaperssharethesametask,referencingevidenceabove"\}\]\}Whennogroupsexist:\{"groups":\[\]\}\#Input<PAPERS\>
### P\.5Multi\-paper Grouping: Problem Dimension
Grouping Problem Prompt[⬇](data:text/plain;base64,IyBSb2xlCllvdSBhcmUgYW4gZXhwZXJ0IHJlc2VhcmNoIGFuYWx5c3QuIFlvdSBhcmUgZ2l2ZW4gPE5fUEFQRVJTPiByZXNlYXJjaCBwYXBlcnMuIFlvdXIgdGFzayBpcyB0byBpZGVudGlmeSB3aGljaCBwYXBlcnMgYWRkcmVzcyB0aGUgc2FtZSAqKnByb2JsZW0qKiBhbmQgZ3JvdXAgdGhlbSBhY2NvcmRpbmdseS4KCiMgRGltZW5zaW9uIERlZmluaXRpb24KYHByb2JsZW1gOiB0aGUgc3BlY2lmaWMgdGVjaG5pY2FsIGNoYWxsZW5nZSwgYm90dGxlbmVjaywgbGltaXRhdGlvbiwgb3IgZ2FwIHRoYXQgYSBwYXBlciBpcyB0cnlpbmcgdG8gc29sdmUg4oCUIHRoZSBwcmVjaXNlIHJlYXNvbiB3aHkgZXhpc3Rpbmcgc29sdXRpb25zIGFyZSBpbnN1ZmZpY2llbnQuCi0gQSBwcm9ibGVtIGRlc2NyaWJlcyAqKndoeSoqIGN1cnJlbnQgYXBwcm9hY2hlcyBmYWlsIG9yIGZhbGwgc2hvcnQsIG5vdCB3aGF0IHRoZSBwYXBlciBkb2VzIChtZXRob2QpIG9yIHdoYXQgaXQgYXBwbGllcyB0byAodGFzaykuCi0gUHJvYmxlbXMgY2FuIGJlIHNoYXJlZCBhY3Jvc3MgZW50aXJlbHkgZGlmZmVyZW50IHRhc2tzIChlLmcuLCAiZGlzdHJpYnV0aW9uIHNoaWZ0IGF0IHRlc3QgdGltZSIgYWZmZWN0cyBib3RoIGltYWdlIGNsYXNzaWZpY2F0aW9uIGFuZCBtYWNoaW5lIHRyYW5zbGF0aW9uKS4KLSBBIGJyb2FkIGdvYWwgb3IgYXBwbGljYXRpb24gY2FwYWJpbGl0eSAoZS5nLiwgImltcHJvdmUgYWNjdXJhY3kiLCAic2NhbGUgdG8gbGFyZ2VyIGRhdGFzZXRzIikgZG9lcyBOT1QgY29uc3RpdHV0ZSBhIHByb2JsZW0uCi0gVGhlIHByb2JsZW0gbXVzdCBpZGVudGlmeSBhIGNvbmNyZXRlIHRlY2huaWNhbCBvYnN0YWNsZSwgY29uc3RyYWludCwgb3IgZ2FwIHRoYXQgbW90aXZhdGVzIHRoZSBzcGVjaWZpYyBjb250cmlidXRpb24uCgojIEdyb3VwaW5nIFJ1bGVzCi0gUGxhY2UgdHdvIG9yIG1vcmUgcGFwZXJzIGluIHRoZSBzYW1lIGdyb3VwIGlmIHRoZXkgYWRkcmVzcyB0aGUgc2FtZSBzcGVjaWZpYyB0ZWNobmljYWwgb2JzdGFjbGUgb3IgYm90dGxlbmVjaywgZXZlbiBpZiB0aGV5IG9wZXJhdGUgb24gZGlmZmVyZW50IHRhc2tzIG9yIGRvbWFpbnMuCi0gRG8gTk9UIGdyb3VwIHBhcGVycyB0aGF0IG9ubHkgc2hhcmUgdGhlIHNhbWUgdGFzayBvciBhcHBsaWNhdGlvbiBkb21haW4gd2l0aG91dCBhIHNoYXJlZCBib3R0bGVuZWNrLgotIERvIE5PVCBncm91cCBwYXBlcnMgd2hvc2UgcHJvYmxlbXMgYXJlIGxvb3NlbHkgcmVsYXRlZCBhdCBhIHN1cmZhY2UgbGV2ZWwgYnV0IHRhcmdldCBjbGVhcmx5IGRpZmZlcmVudCBmYWlsdXJlIG1vZGVzIG9yIGNvbnN0cmFpbnRzLgotIEEgcGFwZXIgY2FuIGJlbG9uZyB0byBhdCBtb3N0IG9uZSBncm91cC4gSWYgYSBwYXBlciBkb2VzIG5vdCBzaGFyZSBhIHByb2JsZW0gd2l0aCBhbnkgb3RoZXIgcGFwZXIsIGl0IGRvZXMgbm90IGFwcGVhciBpbiBhbnkgZ3JvdXAuCi0gVGhlcmUgbWF5IGJlIHplcm8sIG9uZSwgb3IgbXVsdGlwbGUgZ3JvdXBzLgotIFVzZSBvbmx5IHRoZSBwcm92aWRlZCBwYXBlciBjb250ZW50IGFzIGV2aWRlbmNlOyBkbyBub3QgaW52ZW50IGluZm9ybWF0aW9uLgoKIyBPdXRwdXQgRm9ybWF0Ck91dHB1dCB2YWxpZCBKU09OIG9ubHkuIE5vIG1hcmtkb3duLCBubyB0ZXh0IG91dHNpZGUgdGhlIEpTT04gb2JqZWN0LgoKV2hlbiBncm91cHMgZXhpc3Q6CnsKICAiZ3JvdXBzIjogWwogICAgewogICAgICAicGFwZXJfaW5kaWNlcyI6IFsxLCAyXSwKICAgICAgImV2aWRlbmNlIjogewogICAgICAgICIxIjogImNvcHktcGFzdGUgZXhhY3Qgc3BhbiBmcm9tIFBhcGVyIDEgc3RhdGluZyBpdHMgdGVjaG5pY2FsIHByb2JsZW0g4oCUIHZlcmJhdGltIiwKICAgICAgICAiMiI6ICJjb3B5LXBhc3RlIGV4YWN0IHNwYW4gZnJvbSBQYXBlciAyIHN0YXRpbmcgaXRzIHRlY2huaWNhbCBwcm9ibGVtIOKAlCB2ZXJiYXRpbSIKICAgICAgfSwKICAgICAgInJlYXNvbiI6ICIx4oCTMiBzZW50ZW5jZXMgd2h5IHRoZXNlIHBhcGVycyBhZGRyZXNzIHRoZSBzYW1lIGJvdHRsZW5lY2ssIHJlZmVyZW5jaW5nIGV2aWRlbmNlIgogICAgfQogIF0KfQoKV2hlbiBubyBncm91cHMgZXhpc3Q6CnsiZ3JvdXBzIjogW119CgojIElucHV0CjxQQVBFUlM+)\#RoleYouareanexpertresearchanalyst\.Youaregiven<N\_PAPERS\>researchpapers\.Yourtaskistoidentifywhichpapersaddressthesame\*\*problem\*\*andgroupthemaccordingly\.\#DimensionDefinition‘problem‘:thespecifictechnicalchallenge,bottleneck,limitation,orgapthatapaperistryingtosolve\-\-\-theprecisereasonwhyexistingsolutionsareinsufficient\.\-Aproblemdescribes\*\*why\*\*currentapproachesfailorfallshort,notwhatthepaperdoes\(method\)orwhatitappliesto\(task\)\.\-Problemscanbesharedacrossentirelydifferenttasks\(e\.g\.,"distributionshiftattesttime"affectsbothimageclassificationandmachinetranslation\)\.\-Abroadgoalorapplicationcapability\(e\.g\.,"improveaccuracy","scaletolargerdatasets"\)doesNOTconstituteaproblem\.\-Theproblemmustidentifyaconcretetechnicalobstacle,constraint,orgapthatmotivatesthespecificcontribution\.\#GroupingRules\-Placetwoormorepapersinthesamegroupiftheyaddressthesamespecifictechnicalobstacleorbottleneck,eveniftheyoperateondifferenttasksordomains\.\-DoNOTgrouppapersthatonlysharethesametaskorapplicationdomainwithoutasharedbottleneck\.\-DoNOTgrouppaperswhoseproblemsarelooselyrelatedatasurfacelevelbuttargetclearlydifferentfailuremodesorconstraints\.\-Apapercanbelongtoatmostonegroup\.Ifapaperdoesnotshareaproblemwithanyotherpaper,itdoesnotappearinanygroup\.\-Theremaybezero,one,ormultiplegroups\.\-Useonlytheprovidedpapercontentasevidence;donotinventinformation\.\#OutputFormatOutputvalidJSONonly\.Nomarkdown,notextoutsidetheJSONobject\.Whengroupsexist:\{"groups":\[\{"paper\_indices":\[1,2\],"evidence":\{"1":"copy\-pasteexactspanfromPaper1statingitstechnicalproblem\-\-\-verbatim","2":"copy\-pasteexactspanfromPaper2statingitstechnicalproblem\-\-\-verbatim"\},"reason":"1\-\-2sentenceswhythesepapersaddressthesamebottleneck,referencingevidence"\}\]\}Whennogroupsexist:\{"groups":\[\]\}\#Input<PAPERS\>
### P\.6Multi\-paper Grouping: Method Dimension
Grouping Method Prompt[⬇](data:text/plain;base64,IyBSb2xlCllvdSBhcmUgYW4gZXhwZXJ0IHJlc2VhcmNoIGFuYWx5c3QuIFlvdSBhcmUgZ2l2ZW4gPE5fUEFQRVJTPiByZXNlYXJjaCBwYXBlcnMuIFlvdXIgdGFzayBpcyB0byBpZGVudGlmeSB3aGljaCBwYXBlcnMgcHJvcG9zZSBvciB1c2UgdGhlIHNhbWUgKiptZXRob2QqKiBhbmQgZ3JvdXAgdGhlbSBhY2NvcmRpbmdseS4KCiMgRGltZW5zaW9uIERlZmluaXRpb24KYG1ldGhvZGA6IHRoZSBjb3JlIHRlY2huaWNhbCBhcHByb2FjaCwgYWxnb3JpdGhtaWMgaWRlYSwgbW9kZWwgYXJjaGl0ZWN0dXJlLCBvciB0cmFpbmluZy9pbmZlcmVuY2Ugc3RyYXRlZ3kgdGhhdCBhIHBhcGVyIGNvbnRyaWJ1dGVzIG9yIHJlbGllcyBvbi4KLSBBIG1ldGhvZCBkZXNjcmliZXMgKipob3cqKiB0aGUgcGFwZXIgc29sdmVzIGl0cyBwcm9ibGVtIOKAlCB0aGUgc3BlY2lmaWMgdGVjaG5pY2FsIG1lY2hhbmlzbSwgbm90IHRoZSBwcm9ibGVtIGl0c2VsZiBvciB0aGUgdGFzay4KLSBUd28gbWV0aG9kcyBhcmUgc2ltaWxhciBpZiB0aGV5IHNoYXJlIGEgZnVuZGFtZW50YWwgYWxnb3JpdGhtaWMgbWVjaGFuaXNtIG9yIGNvbmNlcHR1YWwgYmFja2JvbmUgYXQgYSBtZWFuaW5nZnVsIGxldmVsIG9mIHNwZWNpZmljaXR5LgotIE1ldGhvZHMgY2FuIGJlIHNoYXJlZCBhY3Jvc3MgZGlmZmVyZW50IHRhc2tzIChlLmcuLCBjb250cmFzdGl2ZSBsZWFybmluZywgUkxIRiwga25vd2xlZGdlIGRpc3RpbGxhdGlvbiBlYWNoIGFwcGVhciBhY3Jvc3MgbWFueSBhcHBsaWNhdGlvbiBhcmVhcykuCi0gQSBicm9hZCB0ZWNobmlxdWUgY2F0ZWdvcnkgKGUuZy4sICJ1c2VzIHRyYW5zZm9ybWVycyIsICJ1c2VzIGRpZmZ1c2lvbiIsICJ1c2VzIHJlaW5mb3JjZW1lbnQgbGVhcm5pbmciKSBpcyBOT1Qgc3VmZmljaWVudCDigJQgcmVxdWlyZSBzaGFyZWQgc3BlY2lmaWMgbWVjaGFuaXNtcy4KCiMgR3JvdXBpbmcgUnVsZXMKLSBQbGFjZSB0d28gb3IgbW9yZSBwYXBlcnMgaW4gdGhlIHNhbWUgZ3JvdXAgaWYgdGhleSBzaGFyZSBhIGZ1bmRhbWVudGFsIHRlY2huaWNhbCBtZWNoYW5pc20sIGNvcmUgYWxnb3JpdGhtaWMgcGlwZWxpbmUsIG9yIHNwZWNpZmljIGFyY2hpdGVjdHVyYWwgZGVzaWduIHByaW5jaXBsZS4KLSAiU2hhcmVkIG1lY2hhbmlzbSIgbWVhbnM6IGlmIHlvdSBkZXNjcmliZWQgdGhlIG1ldGhvZCBvZiBvbmUgcGFwZXIgdG8gYW4gZXhwZXJ0LCB0aGV5IHdvdWxkIHJlY29nbml6ZSB0aGUgY29yZSBpZGVhIGFzIGVzc2VudGlhbGx5IHRoZSBzYW1lIGFzIHRoZSBvdGhlciBwYXBlcidzIG1ldGhvZC4KLSBEbyBOT1QgZ3JvdXAgcGFwZXJzIGJhc2VkIG9uOgogIC0gQm90aCBiZWxvbmdpbmcgdG8gdGhlIHNhbWUgaGlnaC1sZXZlbCB0ZWNobmlxdWUgZmFtaWx5IHdpdGhvdXQgc2hhcmluZyBhIHNwZWNpZmljIG1lY2hhbmlzbSAoZS5nLiwgImJvdGggYXJlIGRpZmZ1c2lvbiBtb2RlbHMiLCAiYm90aCB1c2UgYXR0ZW50aW9uIikuCiAgLSBTaW1pbGFyIHByb2JsZW0gZnJhbWluZyBvciB0YXNrIHNldHRpbmcgd2l0aG91dCBhIHNoYXJlZCB0ZWNobmljYWwgYXBwcm9hY2guCiAgLSBTaGFyZWQgcHJlcHJvY2Vzc2luZywgZXZhbHVhdGlvbiBwcm90b2NvbCwgb3IgZGF0YXNldC4KICAtIEJvdGggZmluZS10dW5pbmcgTExNcyB3aXRob3V0IHNoYXJpbmcgYSBzcGVjaWZpYyBmaW5lLXR1bmluZyBzdHJhdGVneS4KLSBBIHBhcGVyIGNhbiBiZWxvbmcgdG8gYXQgbW9zdCBvbmUgZ3JvdXAuIElmIGEgcGFwZXIgZG9lcyBub3Qgc2hhcmUgYSBtZXRob2Qgd2l0aCBhbnkgb3RoZXIgcGFwZXIsIGl0IGRvZXMgbm90IGFwcGVhciBpbiBhbnkgZ3JvdXAuCi0gVGhlcmUgbWF5IGJlIHplcm8sIG9uZSwgb3IgbXVsdGlwbGUgZ3JvdXBzLgotIFVzZSBvbmx5IHRoZSBwcm92aWRlZCBwYXBlciBjb250ZW50IGFzIGV2aWRlbmNlOyBkbyBub3QgaW52ZW50IGluZm9ybWF0aW9uLgoKIyBPdXRwdXQgRm9ybWF0Ck91dHB1dCB2YWxpZCBKU09OIG9ubHkuIE5vIG1hcmtkb3duLCBubyB0ZXh0IG91dHNpZGUgdGhlIEpTT04gb2JqZWN0LgoKV2hlbiBncm91cHMgZXhpc3Q6CnsKICAiZ3JvdXBzIjogWwogICAgewogICAgICAicGFwZXJfaW5kaWNlcyI6IFsxLCAyXSwKICAgICAgImV2aWRlbmNlIjogewogICAgICAgICIxIjogImNvcHktcGFzdGUgZXhhY3Qgc3BhbiBmcm9tIFBhcGVyIDEgZGVzY3JpYmluZyBpdHMgY29yZSB0ZWNobmljYWwgbWVjaGFuaXNtIOKAlCB2ZXJiYXRpbSIsCiAgICAgICAgIjIiOiAiY29weS1wYXN0ZSBleGFjdCBzcGFuIGZyb20gUGFwZXIgMiBkZXNjcmliaW5nIGl0cyBjb3JlIHRlY2huaWNhbCBtZWNoYW5pc20g4oCUIHZlcmJhdGltIgogICAgICB9LAogICAgICAicmVhc29uIjogIjHigJMyIHNlbnRlbmNlcyBleHBsYWluaW5nIHRoZSBzcGVjaWZpYyBzaGFyZWQgbWVjaGFuaXNtLCByZWZlcmVuY2luZyBldmlkZW5jZSIKICAgIH0KICBdCn0KCldoZW4gbm8gZ3JvdXBzIGV4aXN0Ogp7Imdyb3VwcyI6IFtdfQoKIyBJbnB1dAo8UEFQRVJTPg==)\#RoleYouareanexpertresearchanalyst\.Youaregiven<N\_PAPERS\>researchpapers\.Yourtaskistoidentifywhichpapersproposeorusethesame\*\*method\*\*andgroupthemaccordingly\.\#DimensionDefinition‘method‘:thecoretechnicalapproach,algorithmicidea,modelarchitecture,ortraining/inferencestrategythatapapercontributesorrelieson\.\-Amethoddescribes\*\*how\*\*thepapersolvesitsproblem\-\-\-thespecifictechnicalmechanism,nottheproblemitselforthetask\.\-Twomethodsaresimilariftheyshareafundamentalalgorithmicmechanismorconceptualbackboneatameaningfullevelofspecificity\.\-Methodscanbesharedacrossdifferenttasks\(e\.g\.,contrastivelearning,RLHF,knowledgedistillationeachappearacrossmanyapplicationareas\)\.\-Abroadtechniquecategory\(e\.g\.,"usestransformers","usesdiffusion","usesreinforcementlearning"\)isNOTsufficient\-\-\-requiresharedspecificmechanisms\.\#GroupingRules\-Placetwoormorepapersinthesamegroupiftheyshareafundamentaltechnicalmechanism,corealgorithmicpipeline,orspecificarchitecturaldesignprinciple\.\-"Sharedmechanism"means:ifyoudescribedthemethodofonepapertoanexpert,theywouldrecognizethecoreideaasessentiallythesameastheotherpaper’smethod\.\-DoNOTgrouppapersbasedon:\-Bothbelongingtothesamehigh\-leveltechniquefamilywithoutsharingaspecificmechanism\(e\.g\.,"botharediffusionmodels","bothuseattention"\)\.\-Similarproblemframingortasksettingwithoutasharedtechnicalapproach\.\-Sharedpreprocessing,evaluationprotocol,ordataset\.\-Bothfine\-tuningLLMswithoutsharingaspecificfine\-tuningstrategy\.\-Apapercanbelongtoatmostonegroup\.Ifapaperdoesnotshareamethodwithanyotherpaper,itdoesnotappearinanygroup\.\-Theremaybezero,one,ormultiplegroups\.\-Useonlytheprovidedpapercontentasevidence;donotinventinformation\.\#OutputFormatOutputvalidJSONonly\.Nomarkdown,notextoutsidetheJSONobject\.Whengroupsexist:\{"groups":\[\{"paper\_indices":\[1,2\],"evidence":\{"1":"copy\-pasteexactspanfromPaper1describingitscoretechnicalmechanism\-\-\-verbatim","2":"copy\-pasteexactspanfromPaper2describingitscoretechnicalmechanism\-\-\-verbatim"\},"reason":"1\-\-2sentencesexplainingthespecificsharedmechanism,referencingevidence"\}\]\}Whennogroupsexist:\{"groups":\[\]\}\#Input<PAPERS\>
### P\.7Reasoning Verification \(Stage 3\)
The following prompt is used by the external LLM judge \(GPT\-5\.2 or Gemini\-3\.1\-Pro\) to determine whether a model’s stated reason is logically entailed by its cited evidence strings\. This corresponds to Stage 3 of the cascading evaluation protocol\.
Reasoning Verification Prompt[⬇](data:text/plain;base64,IyBSb2xlCllvdSBhcmUgYSByaWdvcm91cyBsb2dpYyBhdWRpdG9yLiBZb3UgYXJlIGdpdmVuIHR3byBwaWVjZXMgb2YgZXZpZGVuY2UgZXh0cmFjdGVkIHZlcmJhdGltIGZyb20gdHdvIHJlc2VhcmNoIHBhcGVycywgcGx1cyBhIGNvbmNsdXNpb24gYW5kIGEgcmVhc29uIHdyaXR0ZW4gYnkgYW4gYW5hbHlzdC4gWW91ciB0YXNrIGlzIHRvIGp1ZGdlIHdoZXRoZXIgdGhlIHJlYXNvbiBpcyBsb2dpY2FsbHkgc3VwcG9ydGVkIGJ5IHRoZSB0d28gZXZpZGVuY2Ugc3RyaW5ncyBhbG9uZS4KCiMgV2hhdCBZb3UgQXJlIEdpdmVuCi0gKipFdmlkZW5jZSBBKio6IGEgdmVyYmF0aW0gZXhjZXJwdCBmcm9tIFBhcGVyIEEuCi0gKipFdmlkZW5jZSBCKio6IGEgdmVyYmF0aW0gZXhjZXJwdCBmcm9tIFBhcGVyIEIuCi0gKipEaW1lbnNpb24qKjogd2hpY2ggYXNwZWN0IGlzIGJlaW5nIGNvbXBhcmVkIChgdGFza2AsIGBwcm9ibGVtYCwgb3IgYG1ldGhvZGApLgotICoqQ29uY2x1c2lvbioqOiB0aGUgYW5hbHlzdCdzIGJpbmFyeSBqdWRnbWVudCAoYHNpbWlsYXJgIG9yIGBub3Qgc2ltaWxhcmApLgotICoqUmVhc29uKio6IHRoZSBhbmFseXN0J3MgMS0yIHNlbnRlbmNlIGV4cGxhbmF0aW9uIG9mIHdoeSB0aGUgY29uY2x1c2lvbiBob2xkcy4KCiMgRGltZW5zaW9uIEdyYW51bGFyaXR5ClRoZSByZXF1aXJlZCBsZXZlbCBvZiBzcGVjaWZpY2l0eSBkaWZmZXJzIGJ5IGRpbWVuc2lvbjoKLSAqKnRhc2sqKjogQm90aCBwYXBlcnMgbXVzdCB0YXJnZXQgdGhlIHNhbWUgc3BlY2lmaWMgYXBwbGljYXRpb24gc2NlbmFyaW8gb3IgcmVzZWFyY2ggb2JqZWN0aXZlIC0tIG5hcnJvdyBlbm91Z2ggdGhhdCBhIHJlc2VhcmNoZXIgd291bGQgZmlsZSB0aGVtIHVuZGVyIHRoZSBzYW1lIHN1Yi10b3BpYyBoZWFkaW5nLiBOT1Qgc3VmZmljaWVudDogc2FtZSBicm9hZCBmaWVsZC4KLSAqKnByb2JsZW0qKjogQm90aCBwYXBlcnMgbXVzdCBhZGRyZXNzIHRoZSBzYW1lIHNwZWNpZmljIHRlY2huaWNhbCBib3R0bGVuZWNrIG9yIGZhaWx1cmUgbW9kZS4gTk9UIHN1ZmZpY2llbnQ6IHZhZ3VlIHNoYXJlZCBjaGFsbGVuZ2VzLgotICoqbWV0aG9kKio6IEJvdGggcGFwZXJzIG11c3Qgc2hhcmUgYSBzcGVjaWZpYyBhbGdvcml0aG1pYyBtZWNoYW5pc20gb3IgYXJjaGl0ZWN0dXJhbCBkZXNpZ24gcHJpbmNpcGxlLiBOT1Qgc3VmZmljaWVudDogc2FtZSBicm9hZCBwYXJhZGlnbSAoZS5nLiwgImJvdGggdXNlIGF0dGVudGlvbiIpLgoKIyBZb3VyIFRhc2sKSnVkZ2Ugd2hldGhlciBhIHJlYWRlciB3aG8gc2VlcyBvbmx5IEV2aWRlbmNlIEEgYW5kIEV2aWRlbmNlIEIgY2FuIHJlYXNvbmFibHkgYXJyaXZlIGF0IHRoZSBzdGF0ZWQgQ29uY2x1c2lvbiB2aWEgdGhlIHN0YXRlZCBSZWFzb24uCgpTcGVjaWZpY2FsbHksIGFzazoKMS4gRG9lcyBFdmlkZW5jZSBBIGFjdHVhbGx5IHNheSB3aGF0IHRoZSBSZWFzb24gY2xhaW1zIGFib3V0IFBhcGVyIEE/CjIuIERvZXMgRXZpZGVuY2UgQiBhY3R1YWxseSBzYXkgd2hhdCB0aGUgUmVhc29uIGNsYWltcyBhYm91dCBQYXBlciBCPwozLiBEb2VzIHRoZSBSZWFzb24gY29ycmVjdGx5IGNoYXJhY3RlcmlzZSB0aGUgcmVsYXRpb25zaGlwIGJldHdlZW4gdGhlIHR3byBldmlkZW5jZSBzdHJpbmdzPwo0LiBJcyB0aGUgc2ltaWxhcml0eSAqKnNwZWNpZmljIGFuZCBzdWJzdGFudGlhbCoqIGF0IHRoZSBncmFudWxhcml0eSByZXF1aXJlZCBmb3IgdGhlIGdpdmVuIERpbWVuc2lvbj8KCiMgRGVjaXNpb24gUnVsZXMKLSBNYXJrIGBzdXBwb3J0ZWQ6IHRydWVgIGlmIHRoZSBSZWFzb24gZm9sbG93cyBkaXJlY3RseSBmcm9tIHRoZSB0d28gZXZpZGVuY2Ugc3RyaW5ncyBBTkQgbWVldHMgdGhlIGdyYW51bGFyaXR5IGJhci4KLSBNYXJrIGBzdXBwb3J0ZWQ6IGZhbHNlYCBpZjoKICAtIFRoZSBSZWFzb24gYXR0cmlidXRlcyBjb250ZW50IG5vdCBwcmVzZW50IGluIHRoZSBldmlkZW5jZS4KICAtIFRoZSBSZWFzb24gaW50cm9kdWNlcyBmYWN0cyBiZXlvbmQgd2hhdCBlaXRoZXIgZXZpZGVuY2Ugc2F5cy4KICAtIFRoZSBSZWFzb24gY29ycmVjdGx5IGRlc2NyaWJlcyBvbmUgcGFwZXIgYnV0IG1pc3JlcHJlc2VudHMgdGhlIG90aGVyLgogIC0gVGhlIHNoYXJlZCBlbGVtZW50IGlzIHRvbyBnZW5lcmljIGZvciB0aGUgRGltZW5zaW9uLgotIERvIE5PVCB1c2UgeW91ciBvd24ga25vd2xlZGdlIGFib3V0IHRoZSBwYXBlcnMuCgojIE91dHB1dCBGb3JtYXQKT3V0cHV0IHZhbGlkIEpTT04gb25seS4KeyJzdXBwb3J0ZWQiOiB0cnVlL2ZhbHNlLCAiY3JpdGlxdWUiOiAiT25lIHNlbnRlbmNlIGV4cGxhbmF0aW9uLiJ9CgojIElucHV0CkRpbWVuc2lvbjogPERJTUVOU0lPTj4KQ29uY2x1c2lvbjogPENPTkNMVVNJT04+CgpbRXZpZGVuY2UgQV0KPEVWSURFTkNFX0E+CltFdmlkZW5jZSBBIEVORF0KCltFdmlkZW5jZSBCXQo8RVZJREVOQ0VfQj4KW0V2aWRlbmNlIEIgRU5EXQoKUmVhc29uOiA8UkVBU09OPg==)\#RoleYouarearigorouslogicauditor\.Youaregiventwopiecesofevidenceextractedverbatimfromtworesearchpapers,plusaconclusionandareasonwrittenbyananalyst\.Yourtaskistojudgewhetherthereasonislogicallysupportedbythetwoevidencestringsalone\.\#WhatYouAreGiven\-\*\*EvidenceA\*\*:averbatimexcerptfromPaperA\.\-\*\*EvidenceB\*\*:averbatimexcerptfromPaperB\.\-\*\*Dimension\*\*:whichaspectisbeingcompared\(‘task‘,‘problem‘,or‘method‘\)\.\-\*\*Conclusion\*\*:theanalyst’sbinaryjudgment\(‘similar‘or‘notsimilar‘\)\.\-\*\*Reason\*\*:theanalyst’s1\-2sentenceexplanationofwhytheconclusionholds\.\#DimensionGranularityTherequiredlevelofspecificitydiffersbydimension:\-\*\*task\*\*:Bothpapersmusttargetthesamespecificapplicationscenarioorresearchobjective\-\-narrowenoughthataresearcherwouldfilethemunderthesamesub\-topicheading\.NOTsufficient:samebroadfield\.\-\*\*problem\*\*:Bothpapersmustaddressthesamespecifictechnicalbottleneckorfailuremode\.NOTsufficient:vaguesharedchallenges\.\-\*\*method\*\*:Bothpapersmustshareaspecificalgorithmicmechanismorarchitecturaldesignprinciple\.NOTsufficient:samebroadparadigm\(e\.g\.,"bothuseattention"\)\.\#YourTaskJudgewhetherareaderwhoseesonlyEvidenceAandEvidenceBcanreasonablyarriveatthestatedConclusionviathestatedReason\.Specifically,ask:1\.DoesEvidenceAactuallysaywhattheReasonclaimsaboutPaperA?2\.DoesEvidenceBactuallysaywhattheReasonclaimsaboutPaperB?3\.DoestheReasoncorrectlycharacterisetherelationshipbetweenthetwoevidencestrings?4\.Isthesimilarity\*\*specificandsubstantial\*\*atthegranularityrequiredforthegivenDimension?\#DecisionRules\-Mark‘supported:true‘iftheReasonfollowsdirectlyfromthetwoevidencestringsANDmeetsthegranularitybar\.\-Mark‘supported:false‘if:\-TheReasonattributescontentnotpresentintheevidence\.\-TheReasonintroducesfactsbeyondwhateitherevidencesays\.\-TheReasoncorrectlydescribesonepaperbutmisrepresentstheother\.\-ThesharedelementistoogenericfortheDimension\.\-DoNOTuseyourownknowledgeaboutthepapers\.\#OutputFormatOutputvalidJSONonly\.\{"supported":true/false,"critique":"Onesentenceexplanation\."\}\#InputDimension:<DIMENSION\>Conclusion:<CONCLUSION\>\[EvidenceA\]<EVIDENCE\_A\>\[EvidenceAEND\]\[EvidenceB\]<EVIDENCE\_B\>\[EvidenceBEND\]Reason:<REASON\>Similar Articles
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
This paper investigates whether standard benchmarks underestimate LLM performance by re-evaluating hallucination detection datasets using an LLM-first, human-adjudicated assessment method. The study finds that incorporating LLM reasoning into the adjudication process improves agreement and suggests that model-assisted re-evaluation yields more reliable benchmarks for ambiguity-prone tasks.
GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents
GAUGE is a reusable offline protocol that measures the validity of using LLM-as-a-Judge in evaluating task-oriented agents, revealing a satisfaction–success gap and proposing a calibration method to improve accuracy.
On the Limits of LLM-as-Judge for Scientific Novelty Assessment
This paper introduces RQ-Bench, a benchmark to evaluate LLMs' ability to assess the novelty of scientific research questions. It finds that LLM judges consistently rate generated questions as more novel than human experts do, raising concerns about the reliability of using LLMs for scientific novelty evaluation.
GENIE: A Fine-Grained Measure for Novelty
GENIE is a fine-grained evaluation metric that measures the novelty of LLM responses along task-specific features, providing more insight than holistic metrics.
LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization
LongNovel introduces a multi-scale benchmark for evaluating hallucinations in long-context novel summarization, built from Chinese and English novels to improve LLM assessment.