Peerify: Benchmarking Peer-Review Claim Verification
Summary
Peerify is a pipeline for automatically verifying peer-review claims against manuscript evidence, using a benchmark of 800 claims from NeurIPS 2024 and ICLR 2024, demonstrating that retrieval-centered verification outperforms entailment baselines.
View Cached Full Text
Cached at: 09/23/26, 09:06 AM
# Peerify: Benchmarking Peer-Review Claim Verification
Source: [https://arxiv.org/html/2609.25046](https://arxiv.org/html/2609.25046)
###### Abstract
Peer review plays a central role in scholarly publishing, yet verifying whether reviewer claims are supported by manuscript evidence remains a largely manual and time\-consuming process\. We presentPeerify, a pipeline for manuscript\-grounded verification of peer\-review claims\. Given a manuscript and a review comment,Peerifypipeline decomposes reviews into atomic claims, retrieves relevant manuscript evidence, and determines whether each claim is supported by the paper\. To support the development and evaluation of the pipeline, we construct a benchmark of 800 claims derived from authentic peer\-review interactions collected from NeurIPS 2024 and ICLR 2024, including a 300\-claim hand\-labeled subset used to audit the automated supervision\. We evaluate state\-of\-the\-art language models and retrieval strategies withinPeerifypipeline, together with entailment baselines\. Our results demonstrate the importance of retrieval\-centered verification and claim decomposition, while highlighting the challenges posed by ambiguous and interpretive reviewer claims\. Automated labels agree with human consensus on 90\.3% of audited claims \(κ=0\.87\\kappa=0\.87\), while off\-the\-shelf entailment models stay below 0\.24 macro\-F1\.
## 1Introduction
Peer review plays a central role in how scientific knowledge is evaluated and incorporated into the scholarly record and help establish confidence in scientific findings\([Lee et al\., 2013](https://arxiv.org/html/2609.25046#bib.bib23);[Bornmann, 2011](https://arxiv.org/html/2609.25046#bib.bib6);[Smith, 2006](https://arxiv.org/html/2609.25046#bib.bib40);[Arabzadeh et al\., 2024](https://arxiv.org/html/2609.25046#bib.bib4)\)\. As submission volumes grow, editors, area chairs, and program committees must assess more reviews under tighter timelines\. Maintaining review quality and consistency at scale has therefore become an important operational challenge for scholarly publishing\([Ebrahimi et al\., 2025b](https://arxiv.org/html/2609.25046#bib.bib11);[Mulligan et al\., 2013](https://arxiv.org/html/2609.25046#bib.bib32);[Arabzadeh et al\., 2025](https://arxiv.org/html/2609.25046#bib.bib2);[Ebrahimi et al\., 2026](https://arxiv.org/html/2609.25046#bib.bib9)\)\.
A substantial portion of the review process relies on claims made by reviewers which often influence publication outcomes\. However, verifying whether such statements are supported by the contents of the manuscript remains largely a manual process\. In practice, editorial stakeholders rarely have time to systematically examine every review statement against the corresponding paper\. Consequently, unsupported claims and misunderstandings of a manuscript may persist throughout the review process without systematic verification[Ghorbanpour et al\. \(2026\)](https://arxiv.org/html/2609.25046#bib.bib13)\.
The growing adoption of Large Language Models \(LLMs\) within scholarly workflows further amplifies the importance of this problem\([Bender et al\., 2021](https://arxiv.org/html/2609.25046#bib.bib5);[Arabzadeh et al\., 2026b](https://arxiv.org/html/2609.25046#bib.bib3)\)\. LLMs are increasingly used to assist with manuscript preparation, review drafting, and editorial support tasks\([Liang et al\., 2024](https://arxiv.org/html/2609.25046#bib.bib26);[Yuan et al\., 2022](https://arxiv.org/html/2609.25046#bib.bib50);[Sadeghian et al\., 2026](https://arxiv.org/html/2609.25046#bib.bib37)\)\. While these technologies offer substantial opportunities for improving efficiency, they also increase the need for mechanisms that can assess whether statements contained in reviews are faithfully grounded in the manuscript\([Maynez et al\., 2020](https://arxiv.org/html/2609.25046#bib.bib29);[Huang et al\., 2025](https://arxiv.org/html/2609.25046#bib.bib17)\)\. Reliable verification of review claims is therefore becoming an increasingly important component of research integrity within modern publishing ecosystems\.
Although claim verification has been extensively studied in NLP\([Dagan et al\., 2006](https://arxiv.org/html/2609.25046#bib.bib8);[Thorne et al\., 2018](https://arxiv.org/html/2609.25046#bib.bib41);[Wadden et al\., 2020](https://arxiv.org/html/2609.25046#bib.bib44)\), peer\-review claim verification differs in both scope and motivation\. Existing work on LLM\-assisted peer review focuses on review generation or quality assessment\([Yuan et al\., 2022](https://arxiv.org/html/2609.25046#bib.bib50);[Liang et al\., 2024](https://arxiv.org/html/2609.25046#bib.bib26);[Hua et al\., 2019](https://arxiv.org/html/2609.25046#bib.bib16);[Rogers and Augenstein, 2020](https://arxiv.org/html/2609.25046#bib.bib35);[Ebrahimi et al\., 2025a](https://arxiv.org/html/2609.25046#bib.bib10)\)\. However, to the best of our knowledge, verifying whether individual reviewer claims are grounded in the manuscript remains underexplored\. Scientific claims frequently depend on methodological assumptions, dataset characteristics, and experimental conditions that must be interpreted together rather than in isolation\. Furthermore, review statements often combine objectively verifiable assertions with subjective judgments, requiring systems to distinguish between components that can be grounded in evidence and those that reflect reviewer opinion\([Metropolitansky and Larson, 2025](https://arxiv.org/html/2609.25046#bib.bib30);[Pavlick and Kwiatkowski, 2019](https://arxiv.org/html/2609.25046#bib.bib33);[Plank, 2022](https://arxiv.org/html/2609.25046#bib.bib34)\)\. Operationalizing review groundedness thus requires robust document understanding, evidence retrieval, multi\-hop reasoning, and careful treatment of ambiguity\([Pavlick and Kwiatkowski, 2019](https://arxiv.org/html/2609.25046#bib.bib33);[Plank, 2022](https://arxiv.org/html/2609.25046#bib.bib34);[Arabzadeh et al\., 2026a](https://arxiv.org/html/2609.25046#bib.bib1)\)\. We provide a more comprehensive discussion of related work in Appendix[B](https://arxiv.org/html/2609.25046#A2)\.
We identify manuscript\-grounded peer\-review claim verification as an important but underexplored challenge in scholarly publishing\. To address this challenge, we proposePeerify, an end\-to\-end pipeline for peer\-review claim verification\.Peerifypipeline decomposes reviews into atomic, self\-contained claims, retrieves relevant manuscript evidence for each claim, and performs groundedness assessment\.Peerifypipeline was designed with practical deployment requirements in mind, including scalability to large review volumes, transparent evidence attribution, robustness across diverse claim types, and compatibility with existing editorial processes\. We use*production\-ready*in an engineering sense: the pipeline is complete, runs on commodity hardware, and is designed for a human\-in\-the\-loop workflow in whichPeerifytriages reviewer claims and surfaces the supporting passages, and an area chair or author acts on them \(Appendix[E](https://arxiv.org/html/2609.25046#A5)\)\.
To systematically evaluate this task and pipeline, we construct a benchmark of 800 atomic review claims derived from publicly available NeurIPS 2024 and ICLR 2024 submissions, together with their reviews and author–reviewer discussion threads\. The benchmark captures diverse verification scenarios, including explicitly supported claims, claims requiring evidence aggregation, ambiguous claims, and claims that cannot be resolved from the manuscript alone\.
Using this benchmark, we evaluatePeerifypipeline end to end and analyze each of its core components, including claim extraction, evidence retrieval, retrieval configurations, and state\-of\-the\-art language\-model verifiers\. Our results show that retrieval\-centered verification and claim decomposition are important for manuscript\-grounded review verification, but also reveal that ambiguous, interpretive, and partially supported reviewer claims remain challenging for current methods\. We releasePeerifypipeline, the benchmark, labels, prompts, and code to support reproducible research at[https://github\.com/Reviewerly\-Inc/Peerify/](https://github.com/Reviewerly-Inc/Peerify/)\.
Peerifymakes the following contributions:
1. 1\.Task\.We formalize manuscript\-grounded peer\-review claim verification as an intra\-document verification task, where claims from peer reviews are assessed strictly against the reviewed manuscript\.
2. 2\.System\.We proposePeerify, an end\-to\-end verification pipeline that decomposes reviews into atomic claims, retrieves relevant manuscript evidence, and predicts explicit evidence attribution\.
3. 3\.Benchmark and evaluation\.We construct an 800\-claim benchmark from NeurIPS 2024 and ICLR 2024 peer\-review interactions, with six evaluation variants and quality\-controlled slices to evaluatePeerifypipeline, including a 300\-claim human\-audited subset that covers 37\.5% of the benchmark\.
## 2PeerifyPipeline
### 2\.1Problem Definition
LetPPdenote a scientific manuscript andRRdenote an associated peer review\. The objective of review\-groundedness verification is to determine whether verifiable claims contained withinRRare supported by evidence present inPP\. Formally, a reviewRRcan be represented as a collection of atomic claimsC=\{c1,…,cm\}C=\\\{c\_\{1\},\\ldots,c\_\{m\}\\\}\. For each claimcic\_\{i\}, the system identifies a set of supporting evidence candidatesEi=\{e1,…,ek\}E\_\{i\}=\\\{e\_\{1\},\\ldots,e\_\{k\}\\\}from the manuscript and predicts a groundedness labelyi=f\(ci,Ei\)y\_\{i\}=f\(c\_\{i\},E\_\{i\}\)\.
This formulation differs from traditional fact verification and textual entailment settings\([Dagan et al\., 2006](https://arxiv.org/html/2609.25046#bib.bib8);[Thorne et al\., 2018](https://arxiv.org/html/2609.25046#bib.bib41);[Wadden et al\., 2020](https://arxiv.org/html/2609.25046#bib.bib44)\)in several important respects\. First, the evidence space is restricted to a single scientific manuscript rather than an external corpus\. Second, evidence supporting a review claim may be distributed across multiple sections, tables, appendices, and more\. Third, review claims frequently involve methodological choices, experimental design decisions, and interpretations of empirical findings whose validity depends on contextual qualifiers and evidence aggregation\.
### 2\.2Operationalization inPeerify
Peerifyoperationalizes review\-groundedness verification as a three\-stage pipeline:
Step 1: Claim extraction\.Review comments often contain several assertions in a single sentence, mixing factual observations with subjective judgments or recommendations\. Direct verification is difficult because different portions may require distinct evidence\.Peerifytherefore adopts a claim decomposition strategy, breaking each review comment into atomic, self\-contained claims that can be verified independently\. This step preserves the reviewer’s intended meaning while isolating factual assertions to be checked against the manuscript\.
Step 2: Evidence retrieval\.The manuscript is segmented into retrieval units such as paragraphs, figure captions, tables, and appendix passages\. Given a claim,Peerifyretrieves the top\-kkrelevant passages based on semantic similarity between the claim and indexed manuscript segments\. Since supporting evidence may be distributed across multiple locations,Peerifyretrieves a set of evidence candidates rather than a single passage, allowing downstream verification to reason over complementary evidence throughout the document\.
Step 3: Groundedness assessment\.Finally,Peerifyverifies each claim against the retrieved evidence, building on evidence\-based verification paradigms from textual entailment, fact verification, and scientific claim verification\([Dagan et al\., 2006](https://arxiv.org/html/2609.25046#bib.bib8);[Thorne et al\., 2018](https://arxiv.org/html/2609.25046#bib.bib41);[Wadden et al\., 2020](https://arxiv.org/html/2609.25046#bib.bib44)\)\. The verifier jointly reasons over the claim and manuscript evidence to determine the extent of support, assigning one of four labels:Supported,Partially Supported,Not Supported, orNot Determined\. Each prediction is paired with the evidence used for the decision, making the output transparent and traceable for human inspection\.
Table 1:Peerifybenchmark statistics overview\.Dataset\#Papers\#ClaimsPartiallyNotNotICLRNeurIPSRejectAcceptSupportedSupportedSupportedDeterminedNeurIPS 20244236\-\-\-\-\-\-100%86\.0%14\.0%ICLR 20245778\-\-\-\-\-100%\-31\.3%68\.7%PaperSourced\([3\.4](https://arxiv.org/html/2609.25046#S3.SS4)\)10050050000052\.0%48\.0%27\.0%73\.0%RebuttalSourced\([3\.4](https://arxiv.org/html/2609.25046#S3.SS4)\)848002173101898461\.0%39\.0%34\.9%65\.1%LLMJudged\([3\.4](https://arxiv.org/html/2609.25046#S3.SS4)\)848002079314135961\.0%39\.0%34\.9%65\.1%HighAgreement\([3\.4](https://arxiv.org/html/2609.25046#S3.SS4)\)711756039433363\.4%36\.6%37\.1%62\.9%verifiable\-only\([3\.4](https://arxiv.org/html/2609.25046#S3.SS4)\)671315629252164\.9%35\.1%36\.6%63\.4%Human\-Verified\([3\.4](https://arxiv.org/html/2609.25046#S3.SS4)\)7530077107783857\.3%42\.7%34%66%
## 3BenchmarkingPeerify
ThePeerifypipeline requires an evaluation framework that reflects the challenges of real peer\-review workflows\. Existing claim\-verification resources such as FEVER\([Thorne et al\., 2018](https://arxiv.org/html/2609.25046#bib.bib41)\)and SciFact\([Wadden et al\., 2020](https://arxiv.org/html/2609.25046#bib.bib44)\)evaluate claims against external corpora, but do not capture manuscript\-bounded verification, author–reviewer interactions, or the mix of factual and subjective assertions common in peer review \(see Appendix[B](https://arxiv.org/html/2609.25046#A2)\)\. We therefore develop thePeerifybenchmark, which operationalizes review groundedness as an intra\-document verification task and supports the development, validation, and evaluation of review\-groundedness verification systems\. To ensure ecological validity, all claims are derived from authentic peer\-review interactions and evaluated against the manuscripts under review\. The resulting benchmark contains 800 review\-derived claims from publicly available NeurIPS 2024 and ICLR 2024 submissions\. This is at or above standard expert\-annotated claim\-verification sets, with every label grounded in a manuscript \(Appendix[D\.8](https://arxiv.org/html/2609.25046#A4.SS8)\)\.
### 3\.1Source Collection
We collected manuscripts, reviews, author responses, and discussion threads from publicly available submissions in NeurIPS 2024 and ICLR 2024\. These venues provide complete review artifacts, including reviewer comments and author rebuttals, so each instance pairs a manuscript with its review and discussion history and supports multiple sources of evidence about a claim’s validity\. As reported in Table[1](https://arxiv.org/html/2609.25046#S2.T1), the collected papers are well balanced across review outcomes, covering both accepted and rejected submissions\.
### 3\.2Claim Construction
Review comments contain summaries, questions, recommendations, methodological critiques, and subjective opinions\. Since review\-groundedness verification operates at the level of individual factual assertions, raw review text was transformed into atomic claims suitable for independent verification\.
To construct these claims, we evaluated three extraction approaches:Qwen3\-4B\([Yang et al\., 2025](https://arxiv.org/html/2609.25046#bib.bib49)\), prompted to produce atomic, decontextualized assertions; Fenice\([Scirè et al\., 2024](https://arxiv.org/html/2609.25046#bib.bib39)\), which uses dependency parsing and semantic role labeling to identify proposition structures; and Gemma, a general\-purpose instruction\-tuned model prompted to list atomic claims from each passage\.
We run a controlled comparison of the three extractors and adoptQwen3\-4Bfor all benchmark construction, as it gives the best balance of atomicity, decontextualization, and semantic fidelity\. The full evaluation protocol and results are deferred to Appendix[A\.1](https://arxiv.org/html/2609.25046#A1.SS1)\. We also scoredo4\-miniandGPT\-5\-minias extractors: both are stronger on F1 \(0\.895 and 0\.915 against 0\.707\), but the gap is mostly recall rather than precision \(0\.889\), and extraction is the high\-volume stage where API cost and rate limits bite \(Appendix[A\.2](https://arxiv.org/html/2609.25046#A1.SS2)\)\. A human audit of the semantic\-equivalence judge behind these scores gives 89\.3% agreement \(Appendix[A\.3](https://arxiv.org/html/2609.25046#A1.SS3)\)\. Extracted claims were subsequently post\-processed through exact and near\-duplicate removal, string normalization, and filtering of rhetorical questions, speculative statements, and non\-verifiable meta\-commentary\.
### 3\.3Groundedness Annotation
Each extracted claim is assigned one of the four labels ofSupported,Not Supported,Partially SupportedorNot Determined\. Each claim is additionally tagged as*factual*or*subjective*\(automatically, witho4\-mini; criteria in Appendix[A\.5](https://arxiv.org/html/2609.25046#A1.SS5)\), and theverifiable\-onlyslice retains only factual claims\. This ensures groundedness is evaluated on assertions checkable against the manuscript rather than on inherently subjective reviewer opinions\. The split is close to even \(444 objective, 356 subjective\), and objective claims counter\-intuitively need a human tie\-breaker*more*often, 41\.4% against 38\.2% \(Appendix[A\.6](https://arxiv.org/html/2609.25046#A1.SS6)\)\.
### 3\.4Evaluation Variants
Peerifyincludes four core benchmark variants and two quality\-controlled slices, capturing complementary notions of groundedness and label confidence\.
PaperSourced\.PaperSourcedprovides a controlled diagnostic setting in which claims are extracted directly from manuscript content\. These claims are*verifiable by construction*and are expected to be labeledSupported\. We use this subset as a sanity check to isolate errors in retrieval and verification from ambiguity in review\-derived content\.
RebuttalSourced\.This variant derives supervision from realistic author–reviewer discussion threads, assigning each claim a label by interpreting the author response signal \(agreement, correction, clarification, or explicit dispute\)\. Two language models \(GPT\-5\-miniandClaude\-Sonnet\-4\-6\) label each claim independently from the discussion context, with a human arbitrating disagreements to produce the canonical label \(Appendix[A\.10](https://arxiv.org/html/2609.25046#A1.SS10)\)\. This interaction\-grounded supervision, captures how authors themselves judge the correctness of reviewer claims\.
LLMJudged\.This subset provides scalable supervision by judging each claim directly against the manuscript\. The same two labelers \(GPT\-5\-miniandClaude\-Sonnet\-4\-6\) independently assign one of four labels under strict document\-bounded evidence constraints, and a human annotator arbitrates the cases where they disagree to produce the canonical label \(More details in Appendix[A\.10](https://arxiv.org/html/2609.25046#A1.SS10)\)\. We evaluate label quality against theHuman\-Verifiedlabels in Appendix[A\.8](https://arxiv.org/html/2609.25046#A1.SS8)\.
Human\-VerifiedTo provide high\-fidelity reference labels and audit automated supervision, we manually annotated a subset of 300RebuttalSourcedinstances, which is 37\.5% of the benchmark\. Two annotators with formal training in computer science independently labeled each instance following detailed guidelines with concrete examples\. Inter\-annotator reliability was near\-perfect \(Cohen’sκ≈0\.96\\kappa\\approx 0\.96\) for the four\-way nominal task, with disagreements resolved through consensus\. Against these labels the automated supervision is correct on 90\.3% of the 300 claims under strict four\-way exact match \(κ=0\.87\\kappa=0\.87\), 97\.0% allowing a one\-step difference, with hard polarity flips in 1\.0% of cases\. Reliability is flat across claim types, 90\.5% objective against 90\.1% subjective, which is where a labeler\-inherited bias would surface first \(Appendix[A\.9](https://arxiv.org/html/2609.25046#A1.SS9)\)\.
Figure 1:Full\-context verification accuracy across all sixPeerifybenchmark variants\.Quality\-Controlled Sliceswe additionally release slices that improve label reliability\. First, theHighAgreementslice retains only instances where two supervision signals with disjoint evidence,RebuttalSourcedandLLMJudged, agree, yielding a high\-confidence subset less sensitive to labeling noise\. Both signals come from the same two labeler models, so a match is not annotator consensus; what differs is the evidence, sinceRebuttalSourcedreads only the discussion thread andLLMJudgedonly the manuscript, and agreement across disjoint inputs signals claim clarity \(see Limitations\)\. Second, to isolate assertions checkable directly against the manuscript, theverifiable\-onlyslice useso4\-minito filter out subjective or normative claims such as novelty, significance, or impact\. Specifically, the model predicts whether a claim is \(i\) grounded in factual content observable in the manuscript, such as missing experiments or absent figures, or \(ii\) an opinion, recommendation, or forward\-looking judgment not directly verifiable from the paper alone\. Table[1](https://arxiv.org/html/2609.25046#S2.T1)shows the support\-level distributions across different variants\.PaperSourcedis entirelySupported\(500/500\) by construction and serves as a controlled diagnostic\.RebuttalSourcedis far more heterogeneous \(217 Supported, 310 Partially Supported, 189 Not Supported, 84 Not Determined\), reflecting corrections, and unresolved disagreement in author–reviewer exchanges\.LLMJudgedshifts heavily towardNot Determined\(207 Supported, 93 Partially Supported, 141 Not Supported, 359 Not Determined\), indicating that many review statements cannot be decisively validated from the manuscript alone\.
## 4Experimental setup
Table 2:RAG accuracy across five models and four retrievers on all sixPeerifybenchmarks\.ModelRetrieverPaperSourcedRebuttalSourcedLLMJudgedHighAgreementVerifiable\-onlyHumanVerifiedo4\-miniBM250\.8180\.2340\.4910\.5140\.4660\.243BM25\+Reranker0\.8700\.2340\.4960\.5140\.5190\.237Dense0\.8520\.2440\.5140\.5090\.4660\.230Dense\+Reranker0\.9040\.2400\.4990\.4970\.4660\.257GPT\-5\-miniBM250\.7960\.2890\.4120\.5090\.4960\.273BM25\+Reranker0\.8560\.2940\.4170\.4910\.4960\.290Dense0\.8200\.2940\.4140\.5030\.5110\.297Dense\+Reranker0\.8520\.2860\.4050\.4910\.4660\.257Claude\-HaikuBM250\.7860\.2800\.3810\.4400\.3890\.270BM25\+Reranker0\.8220\.2700\.3720\.4510\.4050\.240Dense0\.8080\.2560\.3860\.4290\.3740\.237Dense\+Reranker0\.8440\.2970\.3740\.4800\.4270\.300Qwen3\-8BBM250\.7440\.2740\.3930\.4400\.3820\.257BM25\+Reranker0\.8160\.2810\.3640\.4400\.3890\.297Dense0\.8000\.2870\.3710\.4000\.3210\.283Dense\+Reranker0\.8140\.2650\.3850\.4230\.3740\.277Qwen2\.5\-7BBM250\.8440\.3190\.2110\.2630\.2670\.300BM25\+Reranker0\.8720\.3230\.2110\.2970\.3050\.340Dense0\.8320\.3240\.2090\.2630\.2600\.307Dense\+Reranker0\.8600\.2950\.2280\.2910\.3050\.287
We evaluatePeerifyin two deployment configurations\.*Full\-Context*feeds the whole manuscript and claim to the verifier, serving as an upper bound when the paper fits within the context window; over\-length papers are excluded per model\.*Retrieval\-Augmented*\(RAG\) retrieves evidence per claim and conditions the verifier only on the top\-kkpassages, with or without a cross\-encoder reranker, enabling deployment on long manuscripts\. All prompts are listed in Appendix[C](https://arxiv.org/html/2609.25046#A3)\.
We instantiatePeerifypipeline with five state\-of\-the\-art models spanning proprietary and open\-weight families:o4\-mini\(200K\-context reasoning model\),GPT\-5\-mini\(compact, long\-context\),Claude\-Haiku\-4\.5\(low\-latency, 200K context\), and the open\-weightQwen3\-8BandQwen2\.5\-7B, the last two evaluated mainly under RAG due to context\-window limits\. All models receive identical label definitions and prompts, and outputs are mapped to the four support level labels\.
We evaluate retrieval with recall@kk, nDCG@kk, and MRR onPaperSourced\(BM25 or a dense retriever, optionally reranked by a cross\-encoder; retriever and reranker checkpoints, top\-kk, and chunking are detailed in Appendix[D\.9](https://arxiv.org/html/2609.25046#A4.SS9)\), and verification with Accuracy and Macro\-F1, taking Macro\-F1 as the primary metric since it is robust to label imbalance\. Retrieval metrics are reported only onPaperSourcedbecause gold evidence spans exist only there, so they are inapplicable to the review\-derived variants by construction rather than omitted \(Appendix[D\.11](https://arxiv.org/html/2609.25046#A4.SS11)\)\.
Non\-LLM baselines\.We add three zero\-shot MNLI entailment baselines\([Liu et al\., 2019](https://arxiv.org/html/2609.25046#bib.bib27);[Lewis et al\., 2020a](https://arxiv.org/html/2609.25046#bib.bib24);[He et al\., 2021](https://arxiv.org/html/2609.25046#bib.bib14)\), with the top\-3 retrieved chunks as premise and the claim as hypothesis, run under all four retrieval configurations so a weak result cannot be blamed on one retriever \(Appendix[D\.5](https://arxiv.org/html/2609.25046#A4.SS5)\)\.
## 5Results and Findings
In this section, we evaluatePeerifypipeline from several complementary perspectives\.
### 5\.1Full\-context Verification
Figure[1](https://arxiv.org/html/2609.25046#S3.F1)reports verification accuracy across the six evaluation variants, revealing a clear difficulty gradient\. Performance is highest onPaperSourced, where claims come directly from manuscript content and evidence is explicit;o4\-minireaches 0\.866, and four of five verifiers exceed 0\.75\. Accuracy drops on review\-derived settings such asLLMJudgedandHighAgreement, where claims require reasoning over dispersed evidence and reviewer intent\. Even the strongest verifier reaches only 0\.442 onLLMJudged, showing that review\-groundedness verification goes beyond simple evidence localization\. The hardest setting isRebuttalSourced, where claims come from real author–reviewer interactions and often mix factual observations\.Human\-Verifiedbehaves the same way, with no verifier clearing 0\.38 across its 300 hand\-labeled claims\.
### 5\.2Retrieval\-Centered Verification\.
A key design decision inPeerifyis separating evidence retrieval from groundedness assessment\. Table[2](https://arxiv.org/html/2609.25046#S4.T2)evaluates four retrieval configurations, including BM25, dense retrieval, and their reranked variants\. Retrieval quality substantially affects verification accuracy, but no retriever dominates across all settings\. OnPaperSourced, where evidence is explicit, reranking consistently improves performance\. Dense\+Reranker increaseso4\-minifrom 0\.852 to 0\.904, and BM25\+Reranker increasesGPT\-5\-minifrom 0\.796 to 0\.856\. On review\-derived benchmarks, however, gains are less consistent\. For example, Dense retrieval performs best foro4\-minionLLMJudged\(0\.514\), while BM25\-based retrieval works best forGPT\-5\-mini\(0\.417\) andQwen3\-8B\(0\.393\)\. Similar variability is observed onHighAgreementandHuman\-Verified, indicating that retrieval effectiveness depends not only on evidence quality but also on how each verification engine consumes retrieved context\. Better retrieval therefore helps most when the task requires explicit evidence localization, and much less onRebuttalSourcedandHuman\-Verified, where claims require interpreting reviewer intent and resolving partial support: retrieval is necessary but not sufficient\.
Figure 2:Predicted vs\. ground\-truth \(GT\) labels onRebuttalSourced\. Models exhibit different calibration biases: some over\-predictNot Supported, while others over\-predictPartially Supportedor better match the GT distribution\.Zero\-shot entailment models reach at most 0\.24 macro\-F1, roughly half the weakest LLM verifier\.Swapping the LLM verifier for an off\-the\-shelf NLI model, everything else fixed, does not come close: across three MNLI models and all four retrievers, macro\-F1 never exceeds 0\.24 on any variant, against 0\.45 to 0\.50 for the best LLM configuration on the document\-decidable splits \(Appendix[D\.5](https://arxiv.org/html/2609.25046#A4.SS5)\)\. Two structural causes explain this: three\-way NLI cannot expressPartially Supported, the largest class inRebuttalSourced, and a 512\-token premise limit truncates most of the retrieved evidence\. The retriever moves these numbers by at most two points, so the shortfall belongs to the formulation, which needs long\-context reasoning\.
### 5\.3Error Analysis
The largest error source isRebuttalSourced, where reviewer claims mix factual observations with interpretive judgments, and calibration differences across verification engines \(Figure[2](https://arxiv.org/html/2609.25046#S5.F2)\) show that calibration matters as much as reasoning capability\. A second failure mode emerges onLLMJudgedandHighAgreement, where verification requires multi\-step reasoning over dispersed evidence, limiting even the strongest verifier to 0\.442 and 0\.536 accuracy\. The primary bottlenecks are therefore reasoning over distributed evidence and handling interpretive or partially supported claims, a diagnosis confirmed by full\-context verification, which removes retrieval entirely and still does not lift accuracy on these splits \(Appendix[D\.11](https://arxiv.org/html/2609.25046#A4.SS11)\)\.
Nine in tenPaperSourcederrors are strictness or retrieval, not faulty reasoning\.PaperSourcedis the one split where the gold label is known for every instance without appeal to a judge, which makes it the right place to ask what a residual error actually consists of\. We manually inspected a stratified sample of 150 errors, 30 per verifier under RAG, and assigned each to one of three categories that proved exhaustive\. Overly strict specificity accounts for 83 errors \(55\.3%\): the verifier retrieves the correct passage and then declines to call the claimSupportedbecause one fine detail, an exact count or an appendix pointer, is not restated verbatim\. Retrieval misses account for 54 \(36\.0%\): the supporting passage never enters the top\-kkchunks, so the verifier correctly reports that the evidence in front of it does not contain the claim\. The remaining 13 \(8\.7%\) are malformed outputs that cannot be parsed into a label\.
## 6Conclusion
We presentedPeerify, a pipeline and benchmark for verifying peer reviews\. Our evaluation shows that verifying interpretive claims requires long\-context, four\-way reasoning rather than simple entailment\. Furthermore, retrieval and reasoning create distinct bottlenecks depending on the claim, and raw accuracy is often misleading, making macro\-F1 a necessary evaluation metric\. Our automated labeling proved highly reliable, matching human consensus on 90\.3% of an audited subset\. Designed for a human\-in\-the\-loop editorial workflow,Peerifyaccelerates review verification by triaging questionable claims and surfacing manuscript evidence\. Future extensions include verifying claims against external literature, joint training, and exploring abstention\-aware objectives\. We release the pipeline, benchmark, and all artifacts to support this work\.
## Limitations
Peerifyperformsintra\-documentverification, assessing review claims only against the corresponding manuscript; claims requiring external literature or background knowledge are out of scope\. The benchmark is built from public OpenReview data for NeurIPS 2024 and ICLR 2024, so its distribution may not transfer to other fields or venues with different review norms\. Some variants rely on weak supervision \(author–reviewer discussions and LLM judgments\); we mitigate this with theHighAgreementandverifiable\-onlyslices, but residual ambiguity in review language remains\. Finally, models show strong calibration bias \(e\.g\.,Qwen3\-8BpredictsPartially Supportedfor 60\.5% ofRebuttalSourcedclaims vs\. a 38\.8% ground\-truth rate\), so accuracy alone can mislead; we recommend reporting macro\-F1 alongside accuracy\.
## Ethics Statement
The benchmark uses only publicly accessible OpenReview submissions, reviews, and discussions from NeurIPS 2024 and ICLR 2024; no private or proprietary data are used, and no reviewer or author is identified beyond what is already public\. The 150Human\-Verifiedinstances were labeled by two graduate\-level researchers following explicit guidelines after a calibration round; no crowdsourced or low\-paid labor was involved\.Peerifyis intended as an*assistive*tool that supplements rather than replaces human judgment, and should not be used to rank or target specific reviewers, authors, or submissions\. Because LLM\-based judges and extractors can carry systematic biases \(Section[5](https://arxiv.org/html/2609.25046#S5)\), automated labels should not be treated as ground truth without further validation\. We release all data, labels, prompts, and code under open licenses with documentation of provenance and known limitations\.
## References
- Arabzadeh et al\. \(2026a\)Negar Arabzadeh, Sajad Ebrahimi, Alireza Daghighfarsoodeh, Soroush Sadeghian, Seyed Mohammad Hosseini, Hai Son Le, Mahdi Bashari, and Ebrahim Bagheri\. 2026a\.[From doxa to logos in scientific peer review](https://doi.org/10.1145/3805712.3808494)\.In*Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval*, page 4464–4468\. ACM\.
- Arabzadeh et al\. \(2025\)Negar Arabzadeh, Sajad Ebrahimi, Ali Ghorbanpour, Soroush Sadeghian, Sara Salamat, Muhan Li, Hai Son Le, Mahdi Bashari, and Ebrahim Bagheri\. 2025\.[Building trustworthy peer review quality assessment systems](https://doi.org/10.1145/3746252.3761436)\.In*Proceedings of the 34th ACM International Conference on Information and Knowledge Management*, CIKM ’25, page 6863–6864, New York, NY, USA\. Association for Computing Machinery\.
- Arabzadeh et al\. \(2026b\)Negar Arabzadeh, Sajad Ebrahimi, Soroush Sadeghian, Seyed Mohammad Hosseini, Alireza Daqiq, Hai Son Le, Mahdi Bashari, and Ebrahim Bagheri\. 2026b\.[Can llms uphold research integrity? evaluating the role of llms in peer review quality](https://doi.org/10.1145/3773966.3784970)\.In*Proceedings of the Nineteenth ACM International Conference on Web Search and Data Mining*, WSDM ’26, page 1341–1342, New York, NY, USA\. Association for Computing Machinery\.
- Arabzadeh et al\. \(2024\)Negar Arabzadeh, Sajad Ebrahimi, Sara Salamat, Mahdi Bashari, and Ebrahim Bagheri\. 2024\.[Reviewerly: Modeling the reviewer assignment task as an information retrieval problem](https://doi.org/10.1145/3627673.3679081)\.In*Proceedings of the 33rd ACM International Conference on Information and Knowledge Management*, CIKM ’24, page 5554–5555, New York, NY, USA\. Association for Computing Machinery\.
- Bender et al\. \(2021\)Emily Bender and 1 others\. 2021\.On the dangers of stochastic parrots\.*Proceedings of FAccT*\.
- Bornmann \(2011\)Lutz Bornmann\. 2011\.Scientific peer review\.*Annual Review of Information Science and Technology*\.
- Cohen \(1960\)Jacob Cohen\. 1960\.A coefficient of agreement for nominal scales\.*Educational and Psychological Measurement*, 20\(1\):37–46\.
- Dagan et al\. \(2006\)Ido Dagan, Oren Glickman, and Bernardo Magnini\. 2006\.The pascal recognising textual entailment challenge\.In*Proceedings of the First PASCAL Challenges Workshop on Recognising Textual Entailment*\.
- Ebrahimi et al\. \(2026\)Sajad Ebrahimi, Soroush Sadeghian, Ali Ghorbanpour, Negar Arabzadeh, Sara Salamat, Seyed Mohammad Hosseini, Hai Son Le, Mahdi Bashari, and Ebrahim Bagheri\. 2026\.[Peeriscope: A multi\-faceted framework for evaluating peer review quality](https://doi.org/10.1145/3774905.3793128)\.In*Companion Proceedings of the ACM Web Conference 2026*, WWW Companion ’26, page 168–171, New York, NY, USA\. Association for Computing Machinery\.
- Ebrahimi et al\. \(2025a\)Sajad Ebrahimi, Soroush Sadeghian, Ali Ghorbanpour, Negar Arabzadeh, Sara Salamat, Muhan Li, Hai Son Le, Mahdi Bashari, and Ebrahim Bagheri\. 2025a\.[Rottenreviews: Benchmarking review quality with human and llm\-based judgments](https://doi.org/10.1145/3746252.3761506)\.In*Proceedings of the 34th ACM International Conference on Information and Knowledge Management*, CIKM ’25, pages 5642–5649\.
- Ebrahimi et al\. \(2025b\)Sajad Ebrahimi, Sara Salamat, Negar Arabzadeh, Mahdi Bashari, and Ebrahim Bagheri\. 2025b\.[*exHarmony: Authorship and Citations for Benchmarking the Reviewer Assignment Problem*](https://doi.org/10.1007/978-3-031-88714-7_1), page 1–16\.Springer Nature Switzerland\.
- Evans et al\. \(2023\)Michael Evans, Dominik Soos, Ethan Landers, and Jian Wu\. 2023\.MSVEC: A multidomain testing dataset for scientific claim verification\.In*Proceedings of the 2023 ACM Southeast Conference \(ACM SE\)*, pages 117–123\.
- Ghorbanpour et al\. \(2026\)Ali Ghorbanpour, Soroush Sadeghian, Alireza Daghighfarsoodeh, Sajad Ebrahimi, Negar Arabzadeh, Seyed Mohammad Hosseini, and Ebrahim Bagheri\. 2026\.[Peerispect: Claim verification in scientific peer reviews](https://doi.org/10.1145/3805712.3808368)\.In*Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval*, SIGIR ’26, page 5161–5165, New York, NY, USA\. Association for Computing Machinery\.
- He et al\. \(2021\)Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen\. 2021\.DeBERTa: Decoding\-enhanced BERT with disentangled attention\.In*International Conference on Learning Representations \(ICLR\)*\.
- Hosseini et al\. \(2024\)Mohammad Javad Hosseini, Yang Gao, Tim Baumgärtner, Alex Fabrikant, and Reinald Kim Amplayo\. 2024\.Scalable and domain\-general abstractive proposition segmentation\.*arXiv preprint arXiv:2406\.19803*\.
- Hua et al\. \(2019\)Xinyu Hua, Mitko Nikolov, Nikhil Badugu, and Lu Wang\. 2019\.[Argument mining for understanding peer reviews](https://aclanthology.org/N19-1219/)\.In*Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(NAACL\-HLT\)*, pages 2131–2137\.
- Huang et al\. \(2025\)Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and 1 others\. 2025\.A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions\.*ACM Transactions on Information Systems*, 43\(2\):1–55\.
- Izacard and Grave \(2021\)Gautier Izacard and Edouard Grave\. 2021\.Distilling knowledge from reader to retriever for question answering\.In*Proceedings of ICLR*\.
- Kamalloo et al\. \(2024\)Ehsan Kamalloo, Nandan Thakur, Carlos Lassance, Xueguang Ma, Jheng\-Hong Yang, and Jimmy Lin\. 2024\.[Resources for brewing beir: Reproducible reference models and statistical analyses](https://doi.org/10.1145/3626772.3657862)\.In*Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval*, SIGIR ’24, page 1431–1440, New York, NY, USA\. Association for Computing Machinery\.
- Kang et al\. \(2018\)Dongyeop Kang, Waleed Ammar, Bhavana Dalvi, Madeleine van Zuylen, Sebastian Kohlmeier, Eduard Hovy, and Roy Schwartz\. 2018\.[A dataset of peer reviews \(PeerRead\): Collection, insights and NLP applications](https://doi.org/10.18653/v1/N18-1149)\.In*Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long Papers\)*\. Association for Computational Linguistics\.
- Kotonya and Toni \(2020\)Neema Kotonya and Francesca Toni\. 2020\.Explainable automated fact\-checking for public health claims\.In*Proceedings of EMNLP*\.
- Landis and Koch \(1977\)J\. Richard Landis and Gary G\. Koch\. 1977\.The measurement of observer agreement for categorical data\.*Biometrics*, 33\(1\):159–174\.
- Lee et al\. \(2013\)Carole Lee, Cassidy Sugimoto, Guo Zhang, and Blaise Cronin\. 2013\.Bias in peer review\.*Journal of the American Society for Information Science and Technology*\.
- Lewis et al\. \(2020a\)Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer\. 2020a\.BART: Denoising sequence\-to\-sequence pre\-training for natural language generation, translation, and comprehension\.In*Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics \(ACL\)*, pages 7871–7880\.
- Lewis et al\. \(2020b\)Patrick Lewis, Ethan Perez, Aleksandra Piktus, and 1 others\. 2020b\.Retrieval\-augmented generation for knowledge\-intensive NLP tasks\.In*Proceedings of NeurIPS*\.
- Liang et al\. \(2024\)Weixin Liang, Yuhui Zhang, Hancheng Cao, Binglu Wang, Daisy Ding, Xinyu Yang, Kailas Vodrahalli, and 1 others\. 2024\.[Can large language models provide useful feedback on research papers? a large\-scale empirical analysis](https://arxiv.org/abs/2310.01783)\.*NEJM AI*\.
- Liu et al\. \(2019\)Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov\. 2019\.RoBERTa: A robustly optimized BERT pretraining approach\.*arXiv preprint arXiv:1907\.11692*\.
- Malon \(2018\)Christopher Malon\. 2018\.[Team papelo: Transformer networks at FEVER](https://aclanthology.org/W18-5517/)\.In*Proceedings of the First Workshop on Fact Extraction and Verification \(FEVER\)*\.
- Maynez et al\. \(2020\)Joshua Maynez and 1 others\. 2020\.On faithfulness and factuality in abstractive summarization\.In*ACL*\.
- Metropolitansky and Larson \(2025\)Dasha Metropolitansky and Jonathan Larson\. 2025\.[Towards effective extraction and evaluation of factual claims \(claimify\)](https://arxiv.org/abs/2502.10855)\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(ACL\)*\.
- Min et al\. \(2023\)Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen\-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi\. 2023\.[FActScore: Fine\-grained atomic evaluation of factual precision in long form text generation](https://aclanthology.org/2023.emnlp-main.741/)\.In*Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing \(EMNLP\)*\.
- Mulligan et al\. \(2013\)Adrian Mulligan, Louisa Hall, and Ellen S\. Raphael\. 2013\.[Peer review in a changing world: An international study measuring the attitudes of researchers](https://api.semanticscholar.org/CorpusID:205439770)\.*J\. Assoc\. Inf\. Sci\. Technol\.*, 64:132–161\.
- Pavlick and Kwiatkowski \(2019\)Ellie Pavlick and Tom Kwiatkowski\. 2019\.[Inherent disagreements in human textual inferences](https://aclanthology.org/Q19-1043/)\.*Transactions of the Association for Computational Linguistics*, 7:677–694\.
- Plank \(2022\)Barbara Plank\. 2022\.[The “problem” of human label variation: On ground truth in data, modeling and evaluation](https://doi.org/10.18653/v1/2022.emnlp-main.731)\.In*Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing*, pages 10671–10682, Abu Dhabi, United Arab Emirates\. Association for Computational Linguistics\.
- Rogers and Augenstein \(2020\)Anna Rogers and Isabelle Augenstein\. 2020\.[What can we do to improve peer review in NLP?](https://aclanthology.org/2020.findings-emnlp.112/)In*Findings of the Association for Computational Linguistics: EMNLP 2020*, pages 1256–1262\.
- Saakyan et al\. \(2021\)Arkadiy Saakyan, Tuhin Chakrabarty, and Smaranda Muresan\. 2021\.COVID\-fact: Fact extraction and verification of real\-world claims on COVID\-19 pandemic\.In*Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics \(ACL\)*, pages 2116–2129\.
- Sadeghian et al\. \(2026\)Soroush Sadeghian, Alireza Daghighfarsoodeh, Radin Cheraghi, Sajad Ebrahimi, Negar Arabzadeh, and Ebrahim Bagheri\. 2026\.[Peerprism: Peer evaluation expertise vs review\-writing ai](https://doi.org/10.1145/3805712.3808602)\.In*Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval*, SIGIR ’26, page 3381–3387, New York, NY, USA\. Association for Computing Machinery\.
- Sarrouti et al\. \(2021\)Mourad Sarrouti, Asma Ben Abacha, Yassine M’rabet, and Dina Demner\-Fushman\. 2021\.Evidence\-based fact\-checking of health\-related claims\.In*Findings of the Association for Computational Linguistics: EMNLP 2021*, pages 3499–3512\.
- Scirè et al\. \(2024\)Alessandro Scirè, Karim Ghonim, and Roberto Navigli\. 2024\.[FENICE: Factuality evaluation of summarization based on natural language inference and claim extraction](https://aclanthology.org/2024.findings-acl.841/)\.In*Findings of the Association for Computational Linguistics: ACL 2024*, pages 14148–14161\.
- Smith \(2006\)Richard Smith\. 2006\.Peer review: A flawed process at the heart of science\.*Journal of the Royal Society of Medicine*\.
- Thorne et al\. \(2018\)James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal\. 2018\.FEVER: a large\-scale dataset for fact extraction and verification\.In*Proceedings of NAACL\-HLT*\.
- Tomkins et al\. \(2017\)Andrew Tomkins, Min Zhang, and William D Heavlin\. 2017\.Reviewer bias in single\-versus double\-blind peer review\.*Proceedings of the National Academy of Sciences*, 114\(48\):12708–12713\.
- Ullrich et al\. \(2025\)Herbert Ullrich, Tomáš Mlynář, and Jan Drchal\. 2025\.Claim extraction for fact\-checking: Data, models, and automated metrics\.
- Wadden et al\. \(2020\)David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi\. 2020\.[Fact or fiction: Verifying scientific claims](https://aclanthology.org/2020.emnlp-main.609/)\.In*Proceedings of EMNLP*\.
- Wadden and Lo \(2021\)David Wadden and Kyle Lo\. 2021\.[Overview and insights from the SciVer shared task on scientific claim verification](https://aclanthology.org/2021.sdp-1.16/)\.In*Proceedings of the Second Workshop on Scholarly Document Processing \(SDP\)*, pages 124–129\.
- Williams et al\. \(2018\)Adina Williams, Nikita Nangia, and Samuel Bowman\. 2018\.A broad\-coverage challenge corpus for sentence understanding through inference\.In*Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics \(NAACL\)*, pages 1112–1122\.
- Wright et al\. \(2022\)Dustin Wright, David Wadden, Kyle Lo, Bailey Kuehl, Arman Cohan, Isabelle Augenstein, and Lucy Lu Wang\. 2022\.[Generating scientific claims for zero\-shot scientific fact checking](https://doi.org/10.18653/v1/2022.acl-long.175)\.In*Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 2448–2460, Dublin, Ireland\. Association for Computational Linguistics\.
- Xu et al\. \(2026\)Hang Xu, Ling Yue, Chaoqian Ouyang, Yuchen Liu, Libin Zheng, Shaowu Pan, Shimin Di, and Min\-Ling Zhang\. 2026\.Factreview: Evidence\-grounded reviews with literature positioning and execution\-based claim verification\.*arXiv preprint arXiv:2604\.04074*\.
- Yang et al\. \(2025\)An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, and 1 others\. 2025\.Qwen3 technical report\.*arXiv preprint arXiv:2505\.09388*\.
- Yuan et al\. \(2022\)Weizhe Yuan, Pengfei Liu, and Graham Neubig\. 2022\.[Can we automate scientific reviewing?](https://arxiv.org/abs/2102.00176)*Journal of Artificial Intelligence Research*, 75:171–212\.
## Appendix
##### Appendix Overview\.
The appendix is organized around the three questions a reader is most likely to bring to the paper\.*How was the benchmark built, and can its labels be trusted?*Appendix[A](https://arxiv.org/html/2609.25046#A1)covers the construction pipeline end to end: the extractor comparison and why we run an open\-weight model despite lower F1 \([A\.1](https://arxiv.org/html/2609.25046#A1.SS1),[A\.2](https://arxiv.org/html/2609.25046#A1.SS2)\), the human audit of the semantic\-equivalence judge \([A\.3](https://arxiv.org/html/2609.25046#A1.SS3)\), label definitions and the factual/subjective split \([A\.4](https://arxiv.org/html/2609.25046#A1.SS4),[A\.5](https://arxiv.org/html/2609.25046#A1.SS5),[A\.6](https://arxiv.org/html/2609.25046#A1.SS6)\), and the two reliability analyses that matter most, agreement between the automated labels and human consensus \([A\.9](https://arxiv.org/html/2609.25046#A1.SS9)\) and the two\-labeler arbitration protocol behind the canonical labels \([A\.10](https://arxiv.org/html/2609.25046#A1.SS10)\)\.*What exactly was the model asked?*Appendix[C](https://arxiv.org/html/2609.25046#A3)reproduces every prompt verbatim\.*What do the numbers look like beyond the main tables?*Appendix[D](https://arxiv.org/html/2609.25046#A4)holds the per\-model results \([D\.1](https://arxiv.org/html/2609.25046#A4.SS1),[D\.2](https://arxiv.org/html/2609.25046#A4.SS2),[D\.3](https://arxiv.org/html/2609.25046#A4.SS3),[D\.4](https://arxiv.org/html/2609.25046#A4.SS4)\), the non\-LLM entailment baselines \([D\.5](https://arxiv.org/html/2609.25046#A4.SS5)\), the manual error taxonomy and the calibration effect behind thePaperSourcedordering \([D\.6](https://arxiv.org/html/2609.25046#A4.SS6),[D\.7](https://arxiv.org/html/2609.25046#A4.SS7)\), benchmark size in context \([D\.8](https://arxiv.org/html/2609.25046#A4.SS8)\), retrieval quality and cost \([D\.9](https://arxiv.org/html/2609.25046#A4.SS9),[D\.10](https://arxiv.org/html/2609.25046#A4.SS10)\), and the reason intrinsic retrieval metrics apply to only one split \([D\.11](https://arxiv.org/html/2609.25046#A4.SS11)\)\. Appendix[B](https://arxiv.org/html/2609.25046#A2)gives the extended related work, and Appendix[E](https://arxiv.org/html/2609.25046#A5)closes with the intended deployment setting\.
## Appendix ABenchmark Construction and Annotation
### A\.1Claim Extraction Evaluation
A critical step in thePeerifypipeline is accurately decomposing dense, multi\-faceted reviewer comments into isolated and verifiable statements\. Because reviewer feedback often blends factual assertions with subjective interpretation, the claim extraction module must distill these paragraphs into atomic claims while strictly preserving the original intent of the reviewer\. To determine the most robust approach for this task, table[3](https://arxiv.org/html/2609.25046#A1.T3)compares the performance of three distinct claim extractors \(Fenice\([Scirè et al\., 2024](https://arxiv.org/html/2609.25046#bib.bib39)\), Gemma\([Hosseini et al\., 2024](https://arxiv.org/html/2609.25046#bib.bib15)\), andQwen3\-4B\([Yang et al\., 2025](https://arxiv.org/html/2609.25046#bib.bib49)\)\) across two evaluation axes: semantic alignment with a reference extraction \(LLM\-as\-a\-judge\) and intrinsic structural quality \(reference\-free\)\. The two axes are complementary: reference\-based metrics capture coverage, while reference\-free metrics capture whether individual claims are verification\-ready\. The evaluation focuses on how effectively each model parses authentic review text from our NeurIPS and ICLR 2024 dataset into self\-contained propositions suitable for downstream retrieval and verification\.
#### A\.1\.1Reference\-Based Evaluation\.
To establish a strong comparison signal, we useo4\-minias a high\-quality reasoning model to generate a reference set of claims that are manually verified to be atomic and decontextualized\. We treat this reference extraction as a stronger summarization\-style signal and evaluate how well each candidate extraction method aligns with it\. Agreement is measured using an LLM\-as\-a\-judge protocol\. We employGPT\-5\-minias a semantic judge: for each extracted claim, we identify the most similar reference claim and ask the judge whether the two express the same atomic proposition\. Based on these semantic equivalence decisions, we compute precision, recall, and F1 scores\.
For the reference\-based axis,o4\-miniis used*only*to generate the high\-fidelity reference set against which the three candidate extractors are scored; it is not used to produce any of the 800 benchmark claims\. The benchmark claims are sourced exclusively fromQwen3\-4B, which achieves the best overall balance and is therefore used for all benchmark construction, ensuring extraction quality and reproducibility\.
#### A\.1\.2Reference\-Free Evaluation\.
Regardless of the reference set, a useful claim must meet specific structural and linguistic criteria for downstream verification\([Ullrich et al\., 2025](https://arxiv.org/html/2609.25046#bib.bib43);[Wright et al\., 2022](https://arxiv.org/html/2609.25046#bib.bib47)\)\. We useo4\-minito audit each method’s output against three core dimensions\.Atomicitymeasures whether a claim expresses exactly one checkable assertion\.Faithfulnessassesses whether the claim remains grounded in the original reviewer’s text without introducing hallucinations or scope\-creep\.Decontextualizationchecks whether the claim is self\-contained and does not rely on unresolved references such as “the method” or “Figure 1” that require surrounding context to interpret\.
Table 3:Claim\-extraction quality: LLM\-as\-a\-judge \(GPT\-5\-mini, reference\-based\) and reference\-free intrinsic metrics\.Reference\-BasedReference\-FreeExtractorcountPrec\.Rec\.F1Atom\.Flue\.Deco\.Fenice7390\.8720\.5380\.6650\.3140\.7880\.362Gemma24640\.8680\.7620\.8110\.3790\.7810\.329Qwen3\-4b9230\.8890\.5870\.7070\.4510\.9270\.524
##### Gemma extracts the most claims;Qwen3\-4Bextracts the most usable ones\.
The evaluation reveals distinct operational profiles among the extracted models\. Fenice exhibits a high\-precision, low\-recall regime by extracting 739 claims with a precision of 0\.872 and a recall of 0\.538, effectively minimizing false positives at the expense of comprehensive coverage\. Conversely, Gemma prioritizes recall by achieving a rate of 76\.2% and an F1 score of 0\.811, but it generates a substantially larger volume of 2,464 candidate claims\.Qwen3\-4Boptimizes this balance by producing 923 claims with a precision of 0\.889, thereby maintaining strict semantic control\. In the context of downstream verification, the implications of this tradeoff are asymmetric: while unextracted claims merely reduce total evaluation coverage, poorly formulated claims systematically confound reliable classification\.
##### Intrinsic quality is the discriminating dimension\.
Beyond semantic alignment, intrinsic structural quality emerges as the critical factor determining extractor viability\.Qwen3\-4Bdemonstrates significantly superior performance across essential structural metrics\. Specifically, it achieves an atomicity score of 0\.451, which exceeds the scores of 0\.379 for Gemma and 0\.314 for Fenice\. Furthermore, its fluency score of 0\.927 surpasses both Gemma and Fenice as well\. Finally,Qwen3\-4Brecords a decontextualization score of 0\.524, notably outperforming the respective scores of 0\.329 and 0\.362 achieved by Gemma and Fenice\. These structural properties are foundational to downstream verification accuracy\. Consequently, the high\-fidelity benchmark slices, specificallyVerifiable\-OnlyandHighAgreement, are constructed exclusively usingQwen3\-4Bextractions\. Robust claim extraction is a prerequisite for reliable evaluation, meaning these extraction quality differentials directly determine the diagnostic validity of the resulting benchmark partitions\.
### A\.2Extraction Cost, Latency, and Model Choice
Table[3](https://arxiv.org/html/2609.25046#A1.T3)compares the three open\-weight extractors on quality alone\. That comparison leaves an obvious question open, sinceo4\-miniandGPT\-5\-minialready appear elsewhere in the pipeline: if those models are good enough to generate the reference claims and to judge semantic equivalence, why not use one of them to extract? Table[4](https://arxiv.org/html/2609.25046#A1.T4)answers it with measurements instead of an assertion, scoring all five candidates under the same protocol and adding median latency per review and average API cost per review\.
Table 4:Claim\-extraction quality, latency, and cost for all five candidate extractors, scored against the same reference claim set with theGPT\-5\-minisemantic\-equivalence judge\. Latency is the median wall\-clock time per review, measured locally on a single NVIDIA RTX 3090 for open\-weight models and over the network for API models\. Cost is the average API spend per review\.†GPT\-5\-minialso generated the reference claims, so its F1 is a self\-consistency upper bound rather than a comparable score;o4\-miniis the fair proprietary reference point\.ExtractorPrec\.Rec\.F1Latency \(s\)CostFenice0\.8720\.5380\.6651\.97open\-weightQwen3\-4B0\.8890\.5870\.7078\.20open\-weightGemma0\.8680\.7620\.81121\.76open\-weighto4\-mini0\.9900\.8170\.8959\.62$0\.29GPT\-5\-mini0\.9940\.8490\.915†24\.43$0\.11
The proprietary models are clearly better extractors\.o4\-minireaches 0\.895 F1 andGPT\-5\-mini0\.915, against 0\.707 forQwen3\-4B\. We note that theGPT\-5\-minifigure is not directly comparable, because the same model generated the reference claims, so 0\.915 is a self\-consistency ceiling;o4\-miniat 0\.895 is the honest cross\-model number\.
We still extract withQwen3\-4B, for three reasons that the table makes concrete\. First, cost scales with volume, and extraction is the highest\-volume stage: it runs over every sentence of every review, whereas verification runs once per extracted claim\. At $0\.11 to $0\.29 per review, extracting a full conference cycle through an API is a real budget line and a rate\-limit exposure, and it puts the expensive model at the cheap end of the pipeline\. We would rather spend that budget on verification, which is where the reasoning difficulty actually is\. Second, reproducibility: an open\-weight extractor lets anyone regenerate or extend the benchmark, whereas a versioned API can be updated or deprecated underneath a published dataset\. Third, the quality gap sits mostly in recall \(0\.587 against 0\.817\), not precision \(0\.889 against 0\.990\)\. For a benchmark that evaluates claim*verification*rather than exhaustive claim*mining*, missing a claim reduces coverage while a badly formed claim corrupts a verification unit, so precision is the axis that matters and the local model is close to the frontier there\.
### A\.3Human Validation of the Semantic\-Equivalence Judge
The reference\-based scores in Table[3](https://arxiv.org/html/2609.25046#A1.T3)rest onGPT\-5\-minideciding whether an extracted claim and a reference claim express the same proposition\. Since that decision drives which extractor we adopt, we validate it against human annotation like every other automated component of the benchmark\. An annotator who had not worked on the extraction pipeline independently re\-decided 50 randomly sampled equivalence judgements per extractor, 150 in total, seeing the claim pair but not the judge’s verdict\.
Table 5:Human validation of theGPT\-5\-minisemantic\-equivalence judge used for the reference\-based extraction scores\. An annotator who was not involved in the automated pipeline independently re\-decided 50 randomly sampled equivalence judgements per extractor\.ExtractorAgreementRateFenice45 / 5090\.0%Gemma42 / 5084\.0%Qwen3\-4B47 / 5094\.0%Overall134 / 15089\.3%
Agreement is 89\.3% overall and is consistent across extractors, from 84\.0% on Gemma to 94\.0% onQwen3\-4B\(Table[5](https://arxiv.org/html/2609.25046#A1.T5)\)\. This is close to the rate at which our groundedness labels match human consensus \(Appendix[A\.9](https://arxiv.org/html/2609.25046#A1.SS9)\), which suggests the judge is a reasonable stand\-in for a human on this narrow decision, and it supports applying it across the whole extraction evaluation rather than reporting a hand\-checked subset\. The Gemma figure is the lowest of the three, which is consistent with its output profile: it produces 2,464 claims, many of them long or under\-decontextualized, and those are exactly the pairs where equivalence is a judgement call rather than a lookup\.
### A\.4Label Definitions and Examples
Each review claim receives one of the four labels defined in Section[2\.1](https://arxiv.org/html/2609.25046#S2.SS1), determined by whether the manuscript provides sufficient evidence under a strict document\-bounded grounding criterion\. Worked examples follow\.
- •Supported: The manuscript provides clear and sufficient evidence that fully substantiates the claim\.*Example:*A reviewer states “The authors evaluate on CIFAR\-100,” and the paper explicitly reports CIFAR\-100 experiments\.
- •Not Supported: The manuscript contradicts the claim or the asserted content is demonstrably absent\.*Example:*A reviewer states “No ablation study is provided,” but the paper contains a dedicated ablation section\.
- •Partially Supported: The manuscript supports only a qualified or incomplete version of the claim\.*Example:*A reviewer states “The method outperforms all baselines,” but the paper shows it outperforms most but not all\.
- •Not Determined: The manuscript does not contain sufficient information to resolve the claim in either direction\. This label applies when \(i\) the claim references external knowledge or prior work not discussed in the paper, \(ii\) the claim is too vague or ambiguous to verify, or \(iii\) the relevant evidence is absent without any contradicting assertion\.
### A\.5Factual vs\. Subjective Claims
Each claim is tagged as*factual*or*subjective*usingo4\-mini:*factual*if its validity is grounded in content observable in the manuscript \(e\.g\., a missing experiment or an absent figure\), and*subjective*if it is an opinion, recommendation, or forward\-looking judgment about novelty, significance, or impact that cannot be resolved from the paper alone\. Theverifiable\-onlyslice retains only factual claims, so verification on that slice targets assertions checkable against the manuscript rather than reviewer opinion\.
### A\.6Claim Types and Arbitration
We usedGPT\-5\-minito sort the 800 benchmark claims into objective claims, whose truth follows from something written in the manuscript, and subjective claims, which are judgements about novelty, significance, or direction\. The split is 444 objective and 356 subjective\. We then measured how often each type required human arbitration, meaning the two automated labelers disagreed and a person had to break the tie\.
Table 6:Arbitration onRebuttalSourcedbroken down by claim type\.*Arbitrated*counts claims where the two automated labelers disagreed and a human tie\-breaker was needed\.*Neither*counts disagreements where the human annotator supplied a third label because both candidates were wrong\.Claim typeTotalArbitratedArbitration rateObjective44418441\.4%Subjective35613638\.2%Overall80032040\.0%
Arbitration outcomeObjectiveSubjectiveLabelers agree \(no arbitration\)260 \(58\.6%\)220 \(61\.8%\)GPT\-5\-miniupheld93 \(21\.0%\)65 \(18\.3%\)Claude\-Sonnet\-4\-6upheld79 \(17\.8%\)57 \(16\.0%\)Neither upheld12 \(2\.7%\)14 \(3\.9%\)
The result runs against intuition\. Objective claims are arbitrated*more*often than subjective ones, 41\.4% against 38\.2% \(Table[6](https://arxiv.org/html/2609.25046#A1.T6)\)\. The explanation is that objectivity in a peer review is not the same as ease of verification\. Settling “the paper does not compare against baseline X” or “the reported IoU is 44\.91” requires finding one specific fact somewhere in a long manuscript, and two labelers that retrieve different passages will reach different conclusions\. Subjective claims are vaguer, but that vagueness gives the labelers less to disagree about: both tend to land onPartially SupportedorNot Determined\.
The arbitration outcomes point the same way\.GPT\-5\-miniis upheld somewhat more often thanClaude\-Sonnet\-4\-6in both categories \(21\.0% against 17\.8% on objective claims, 18\.3% against 16\.0% on subjective ones\), so neither labeler dominates and the arbitration is not a systematic correction of one model\. Cases where the human annotator had to supply a third label because both candidates were wrong are rare, 2\.7% on objective claims and 3\.9% on subjective ones\. The slightly higher rate on subjective claims is the expected direction: when a reviewer’s remark is genuinely interpretive, there is sometimes no label either model would have proposed\.
### A\.7Benchmark Label Distributions
Table[1](https://arxiv.org/html/2609.25046#S2.T1)in Section[2](https://arxiv.org/html/2609.25046#S2)reports the fullPeerifybenchmark statistics overview, including label distributions across all six benchmark variants and dataset provenance \(venue and acceptance decision\)\. ThePaperSourcedvariant is all\-Supported by construction; the remaining variants show progressively more heterogeneous label distributions as claims become more interaction\-grounded and ambiguous\. The fourHuman\-Verifiedlabel counts are the human\-consensus totals of Table[7](https://arxiv.org/html/2609.25046#A1.T7)\.
### A\.8Inter\-Source Agreement
TheHighAgreementslice is constructed by retaining only instances whereRebuttalSourcedandLLMJudgedlabels agree\. This section describes the agreement analysis underlying this design choice\.
##### Automated supervision sources\.
RebuttalSourcedlabels are derived from author–reviewer interaction dynamics: the model infers whether the author’s response confirms or refutes the reviewer’s claim\.LLMJudgedlabels are derived from direct document\-bounded verification: the model checks whether the manuscript supports the claim without consulting the discussion thread\. These two sources are therefore*independent*in both mechanism and evidence source, making their agreement a meaningful signal of claim clarity and label reliability\.
Across the 800 core review\-derived claims, the canonicalRebuttalSourcedandLLMJudgedlabels agree on 175 instances \(21\.9%\), forming theHighAgreementslice\. The relatively low raw agreement rate reflects the genuine difficulty of peer\-review claims and the different perspectives captured by the two supervision sources: interaction\-grounded labels capture how authors interpret their own work, while document\-grounded labels capture what the manuscript literally supports\.
##### Automated supervision vs\. human labels\.
To validate automated supervision quality, we compare bothRebuttalSourcedandLLMJudgedlabels against theHuman\-Verifiedconsensus labels\. Inter\-annotator agreement between the two human annotators is near\-perfect \(Cohen’sκ≈0\.96\\kappa\\approx 0\.96\), indicating that the high human–automated agreement is not an artifact of low human reliability\. These results support the use ofRebuttalSourcedandLLMJudgedas reliable large\-scale supervision sources, withHuman\-Verifiedserving as a gold\-standard reference for validation and error analysis\. Appendix[A\.9](https://arxiv.org/html/2609.25046#A1.SS9)gives the full confusion matrix and a per\-claim\-type breakdown\.
### A\.9Human\-Verified Label Quality
Most of the benchmark is labeled automatically, so the question that matters is not whether those labels look plausible but how far they can be trusted\. TheHuman\-Verifiedsubset exists to answer it: 300 claims, 37\.5% of the benchmark, labeled by hand under the guidelines of Section[3\.4](https://arxiv.org/html/2609.25046#S3.SS4)and used as the reference against which the automated supervision is scored\.
Table 7:Confusion matrix between automated \(LLMJudged\) labels and human consensus labels on the expandedHuman\-Verifiedsubset \(n=300n=300\)\. Rows are automated labels, columns are human labels\. Exact four\-way agreement is the diagonal,271/300=90\.3%271/300=90\.3\\%\(Cohen’sκ=0\.87\\kappa=0\.87\)\.HumanAutomatedSupp\.Part\. Supp\.Not Supp\.Not Det\.Supported65211Partially Supported810270Not Supported23692Not Determined20135
##### Automated labels match human consensus on 90\.3% of claims \(κ=0\.87\\kappa=0\.87\)\.
Under a strict four\-way exact\-match criterion, the automated labels agree with human consensus on 271 of 300 claims, 90\.3%, with Cohen’sκ=0\.87\\kappa=0\.87\([Cohen, 1960](https://arxiv.org/html/2609.25046#bib.bib7)\), which falls in the “almost perfect” band on the conventional scale\([Landis and Koch, 1977](https://arxiv.org/html/2609.25046#bib.bib22)\)\. Table[7](https://arxiv.org/html/2609.25046#A1.T7)shows where the 29 disagreements sit, and their shape matters more than their count\. They cluster on thePartially Supportedboundary: 8 claims the humans calledSupportedwere labeledPartially Supported, and 7 they calledNot Supportedwere labeled the same way\. Under an ordinal\-tolerant criterion, where a one\-step difference on theSupported,Partially Supported,Not Supportedscale is acceptable, agreement rises to 97\.0%\. Hard polarity flips, where one side says clearly supported and the other says clearly unsupported, occur 3 times in 300, or 1\.0%\. The automated supervision is therefore not making a different kind of judgement from the humans; it draws the partial\-support line in a slightly different place, which is the same place two human annotators need a consensus discussion\.
Table 8:Agreement between automated labels and human consensus on the expandedHuman\-Verifiedsubset, broken down by claim type\. Reliability is essentially unchanged between objective and subjective claims, which is the case an annotation\-bias account would predict to diverge\.SlicenAgreementCohen’sκ\\kappaOverall30090\.3%0\.866Objective / factual15890\.5%0\.866Subjective14290\.1%0\.858
##### Reliability is the same on objective and subjective claims\.
A bias inherited from the labeler models would most plausibly show up as degraded reliability on exactly the claims where reviewer language is interpretive\. It does not\. Table[8](https://arxiv.org/html/2609.25046#A1.T8)splits the subset by claim type: 90\.5% agreement andκ=0\.866\\kappa=0\.866on the 158 objective claims, 90\.1% andκ=0\.858\\kappa=0\.858on the 142 subjective ones\. A 0\.4\-point difference at this sample size is not a signal\. The automated labels track human judgement about equally well on both, which is the evidence we can offer that they reflect claim groundedness rather than an artifact of the annotation procedure\. It does not prove the absence of a bias that the models and the guidelines share, and we say so in the Limitations section\.
Table 9:RAG accuracy on the expandedHuman\-Verifiedsubset \(n=300n=300\), by retriever\. These are the values reported in theHuman\-Verifiedcolumn of Table[2](https://arxiv.org/html/2609.25046#S4.T2)\. No model clears0\.350\.35under any retriever, and swapping retrievers moves accuracy by at most a few points, which is the pattern we use to argue that reasoning rather than retrieval is the binding constraint on this split\.ModelBM25BM25\+RDenseDense\+Ro4\-mini0\.2430\.2370\.2300\.257GPT\-5\-mini0\.2730\.2900\.2970\.257Claude\-Haiku0\.2700\.2400\.2370\.300Qwen3\-8B0\.2570\.2970\.2830\.277Qwen2\.5\-7B0\.3000\.3400\.3070\.287
##### No verifier clears 0\.38 on the hand\-labeled claims\.
Human\-Verifiedis among the hardest splits in the benchmark\. No verifier clears 0\.38 in full\-context \(Table[14](https://arxiv.org/html/2609.25046#A4.T14)\) or 0\.34 under any retriever \(Table[9](https://arxiv.org/html/2609.25046#A1.T9)\)\.Qwen3\-8Bis the strongest full\-context model here at 0\.377, which is consistent with the subset being dominated byPartially Supportedclaims \(107 of 300\) and with that model’s tendency to over\-predict exactly that label \(Figure[2](https://arxiv.org/html/2609.25046#S5.F2)\)\. Because the human labels are the reference rather than a proxy, these numbers are the cleanest available estimate of how far current verifiers are from the task\.
### A\.10Annotator Disagreement and Arbitration
The canonical labels forRebuttalSourcedandLLMJudgedare produced by a two\-annotator protocol with human arbitration\. Each of the 800 atomic reviewer claims is labeled twice and independently:GPT\-5\-mini\(Annotator A\) andClaude\-Sonnet\-4\-6\(Annotator B\) receive the same claim and the same source material \(the author response forRebuttalSourcedand the paper content forLLMJudged\) under identical prompts\. For every record where the two annotators disagree, a human annotator is shown the claim, the relevant source material, and both candidate labels \(with model identity anonymized and presentation order randomized per record\), and either selects the correct label or assigns a new one when both candidates are wrong\. This arbitrated label is taken as canonical\.
##### Inter\-annotator agreement before arbitration\.
Table[10](https://arxiv.org/html/2609.25046#A1.T10)reports raw agreement between the two annotators\.Claude\-Sonnet\-4\-6is systematically more conservative thanGPT\-5\-mini: onRebuttalSourcedit reclassifies manySupportedcalls asPartially SupportedorNot Supported, and onLLMJudgedit shifts confident labels intoNot Determined\. The two models agree almost entirely onNot Determineditself \(52/67 onRebuttalSourced, 193/251 onLLMJudged\) but diverge on the confident labels\.
Table 10:Inter\-annotator agreement \(GPT\-5\-minivs\.Claude\-Sonnet\-4\-6\) before arbitration\.BenchmarknAnnotators agreeRateRebuttalSourced80048060\.0%LLMJudged80045957\.4%
##### Arbitration verdicts\.
Table[11](https://arxiv.org/html/2609.25046#A1.T11)shows, among the disagreements, how often each annotator was upheld by the human annotator\.GPT\-5\-miniis correct slightly more often \(∼\\sim49%\) thanClaude\-Sonnet\-4\-6\(∼\\sim41%\), but in roughly one in ten disagreements neither candidate was correct and the annotator supplied a third label\.
Table 11:Quantitative breakdown of arbitration verdicts on annotator disagreements\. Annotator A and B denote theGPT\-5\-miniandClaude\-Sonnet\-4\-6, respectively\.Benchmark\#DisagreeA correctB correctNeitherRebuttalSourced320158 \(49\.4%\)136 \(42\.5%\)26 \(8\.1%\)LLMJudged341168 \(49\.3%\)137 \(40\.2%\)35 \(10\.3%\)
##### Canonical distribution and derived slices\.
Table[12](https://arxiv.org/html/2609.25046#A1.T12)gives the arbitrated label distribution\. TheLLMJudgeddistribution is heavily weighted towardNot Determined, reflecting that a meaningful share of reviewer claims have no verifiable answer in the paper text alone\. After arbitration, theHighAgreementslice \(canonicalRebuttalSourcedlabel equals canonicalLLMJudgedlabel\) contains 175 instances, and theverifiable\-onlyslice contains 131\.
Table 12:Canonical label distribution \(post\-arbitration\)\.LabelRebuttalSourcedLLMJudgedSupported217207Partially Supported31093Not Supported189141Not Determined84359
## Appendix BRelated Works
##### LLMs for Peer Review\.
A growing body of work explores the use of large language models to support scholarly peer review through tasks such as review generation, review assistance, reviewer recommendation, and manuscript assessment\([Yuan et al\., 2022](https://arxiv.org/html/2609.25046#bib.bib50);[Liang et al\., 2024](https://arxiv.org/html/2609.25046#bib.bib26);[Tomkins et al\., 2017](https://arxiv.org/html/2609.25046#bib.bib42)\)\. A related line analyzes peer\-review content and dynamics, including review quality, argumentation, sentiment, and reviewer behavior\([Hua et al\., 2019](https://arxiv.org/html/2609.25046#bib.bib16);[Rogers and Augenstein, 2020](https://arxiv.org/html/2609.25046#bib.bib35);[Kang et al\., 2018](https://arxiv.org/html/2609.25046#bib.bib20);[Ebrahimi et al\., 2025a](https://arxiv.org/html/2609.25046#bib.bib10)\)\. These efforts demonstrate that state\-of\-the\-art language models can generate meaningful feedback and shed light on how reviews are written and how they function socially\. However, their primary objective is to produce, evaluate, or characterize reviews rather than to determine whether specific claims made within a review are supported by evidence contained in the reviewed manuscript\. As a result, the problem of manuscript\-grounded review verification remains largely unexplored\.
##### Scientific Claim Verification\.
Peerifyis also related to scientific claim verification and evidence\-based fact checking, studied extensively through textual entailment and natural language inference, from early RTE\-style formulations\([Dagan et al\., 2006](https://arxiv.org/html/2609.25046#bib.bib8)\)to large\-scale benchmarks such as FEVER\([Thorne et al\., 2018](https://arxiv.org/html/2609.25046#bib.bib41)\), SciFact\([Wadden et al\., 2020](https://arxiv.org/html/2609.25046#bib.bib44)\), PubHealth\([Kotonya and Toni, 2020](https://arxiv.org/html/2609.25046#bib.bib21)\), and scientific verification tasks\([Malon, 2018](https://arxiv.org/html/2609.25046#bib.bib28);[Wadden and Lo, 2021](https://arxiv.org/html/2609.25046#bib.bib45)\)\. These datasets have enabled systems that retrieve evidence and determine whether a claim is supported, contradicted, or unsupported, and retrieval\-augmented approaches\([Lewis et al\., 2020b](https://arxiv.org/html/2609.25046#bib.bib25);[Izacard and Grave, 2021](https://arxiv.org/html/2609.25046#bib.bib18)\)have further improved evidence\-grounded and multi\-hop reasoning\. While these settings share similarities with review verification, they typically treat claims as standalone statements and assume open\-domain retrieval or citation\-based reasoning across multiple sources, with evaluation centered on short claims whose supporting evidence is relatively concentrated\. In contrast, peer\-review claims often contain methodological observations, novelty assessments, and experimental critiques whose supporting evidence may be distributed across multiple sections of a single manuscript\. Review verification is therefore a document\-bounded grounding problem that requires reasoning over long, highly structured scientific documents while accounting for contextual qualifiers, experimental assumptions, and evidence aggregation\.
##### Claim Decomposition\.
A further challenge arises from the fact that review comments frequently contain multiple assertions expressed within a single sentence or paragraph\. Recent work on claim decomposition has shown that breaking complex statements into atomic claims improves the reliability and interpretability of downstream verification\([Metropolitansky and Larson, 2025](https://arxiv.org/html/2609.25046#bib.bib30);[Min et al\., 2023](https://arxiv.org/html/2609.25046#bib.bib31);[Scirè et al\., 2024](https://arxiv.org/html/2609.25046#bib.bib39)\)\. Peer reviews complicate this step because they often mix subjective assessment with factual assertions and rely on hedging or implicit references\.Peerifyadopts a similar atomic\-claim perspective and adapts extraction using multiple supervision signals, enabling more precise evidence attribution and systematic, manuscript\-grounded groundedness assessment in an end\-to\-end evaluation\.
##### Concurrent Work\.
A closely related concurrent system is FactReview\([Xu et al\., 2026](https://arxiv.org/html/2609.25046#bib.bib48)\), which also targets evidence\-grounded peer review assessment\. FactReview focuses on author\-written claims extracted from the submitted manuscript, using literature retrieval and code execution to verify whether reported results are reproducible and positioned correctly relative to prior work\.Peerifyis complementary in scope and emphasis: we focus on*reviewer*\-authored claims, statements made in the review about the paper rather than by the paper, and derive supervision from author–reviewer interaction dynamics and multiple independent label sources\. This distinction matters because reviewer claims introduce interpretive and normative language absent from author\-written text, and their groundedness must be assessed against a single manuscript rather than the broader literature\. Our multi\-benchmark design, spanning six supervision variants with different reliability and ambiguity profiles, enables systematic evaluation of the task difficulty spectrum\.
## Appendix CPrompts
This section lists every prompt used in the pipeline verbatim\. Figure[3](https://arxiv.org/html/2609.25046#A3.F3)gives the groundedness verification prompt shared by all five verifiers, Figure[4](https://arxiv.org/html/2609.25046#A3.F4)the claim\-extraction prompt, and Figure[5](https://arxiv.org/html/2609.25046#A3.F5)the two labeling prompts behindRebuttalSourcedandLLMJudged\.
### C\.1Verification Prompt Template
We use the following prompt template for all LLM verifiers in full\-context mode\. The same label definitions are used across all dataset variants\. For RAG\-based verification, the\{manuscript\}field is replaced with the concatenation of top\-kkretrieved passages, ranked by the chosen retrieval configuration\.
Groundedness Verification Prompt \(full\-context and RAG\)You are a scientific claim verifier\. Given a peer review claim and the full text of the reviewed manuscript, determine whether the manuscript supports the claim\.Claim:\{claim\}Manuscript:\{manuscript\}Classify the claim as exactly one of the following labels based solely on the content of the manuscript:•Supported:The manuscript provides clear and sufficient evidence that fully substantiates the claim\.•Not Supported:The manuscript contradicts the claim, or the asserted content is demonstrably absent from the manuscript\.•Partially Supported:The manuscript supports only a qualified or incomplete version of the claim \(e\.g\., missing conditions, limited scope, or omitted caveats\)\.•Not Determined:The manuscript does not contain sufficient information to resolve the claim in either direction\. Use this label when \(1\) the claim references external knowledge, \(2\) the claim is too vague or ambiguous to verify, or \(3\) the relevant evidence is absent without any contradicting assertion\.Figure 3:Groundedness verification prompt, shared by all five verifiers\. Under RAG the\{manuscript\}field carries the top\-kkretrieved passages instead of the full paper\.
### C\.2Claim Extraction Prompt
The prompted extractor \(Qwen3\-4B\) decomposes each review into atomic claims with the prompt in Figure[4](https://arxiv.org/html/2609.25046#A3.F4)\.
Claim Extraction Prompt \(Qwen3\-4B\)System:You are an expert at extracting factual claims from academic reviews\.User:Extract factual claims from the review text below\. A claim is a specific, verifiable statement about the paper\.Review Text:\{review\_text\}Each claim should be \(1\) a single atomic statement, \(2\) verifiable against the paper or author response, and \(3\) specific and factual \(not an opinion\)\.Figure 4:Claim\-extraction prompt used to decompose each review comment into atomic, self\-contained claims\.Table 13:PaperSourcedDense\+R accuracy vs\. the Oracle upper bound \(gold paragraph fed directly\) for all five models\. Full four\-retriever results in Table[2](https://arxiv.org/html/2609.25046#S4.T2)\.ModelDense\+RerankerOracleo4\-mini0\.9040\.988GPT\-5\-mini0\.8520\.958Claude\-Haiku0\.8440\.860Qwen3\-8B0\.8140\.872Qwen2\.5\-7B0\.8600\.816
### C\.3Labeling Prompts
The two automated annotators \(GPT\-5\-miniandClaude\-Sonnet\-4\-6\) use the prompts in Figure[5](https://arxiv.org/html/2609.25046#A3.F5)for theRebuttalSourcedandLLMJudgedvariants\. The rawContradictedlabel corresponds toNot Supportedin the four\-way scheme\.
RebuttalSourcedLabeling Prompt \(evidence: the discussion thread\)System:You are an expert at analyzing academic discourse\.User:Given a reviewer’s claim and the author’s response, determine how the authors address the claim\.Guidelines:*Supported*: authors clearly agree with or confirm the claim;*Partially Supported*: authors acknowledge some validity but not fully;*Contradicted*: authors explicitly disagree or provide counter\-evidence;*Not Determined*: authors do not address the claim, or it is unclear\.Reviewer Claim:\{claim\}Author’s Response:\{author\_response\}
LLMJudgedLabeling Prompt \(evidence: the manuscript\)System:You are an expert at analyzing academic papers\.User:Given a reviewer’s claim about a paper and the paper content, determine whether the claim is true according to the paper\.Guidelines:*Supported*: the claim is true according to the paper;*Partially Supported*: partially true;*Contradicted*: the paper contradicts the claim;*Not Determined*: the paper does not address the topic, or there is insufficient information\. For example, if the claim states “the paper lacks comparison with baseline X” and the paper indeed omits it, the label is*Supported*\.Reviewer Claim:\{claim\}Paper Content:\{paper\_content\}
Figure 5:The two labeling prompts\. They differ only in the evidence they see:RebuttalSourcedreads the author–reviewer thread and never the manuscript,LLMJudgedreads the manuscript and never the thread\. This disjointness is what makes their agreement informative \(Appendix[A\.8](https://arxiv.org/html/2609.25046#A1.SS8)\)\.
## Appendix DAdditional Results
### D\.1Full\-Context Verification Results
Table 14:Full\-context accuracy \(five models, six benchmarks\)\.Qwen2\.5\-7Bexcludes 9 over\-length papers \(Section[2](https://arxiv.org/html/2609.25046#S4.T2)\)\.Human\-Verifiedis the expanded 300\-instance subset; Tablecompares it against the original 150 instances\.Benchmarko4\-miniGPT\-5\-miniClaude\-HaikuQwen3\-8BQwen2\.5\-7BPaperSourced0\.8660\.7920\.7640\.6180\.753RebuttalSourced0\.2450\.2800\.2950\.2760\.278LLMJudged0\.4420\.3790\.2950\.1710\.170HighAgreement0\.5360\.5130\.4450\.3020\.213verifiable\-only0\.4950\.5180\.3880\.3200\.250Human\-Verified0\.2670\.3170\.3000\.3770\.256
Table[14](https://arxiv.org/html/2609.25046#A4.T14)reports the precise full\-context verification accuracy for all five models across the six benchmarks, complementing Figure[1](https://arxiv.org/html/2609.25046#S3.F1)\.
### D\.2Accuracy vs\. Macro\-F1
Table 15:Accuracy vs\. macro\-F1 \(full\-context FC and best\-retriever RAG\) onRebuttalSourcedandLLMJudged\. A large Acc−\-F1 gap reflects label\-distribution bias, not reasoning\.Qwen2\.5\-7BFC excludes 9 over\-length papers\.Full\-ContextRAG \(best retriever\)RebuttalSourcedLLMJudgedRebuttalSourcedLLMJudgedModelAccF1AccF1AccF1AccF1o4\-mini0\.2450\.2460\.4420\.4160\.244D\{\}\_\{\\text\{D\}\}0\.2360\.514D\{\}\_\{\\text\{D\}\}0\.454GPT\-5\-mini0\.2800\.2750\.3790\.3810\.294D\{\}\_\{\\text\{D\}\}0\.2790\.414D\{\}\_\{\\text\{D\}\}0\.399Claude\-Haiku0\.2950\.2940\.2950\.2930\.297DR\{\}\_\{\\text\{DR\}\}0\.2750\.374DR\{\}\_\{\\text\{DR\}\}0\.336Qwen3\-8B0\.2760\.2680\.1710\.1920\.287D\{\}\_\{\\text\{D\}\}0\.2710\.371D\{\}\_\{\\text\{D\}\}0\.336Qwen2\.5\-7B0\.2780\.2690\.1700\.1900\.324D\{\}\_\{\\text\{D\}\}0\.2770\.209D\{\}\_\{\\text\{D\}\}0\.200
Table[15](https://arxiv.org/html/2609.25046#A4.T15)reports accuracy and macro\-F1 for both full\-context and RAG paradigms\. In the RAG setting,Qwen2\.5\-7BDense achieves the highestRebuttalSourcedaccuracy \(0\.324, macro\-F1 0\.277\), followed byGPT\-5\-miniDense \(Acc 0\.294, F1 0\.279\) ando4\-miniDense \(Acc 0\.244, F1 0\.236\)\. The large accuracy–macro\-F1 gaps for some models reflect calibration bias rather than genuine reasoning, motivating macro\-F1 as a complementary metric robust to label imbalance\.
### D\.3PaperSourced RAG Results
Table[13](https://arxiv.org/html/2609.25046#A3.T13)reports thePaperSourcedOracle upper bound alongside Dense\+R for all five models; the full four\-retrieverPaperSourcedresults are in the main RAG table \(Table[2](https://arxiv.org/html/2609.25046#S4.T2)\)\. Oracle feeds the gold\-standard supporting paragraph directly to the verifier, providing an upper bound under perfect retrieval\.
The Oracle upper bound reveals thatPaperSourcedis nearly solvable when retrieval is perfect:o4\-minireaches 0\.988,GPT\-5\-mini0\.958,Qwen3\-8B0\.872, andClaude\-Haiku0\.860\. The gap between Oracle and Dense\+R directly quantifies accuracy lost to retrieval imperfection: 0\.084 foro4\-mini\(0\.988→\\to0\.904\) and 0\.106 forGPT\-5\-mini\. One anomaly:Qwen2\.5\-7BDense\+R \(0\.860\) exceeds its own Oracle \(0\.816\), because Dense\+R supplies ten passages \(including adjacent context\) while Oracle supplies only the single gold paragraph\. For a model with limited long\-context capacity, richer surrounding context improves its ability to confirm supported claims even when the exact evidence is present in both conditions\.
### D\.4Extended RAG Analysis
This section collects secondary RAG findings deferred from Section[5](https://arxiv.org/html/2609.25046#S5)\.
##### Performance gap narrows on interpretive claims\.
OnRebuttalSourced, all retrieval configurations cluster between 0\.234 and 0\.244 foro4\-mini, closely tracking the full\-context result of 0\.245\. ForLLMJudged, the picture inverts:Qwen2\.5\-7Baccuracy collapses to 0\.209–0\.228 with macro\-F1 close to accuracy \(e\.g\., 0\.200 for Dense\), confirming that itsRebuttalSourcedaccuracy advantage does not transfer to semantically harder claims\.
##### Qwen2\.5\-7Bcollapses onLLMJudgedyet leads onHuman\-Verified\.
OnLLMJudged, among the models scored by re\-scoringRebuttalSourcedpredictions,GPT\-5\-minileads \(best: BM25\-R 0\.417\), whileQwen2\.5\-7Bcollapses to 0\.209–0\.228, a 2×\\timesgap confirming a hard capability boundary on semantic entailment\. OnHuman\-Verified,Qwen2\.5\-7Brecovers sharply: BM25\-R reaches 0\.367, the highestHuman\-Verifiedscore across all models and retrievers, consistent with its tendency to over\-predict Supported and Partially Supported, the dominant categories in the human\-annotated set\.
##### Retriever choice matters most for grounded claims\.
AcrossPaperSourcedandverifiable\-only, BM25\+Reranker is usually the strongest retriever, beating pure BM25 and dense retrieval for every model onPaperSourcedand for most onverifiable\-only\(the exception isGPT\-5\-mini, where dense retrieval is marginally higher\)\. Reranking improves precision at low ranks, which is critical when the verifier’s context is restricted to a small passage set\. ForLLMJudged, dense retrieval surpasses the reranker foro4\-mini\(0\.514 vs\. 0\.499\), suggesting that semantic embeddings better capture paraphrased or implicit evidence that lexical matching misses\.
### D\.5Classical Non\-LLM Baselines
A fair question about any LLM\-based system is whether the LLM is doing work that a smaller, cheaper model could do\. We tested this directly with three zero\-shot entailment models trained on MultiNLI\([Williams et al\., 2018](https://arxiv.org/html/2609.25046#bib.bib46)\): RoBERTa\-large\-MNLI\([Liu et al\., 2019](https://arxiv.org/html/2609.25046#bib.bib27)\), BART\-large\-MNLI\([Lewis et al\., 2020a](https://arxiv.org/html/2609.25046#bib.bib24)\), and DeBERTa\-large\-MNLI\([He et al\., 2021](https://arxiv.org/html/2609.25046#bib.bib14)\)\. For each claim, the premise is the concatenation of the top\-3 retrieved chunks and the hypothesis is the claim, and the three\-way output is mapped onto our schema \(entailment toSupported, contradiction toNot Supported, neutral toNot Determined\)\. We ran every model under all four retrieval configurations, so that a weak result could not be attributed to one bad retriever\.
Table 16:Zero\-shot NLI baselines under all four retrieval configurations, reported as Accuracy / macro\-F1 on all sixPeerifybenchmarks\.Best LLMis the single highest\-accuracy model–retriever pair per benchmark, taken from Table[2](https://arxiv.org/html/2609.25046#S4.T2)\. No NLI configuration exceeds a macro\-F1 of0\.240\.24on any benchmark, and the ranking is stable across retrievers, so the gap is a property of the task formulation rather than of a particular retriever\.‡NLI baselines and the matchedBest LLMreference onHuman\-Verifiedare computed on the original 150\-instance subset\.MethodRetrieverPaperSourcedRebuttalSourcedLLMJudgedHighAgreementVerifiable\-onlyHumanVerified‡RoBERTa\-MNLIBM250\.070 / 0\.0440\.147 / 0\.1160\.401 / 0\.2120\.234 / 0\.1620\.206 / 0\.1520\.167 / 0\.106BM25\+Reranker0\.072 / 0\.0450\.136 / 0\.1080\.386 / 0\.2050\.206 / 0\.1430\.168 / 0\.1260\.147 / 0\.103Dense0\.066 / 0\.0410\.145 / 0\.1210\.384 / 0\.2060\.206 / 0\.1520\.168 / 0\.1320\.153 / 0\.120Dense\+Reranker0\.068 / 0\.0420\.136 / 0\.1080\.390 / 0\.2050\.194 / 0\.1370\.168 / 0\.1310\.147 / 0\.106BART\-MNLIBM250\.426 / 0\.1990\.176 / 0\.1290\.359 / 0\.2320\.303 / 0\.2240\.260 / 0\.2040\.153 / 0\.124BM25\+Reranker0\.490 / 0\.2190\.191 / 0\.1430\.356 / 0\.2300\.326/0\.2380\.275 / 0\.2130\.167 / 0\.134Dense0\.372 / 0\.1810\.175 / 0\.1340\.367 /0\.2380\.314 / 0\.2360\.275 / 0\.2180\.173 /0\.145Dense\+Reranker0\.512/0\.2260\.188 / 0\.1420\.351 / 0\.2280\.314 / 0\.2370\.267 / 0\.2150\.173 / 0\.141DeBERTa\-MNLIBM250\.184 / 0\.1040\.166 / 0\.1330\.406 / 0\.2260\.280 / 0\.2050\.229 / 0\.1780\.187/ 0\.140BM25\+Reranker0\.216 / 0\.1180\.149 / 0\.1210\.422/ 0\.2330\.257 / 0\.1860\.214 / 0\.1660\.147 / 0\.114Dense0\.176 / 0\.1000\.136 / 0\.1090\.374 / 0\.1980\.211 / 0\.1530\.206 / 0\.1620\.133 / 0\.105Dense\+Reranker0\.290 / 0\.1500\.135 / 0\.1050\.400 / 0\.2210\.240 / 0\.1760\.191 / 0\.1460\.153 / 0\.118Best LLMbest per benchmark0\.904 / 0\.2370\.324 / 0\.1570\.514 / 0\.4540\.514 / 0\.4990\.519 / 0\.4860\.367 / 0\.300
Table[16](https://arxiv.org/html/2609.25046#A4.T16)reports all 72 configurations\. Macro\-F1 never exceeds 0\.24 on any benchmark under any retriever, against 0\.45 to 0\.50 for the best LLM configuration on the document\-decidable splits\. Two structural causes explain the gap, and neither is a tuning problem\. Three\-way NLI has no output forPartially Supported, which is the single largest class inRebuttalSourcedat 310 of 800 claims, so the ceiling is imposed by the label space before inference begins\. And MNLI models cap the premise at 512 tokens, which truncates most of a 10\-chunk evidence set, so the model frequently decides on evidence it never read\.
The accuracy column occasionally looks less bad than the macro\-F1 column, and that discrepancy is itself informative\. BART\-MNLI reaches 0\.512 accuracy onPaperSourcedwith a macro\-F1 of 0\.226, and RoBERTa\-MNLI reaches 0\.401 accuracy onLLMJudgedat 0\.212 macro\-F1\. Both come from predicting one label heavily on a skewed split, the same distribution\-matching effect we discuss forQwen2\.5\-7Bin Appendix[D\.7](https://arxiv.org/html/2609.25046#A4.SS7)\. Read on macro\-F1, the ranking is stable: off\-the\-shelf entailment is not a viable substitute here, and the value of the LLM verifiers lies in long\-context, four\-way reasoning performed zero\-shot, not in raw entailment ability\.
### D\.6Error Analysis onPaperSourced
PaperSourcedis the only split where we know the correct label for every instance without appeal to a judge, since each claim is extracted from a manuscript passage and is genuinelySupportedby construction\. That makes it a clean setting for asking what the remaining errors actually are, rather than only how many there are\. We sampled 150 errors, stratified as 30 per verifier under the RAG setting, and read each one against the retrieved evidence and the source passage\. Table[17](https://arxiv.org/html/2609.25046#A4.T17)gives the breakdown; the categories are described below\.
Table 17:Error taxonomy onPaperSourcedunder RAG, from manual inspection of a stratified sample of 150 errors \(30 per verifier\)\. Every claim in this split is genuinelySupported, so any other predicted label counts as an error\.Error typeShareOverly strict specificity \(downgrade toPart\. Supp\.\)55\.3%Retrieval and evidence localization miss36\.0%Verifier output or parse failure8\.7%
- •Overly strict specificity \(55\.3%\)\.The verifier retrieves the right passage, recognizes the substance of the claim, and then downgrades it toPartially Supportedbecause one detail is not restated verbatim: an exact count, a pointer to a theorem or appendix, a numeric threshold\. The evidence supports the claim; the grader is stricter than the label definition intends\.
- •Retrieval and localization miss \(36\.0%\)\.The supporting passage is not among the top\-10 retrieved chunks, so the verifier correctly reports that the evidence in front of it does not contain the claim and hedges toNot Determined\. This is a retrieval failure presented as a verification error, and it is consistent with Recall@10 staying below 0\.55 for every retriever \(Appendix[D\.9](https://arxiv.org/html/2609.25046#A4.SS9)\)\.
- •Output or parse failure \(8\.7%\)\.The verifier emits a malformed response that cannot be mapped to a label, and the harness defaults toNot Determined\.
Taken together, 91\.3% of the error mass is strictness or retrieval rather than faulty inference on evidence the model actually saw\. The practical consequence is that rawPaperSourcedaccuracy overstates the reasoning error rate, and that two different fixes are indicated: better retrieval for the second category, and a calibration or rubric adjustment for the first\. Neither requires a stronger reasoner\.
### D\.7Predicted\-Label Distribution onPaperSourced
Qwen2\.5\-7BoutperformingGPT\-5\-minionPaperSourcedunder RAG looks anomalous next to the rest of the results\. It is not, and the mechanism is worth making explicit because it applies to any single\-class evaluation\.
Table 18:Predicted\-label distribution \(%\) onPaperSourcedunder RAG, aggregated over all four retrievers\. Every gold label in this split isSupported, so this is a single\-row confusion matrix and theSupportedcolumn equals accuracy\.ModelSupportedPart\. Supp\.Not Det\.Not Supp\.GPT\-5\-mini83\.19\.55\.12\.3Qwen2\.5\-7B85\.27\.05\.42\.4
Every gold label inPaperSourcedisSupported, so the predicted\-label distribution is a one\-row confusion matrix and theSupportedcolumn*is*the accuracy\. Table[18](https://arxiv.org/html/2609.25046#A4.T18)shows the two models side by side, aggregated over all four retrievers\.Qwen2\.5\-7BsaysSupported85\.2% of the time;GPT\-5\-minisays it 83\.1% of the time and hedges toPartially SupportedorNot Determinedmore often\. The 2\.1\-point difference is the whole margin\. On a split where the correct answer is always the confident one, willingness to commit is rewarded and caution is punished, even when caution is the better\-calibrated behavior\. Two claims from our sample illustrate the pattern: for a claim quoting an exact mean IoU of 44\.91%,GPT\-5\-minireturnsNot Determinedbecause it will not certify the precise figure, and for a claim describing several cache\-compression techniques evaluated under five settings it returnsPartially Supportedbecause it reads the enumeration as incomplete\. Both are goldSupported\.
The same prior is costly wherever the label set is not degenerate\. In full\-context evaluation,GPT\-5\-minileadsQwen2\.5\-7B0\.379 to 0\.170 onLLMJudged, 0\.513 to 0\.213 onHighAgreement, and 0\.518 to 0\.250 onverifiable\-only\(Table[14](https://arxiv.org/html/2609.25046#A4.T14)\)\. So thePaperSourcedordering reflects aSupported\-prior meeting an all\-Supportedsplit, not stronger verification, and it is the reason we report macro\-F1 alongside accuracy throughout\.
### D\.8Benchmark Size in Context
At 800 claims,Peerifyis small next to open\-domain fact\-verification corpora, and it is reasonable to ask whether conclusions drawn from it are stable\. The relevant comparison is not to web\-scale datasets but to claim\-verification benchmarks with expert, evidence\-grounded annotation, where size is bounded by annotation cost rather than by data availability\.
Table 19:Evaluation size and evidence scope ofPeerifyrelative to expert\-annotated scientific claim\-verification benchmarks\. Evidence scope matters as much as raw count: our labels are grounded in the full manuscript rather than an abstract or a single reference passage\.Dataset\#ClaimsAnnotationEvidence scopeMSVEC\([Evans et al\., 2023](https://arxiv.org/html/2609.25046#bib.bib12)\)200ExpertReference paperSciFact \(test\)\([Wadden et al\., 2020](https://arxiv.org/html/2609.25046#bib.bib44)\)300ExpertAbstractPeerify\(Human\-Verified\)300Human\-auditedFull manuscriptPeerify\(full\)800LLM \+ human\-auditedFull manuscript
Table[19](https://arxiv.org/html/2609.25046#A4.T19)places our benchmark against that reference class\. The full 800\-claim set exceeds the SciFact test split \(300\)\([Wadden et al\., 2020](https://arxiv.org/html/2609.25046#bib.bib44)\)and the MSVEC evaluation corpus \(200\)\([Evans et al\., 2023](https://arxiv.org/html/2609.25046#bib.bib12)\), and the hand\-audited 300\-claim subset matches them exactly\. Larger corpora in this space, such as COVID\-Fact\([Saakyan et al\., 2021](https://arxiv.org/html/2609.25046#bib.bib36)\)and HealthVer\([Sarrouti et al\., 2021](https://arxiv.org/html/2609.25046#bib.bib38)\), reach scale through crowdsourced or automatically derived labels over abstracts and web text; we deliberately traded that scale for manuscript\-grounded reliability, since a label for a review claim is only meaningful with respect to the full paper the reviewer read\. Annotating against a whole manuscript is substantially more expensive per claim than annotating against an abstract, which is the constraint the size reflects\.
### D\.9Retrieval Quality Analysis
##### Retrieval configuration\.
The dense retriever usesall\-MiniLM\-L6\-v2sentence embeddings, and the cross\-encoder reranker usescross\-encoder/ms\-marco\-MiniLM\-L\-6\-v2\. Manuscripts are segmented with Docling structure\-aware \(section/heading\-based\) chunking, producing variable\-length chunks with no fixed token window \(median 133 words per chunk, mean 132, p90 188\)\. For each claim we retrieve an initial pool ofk=20k\\\!=\\\!20candidate passages; when reranking is enabled, these are reranked down to the final top\-1010passages fed to the verifier, and non\-reranked configurations directly take the top\-1010\.
Table 20:Retrieval quality onPaperSourced\.RetrieverRecall@5Recall@10nDCG@10MRR@10BM250\.3980\.4740\.5700\.289BM25 \+ Reranker0\.4340\.4920\.6080\.331Dense0\.4200\.5280\.6440\.299Dense \+ Reranker0\.4520\.5300\.6840\.346
Table[20](https://arxiv.org/html/2609.25046#A4.T20)compares four retrieval strategies onPaperSourced, where the ground\-truth supporting section is known, enabling intrinsic evaluation\. Dense\+Reranker achieves the best performance across all metrics \(Recall@5: 0\.452, Recall@10: 0\.530, nDCG@10: 0\.684, MRR: 0\.346\), while BM25 alone is weakest\.
Consistent with prior work\([Kamalloo et al\., 2024](https://arxiv.org/html/2609.25046#bib.bib19)\), reranking substantially boosts BM25\. BM25\+Reranker surpasses standalone dense retrieval at Recall@5 \(0\.434 vs\. 0\.420\) but lags at Recall@10 \(0\.492 vs\. 0\.528\), suggesting that lexical matching excels at placing relevant chunks at shallow ranks while dense representations improve deeper recall\. Despite this, Recall@10 remains below 0\.55 across all configurations: the ground\-truth section is absent from the top\-10 retrieved passages in nearly half of all cases\. The Oracle upper bound \(Table[13](https://arxiv.org/html/2609.25046#A3.T13)\) makes this bottleneck concrete: the 8\.4\-point gap between Oracle \(0\.988\) and Dense\+R \(0\.904\) foro4\-miniis attributable entirely to retrieval imperfection\.
### D\.10Token Efficiency and Operating Cost
RAG\-based verification substantially reduces token consumption relative to full\-context processing\. Tables[21](https://arxiv.org/html/2609.25046#A4.T21),[22](https://arxiv.org/html/2609.25046#A4.T22), and[23](https://arxiv.org/html/2609.25046#A4.T23)report average token usage per claim across all six benchmarks in both settings\.
Table 21:Tokens per claim, RAG verification \(four retrievers, six benchmarks\)\.BenchmarkBM25FAISSDense\+RBM25\+RPaperSourced2,235\.852,231\.822,384\.762,291\.40RebuttalSourced2,156\.782,128\.872,371\.782,236\.87LLMJudged2,156\.782,128\.872,371\.782,236\.87HighAgreement2,112\.562,085\.182,262\.942,205\.61verifiable\-only2,153\.542,240\.622,253\.422,244\.24Human\-Verified1,922\.491,986\.741,976\.661,945\.26Average2,123\.002,133\.682,270\.222,193\.38
Table 22:Tokens per claim, full\-context verification \(six benchmarks\)\.BenchmarkTokens per ClaimPaperSourced21,648\.85RebuttalSourced21,130\.27LLMJudged21,130\.27HighAgreement22,323\.60verifiable\-only21,333\.30Human\-Verified18,912\.43Average21,079\.79
Table 23:Token efficiency, RAG vs\. full\-context\. Reduction = full\-context/RAG token ratio\.BenchmarkFull\-ContextRAG \(Avg\)ReductionPaperSourced21,648\.852,285\.969\.47×\\timesRebuttalSourced21,130\.272,223\.579\.50×\\timesLLMJudged21,130\.272,223\.579\.50×\\timesHighAgreement22,323\.602,166\.5710\.30×\\timesverifiable\-only21,333\.302,222\.969\.60×\\timesHuman\-Verified18,912\.431,957\.799\.66×\\timesOverall Average21,079\.792,180\.079\.67×\\times
On average, RAG reduces token consumption by roughly9\.7×9\.7\\timesrelative to full\-context verification, consistently across all six benchmarks, from9\.47×9\.47\\timesonPaperSourcedto10\.30×10\.30\\timesonHighAgreement\. This is what makes the pipeline runnable at venue scale rather than only on a sample, and it is also what lets open\-weight models with restricted context windows participate at all\. The saving does not cost accuracy on grounded benchmarks, and onPaperSourcedit improves it: Dense\+R reaches 0\.904 foro4\-miniagainst 0\.866 in full\-context, because a focused evidence set is easier to reason over than a whole manuscript\.
Together with the extraction costs in Table[4](https://arxiv.org/html/2609.25046#A1.T4), this fixes the operating cost of a deployment\. Extraction runs locally at no inference cost and a median of 8\.2 seconds per review, and verification consumes about 2\.2K tokens per claim rather than 21K\. These are the measurements behind the deployment discussion in Appendix[E](https://arxiv.org/html/2609.25046#A5)\.
### D\.11Where Intrinsic Retrieval Metrics Apply
Recall@kk, nDCG@kk, and MRR are reported only onPaperSourced, and the Oracle upper bound is computed there as well\. This is a constraint of the data rather than a choice about what to report\. All three metrics need a gold evidence span, and spans exist only where a claim was drawn from a known passage\. ForRebuttalSourced,LLMJudged, andHuman\-Verified, the label comes from an author–reviewer exchange or from a whole\-manuscript reading, and neither procedure identifies the paragraph that settles the claim\. Some claims are settled by an absence, such as “no ablation is reported,” which has no supporting span anywhere in the document\. Annotating spans for those splits would mean inventing a ground truth we do not have\.
The question behind the metric, whether retrieval or reasoning limits performance, can still be answered on those splits, by removing retrieval instead of measuring it\. Full\-context verification passes the entire manuscript to the verifier and skips the retrieval stage, so it is a retrieval\-free upper bound\. If retrieval were binding, accuracy should rise sharply\. It does not: onHuman\-Verified,GPT\-5\-minireaches 0\.317 in full\-context against 0\.297 with its best retriever, ando4\-minireaches 0\.267 against 0\.257 \(Tables[14](https://arxiv.org/html/2609.25046#A4.T14)and[9](https://arxiv.org/html/2609.25046#A1.T9)\)\. Perfect retrieval buys a point or two, so what limits these splits is verification reasoning\.PaperSourcedbehaves in exactly the opposite way, where the 8\.4\-point Oracle\-to\-Dense\+R gap foro4\-miniis retrieval and nothing else, which is why we keep both regimes in the evaluation\.
## Appendix EDeployment Setting and Intended Use
##### The pipeline runs at the scale of a real venue\.
All three stages are implemented end to end and are cheap enough to apply to an entire submission cycle\. Claim extraction runs locally on a single consumer GPU at a median of 8\.2 seconds per review with no API cost \(Appendix[A\.2](https://arxiv.org/html/2609.25046#A1.SS2)\), and retrieval\-augmented verification consumes roughly9\.7×9\.7\\timesfewer tokens per claim than passing the verifier a whole manuscript \(Appendix[D\.10](https://arxiv.org/html/2609.25046#A4.SS10)\)\. Because both figures are measured rather than estimated, a venue can compute the cost of a deployment in advance\. This is what we mean when we describePeerifyas production\-ready: the engineering is complete, the throughput is known, and nothing in the design assumes a research\-scale workload\.
##### Every prediction ships with its evidence\.
The output of the pipeline is not a bare label\. Each claim is returned together with the manuscript passages the verifier conditioned on, so a decision can always be traced back to the text that produced it\. This is what makes the system useful in an editorial workflow: a reader can confirm or overturn any prediction in seconds by looking at the passages already retrieved for them, without opening the paper and searching from scratch\.
##### The intended workflow is human\-in\-the\-loop by design\.
Peerifyis built to be an assistant to editorial judgement rather than a replacement for it\. An area chair opens a review and sees which reviewer claims the system flagged as needing a closer look, each attached to the relevant manuscript passages, and directs attention accordingly\. Authors can use the same view when preparing a rebuttal, and reviewers can use it to self\-check a draft review before submitting\. The value the system delivers is triage and evidence retrieval at scale, which is precisely the part of the task that does not scale for humans, while the final judgement stays where it belongs\.
##### The benchmark is deliberately demanding\.
Peerifyships with an evaluation suite built to remain informative as models improve\. The variants span the full difficulty range, fromPaperSourced, where evidence is explicit and the best verifier reaches 0\.904 under RAG, through toRebuttalSourced, where claims are drawn from live author–reviewer disputes and mix factual observation with interpretation\. Setting the harder end of that range beyond what current frontier models saturate is a design goal rather than a shortcoming: it gives the community headroom to measure progress on manuscript\-grounded verification for several model generations, and it keeps the assistive framing above honest about which claims a human should look at first\.Similar Articles
Do Methods Support the Claims? Intra-Paper Verification for Peer Review
This paper introduces intra-paper claim verification, a framework that uses LLMs to evaluate whether novelty claims in a paper are supported by its methodological evidence, addressing a gap in existing automated peer review systems. Human evaluation shows significant alignment with human reviewer concerns, especially for novelty-related issues.
Beyond Final Decisions: A Process-Centric Benchmark for Transparent AI-Assisted Peer Review
This paper introduces a process-centric benchmark for evaluating AI-assisted peer review systems, aiming to improve transparency and reliability beyond final decision accuracy.
Veriphi: Attack-Guided Neural Network Verification with Dataset-Dependent Training Methods
Veriphi is a GPU-accelerated neural network verification system that combines adversarial attacks with formal certification. It demonstrates that the effectiveness of training methods (standard, adversarial, certified) depends heavily on dataset complexity, with IBP dominating on simple MNIST and PGD on complex CIFAR-10, and achieves 5x verification speedup.
Benchmarking Agentic Review Systems
This paper benchmarks agentic review systems for peer review, evaluating open-source and proprietary systems on research papers. The best configuration achieves 83.0% pairwise accuracy and catches 71.6% of injected errors, but user feedback highlights issues with false positives and nitpicks.
HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews
This paper introduces HalluPeer, a taxonomy-driven benchmark for detecting hallucinations in scientific peer reviews, providing annotated data to evaluate and improve detection methods.