A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review
Summary
The paper introduces Rhetorical Robustness as an evaluation target for AI scientific reviewers, proposes the RobustReview benchmark with 1,260 manuscript versions, and presents SciCore, a dual-branch reviewer that combines full-manuscript assessment with a science-core-extracted judgment to improve stability while maintaining human alignment.
View Cached Full Text
Cached at: 10/01/26, 09:46 AM
# A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review
Source: [https://arxiv.org/html/2609.39027](https://arxiv.org/html/2609.39027)
Ming LiCo\-first AuthorChengrui FanJianpeng ChenHan ChenTianyi ZhouDawei Zhou
###### Abstract
AI reviewers can assign different judgments to manuscripts that report the same science in different wording, potentially rewarding rhetorical optimization over scientific improvement\. We formulateRhetorical Robustnessas the joint requirement of stability across content\-preserving rewrites and discrimination across papers\. We introduceRobustReview, a controlled full\-manuscript benchmark with 1,260 manuscript versions, and evaluate 30 reviewer configurations\. The benchmark reveals*false robustness*, where low rewrite sensitivity coincides with score collapse across papers, and shows that human alignment and rhetorical robustness rank reviewers differently\. Moreover, the evaluated content\-focused prompting protocol does not consistently improve robustness across backbones\. Motivated by these findings, we introduceSciCore, a dual\-branch reviewer that averages a full\-manuscript judgment with a judgment based on an extracted, structured science core\. This design combines manuscript\-level assessment with a content\-normalized view intended to reduce rhetorical sensitivity\. In our primary GPT\-5\.5 comparison,SciCoreachieves a leading joint stability\-discrimination profile among the benchmarked reviewers while maintaining competitive human alignment\. These results identify rhetorical robustness as a distinct evaluation target and demonstrate the potential of science\-core review to improve it\.
††Project Page:[https://github\.com/c\-steve\-wang/Robust\_Review](https://github.com/c-steve-wang/Robust_Review)## 1Introduction
Large language models \(LLMs\) are increasingly being used to support scientific peer review, from generating manuscript feedback to assisting with review and decision making\([Wang et al\., 2020](https://arxiv.org/html/2609.39027#bib.bib46);[Liang et al\., 2023](https://arxiv.org/html/2609.39027#bib.bib30);[Thakkar et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib41);[Chen et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib7)\)\. Because review judgments shape which work is accepted, revised, and disseminated, the reliability of AI reviewers matters not only for individual manuscripts but also for the broader scientific record\([Fytas et al\., 2021](https://arxiv.org/html/2609.39027#bib.bib15);[Li et al\., 2025b](https://arxiv.org/html/2609.39027#bib.bib28)\)\. Existing evaluations commonly ask whether AI\-generated reviews are useful, resemble expert feedback, or reproduce human scores and decisions\([Liang et al\., 2023](https://arxiv.org/html/2609.39027#bib.bib30);[Zhou et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib53);[Li et al\., 2025a](https://arxiv.org/html/2609.39027#bib.bib27);[Chen et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib7)\)\. These criteria are necessary, but they do not fully characterize a trustworthy AI reviewer\. Such a reviewer should also preserve its scientific judgments when the same reported science is expressed in rhetorically different ways\. We call this propertyRhetorical Robustness, and argue that it is an important requirement for trustworthy AI reviewers\. As LLMs make it increasingly easy to rewrite and strategically optimize manuscripts at low cost,review judgments that can be manipulated through wording alone would reward rhetorical optimization over scientific improvement and undermine the credibility of AI\-based review\([Kaneko, 2026](https://arxiv.org/html/2609.39027#bib.bib20);[Li et al\., 2026b](https://arxiv.org/html/2609.39027#bib.bib25);[Li et al\., 2026c](https://arxiv.org/html/2609.39027#bib.bib26)\)\. Presentation may legitimately shape assessments of clarity and communicative quality, but it should not unduly alter judgments of scientific merit when the scientific content is preserved\([James et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib19)\)\.
Existing work shows that AI review and LLM judging systems can be manipulated by overt instructions, adversarial phrasing, and strategically constructed text\([Ye et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib50);[Lin et al\., 2025](https://arxiv.org/html/2609.39027#bib.bib31);[Collu et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib8)\)\. More importantly for scientific review, visible and meaning\-preserving revisions to titles, abstracts, and full manuscripts can also alter automated evaluations to a large extent\([Du, 2025](https://arxiv.org/html/2609.39027#bib.bib9);[Kaneko, 2026](https://arxiv.org/html/2609.39027#bib.bib20);[Li et al\., 2026b](https://arxiv.org/html/2609.39027#bib.bib25);[Baumann et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib3);[Yang et al\., 2026b](https://arxiv.org/html/2609.39027#bib.bib49);[Li et al\., 2026c](https://arxiv.org/html/2609.39027#bib.bib26)\)\. However, the evaluation and design of AI reviewers have generally not treated rhetorical robustness as a joint requirement of stable judgment and scientific discrimination\. Existing evaluations have extensively examined human alignment\([Liang et al\., 2023](https://arxiv.org/html/2609.39027#bib.bib30);[Chen et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib7)\), while rhetorical robustness has remained a comparatively neglected dimension\.
We therefore formulate rhetorical robustness through two complementary requirements:within\-paper stabilityacross rhetorical variants designed to preserve reported scientific content, andbetween\-paper discrimination\. Stability alone is insufficient: a reviewer assigning nearly identical scores to every paper would appear robust while failing to distinguish papers\. The joint requirement asks whether reviewers can resist rhetorical variation without collapsing differences across papers\. Human alignment captures a separate property, since agreement with human judgments on original manuscripts does not establish stability across their rhetorical variants\. This motivates our central question:Can AI reviewers remain stable under rhetorical variation while retaining paper\-level discrimination?
We operationalize this requirement throughRobustReview, a controlled full\-manuscript benchmark built from 60 anonymized ICLR 2026 submissions, sampled equally from six mean human\-review score intervals to cover different levels of human assessment\. We retain each original manuscript and construct matched variants under 10 rhetorical conditions using two independent LLM systems, yielding 1,260 manuscripts in total\. We evaluate 30 reviewer configurations spanning rubric\-instructed LLMs, specialized scientific\-review models, and agentic review systems under multiple review protocols\. Our evaluation combines rhetorical stability and paper discrimination with agreement with human review scores\.
The evaluated content\-focused prompting protocol does not consistently improve robustness across backbones\. We therefore introduceSciCore, a dual\-branch framework\. Because rhetorical robustness requires greater stability without sacrificing meaningful paper\-level discrimination,SciCoreuses a content\-normalized scientific judgment to complement manuscript\-level assessment\. One branch reviews the complete manuscript\. The other extracts a structured record of the manuscript’s reported problem, claims, methods, assumptions, evidence, results, and limitations, and then reviews that record directly with an adapted protocol\. We call this record a*science core*\. The final overall assessment is the mean of the two branch scores\. The science\-core branch is motivated by invariant representation theory\([Eaton, 1989](https://arxiv.org/html/2609.39027#bib.bib13);[Dubois et al\., 2021](https://arxiv.org/html/2609.39027#bib.bib10)\): the extracted record should vary little across rhetorical realizations of the same reported science while preserving distinctions across papers\. In our experiments,SciCoreachieves a leading joint stability\-discrimination profile while maintaining competitive human alignment\. These results suggest that augmenting conventional review with a content\-normalized scientific judgment can improve rhetorical robustness without discarding manuscript\-level assessment\.
Our contributions are threefold:
- •We argue forRhetorical Robustnessas an important requirement for trustworthy AI reviewers and formulate it jointly through within\-paper stability and between\-paper discrimination, with human alignment evaluated as a distinct property\.
- •We introduceRobustReview, a controlled full\-manuscript benchmark spanning rewrite strategies, reviewer families, and review protocols\. It shows that content\-focused review is nontrivial, exposes widespread configuration\-dependent sensitivity, and identifies false robustness as a central measurement failure\.
- •We introduceSciCore, a dual\-branch framework that averages a conventional full\-manuscript judgment with a content\-normalized judgment from an extracted science core\. The method improves the joint robustness profile while maintaining competitive human alignment in our primary evaluation\.
Figure 1:Overview of the benchmark and method\.RobustReviewevaluates rhetorical robustness across controlled rhetorical variants, whileSciCoreaverages a full\-manuscript judgment with a content\-normalized judgment obtained by extracting and reviewing the science core\.
## 2RobustReview: Benchmark Design
### 2\.1Rhetorical Robustness
We distinguish judgments of reported science from assessments of presentation\. Clarity and style may legitimately affect the latter, but judgments of scientific contribution should not be unduly altered by rhetorical rewrites designed to preserve the reported scientific content\. We therefore define*rhetorical robustness*as the joint ability of an AI reviewer to maintain stable scientific judgments under such variation while retaining sensitivity to differences in reported scientific content across manuscripts\.
Formally, for paperii, letxi0x\_\{i0\}denote the original manuscript\. Each controlled rewrite settingkk, which specifies both a rhetorical condition and a rewrite producer, generates a variantxikx\_\{ik\}designed to preserve the paper’s reported claims, methods, evidence, results, and conclusions while changing how this content is communicated\. Such variation may involve claim stance, evidence framing, contribution organization, technical register, or lexical and syntactic realization\. A reviewer configurationmmspecifies both the reviewer model and review protocol\. The resulting rewrite and review process is
xik\\displaystyle x\_\{ik\}=Rewritek\(xi0\),\\displaystyle=\\operatorname\{Rewrite\}\_\{k\}\(x\_\{i0\}\),k=1,…,K,\\displaystyle k=1,\\ldots,K,\(1\)ymik\\displaystyle y\_\{mik\}=Reviewm\(xik\),\\displaystyle=\\operatorname\{Review\}\_\{m\}\(x\_\{ik\}\),k=0,…,K\.\\displaystyle k=0,\\ldots,K\.Here,i=1,…,Ni=1,\\ldots,Nindexes papers,xikx\_\{ik\}is variantkkof paperii, andymiky\_\{mik\}is the judgment assigned by reviewer configurationmmto that presentation\. A rhetorically robust reviewer should jointly satisfy two complementary requirements\. First,within\-paper stabilityrequires judgments to remain consistent across rhetorical variants of the same paper\. This condition alone is insufficient because indiscriminately constant scores would be maximally stable\. Second,between\-paper discriminationrequires the reviewer to preserve distinctions based on the reported scientific content of different manuscripts, thereby ruling out this collapse\. This requirement does not treat cross\-paper score variation as ground\-truth scientific merit; it asks whether paper differentiation remains large enough relative to rewrite\-induced variation to be meaningful\.
Rhetorical robustness therefore constitutes ajoint stability\-discrimination requirement: it requires within\-paper stability without sacrificing between\-paper discrimination\. Section[2\.4](https://arxiv.org/html/2609.39027#S2.SS4)operationalizes this joint requirement with two direct within\-paper metrics and three joint stability\-discrimination metrics\. The joint metrics do not measure between\-paper behavior in isolation: each relates within\-paper consistency to cross\-paper variation or separation\. Alignment with human reviewer judgments remains a distinct evaluation dimension, which we measure separately\.
### 2\.2Benchmark Construction
RobustReviewoperationalizes rhetorical robustness through*matched paper families*, each containing an original manuscript and rhetorical variants designed to preserve its reported scientific content\. Comparisons within each family directly measure rewrite\-induced instability\. Comparisons across families then provide the reference needed to determine whether that within\-paper consistency coexists with differentiation among papers\. Human judgments on the original manuscripts provide a separate reference for alignment\.
We constructRobustReviewfrom 60 anonymized ICLR 2026 submissions with matched arXivLaTeXsources\. We stratify eligible papers by their mean human overall\-assessment rating and randomly sample 10 papers from each of six score intervals\. This balanced sampling covers papers with different human\-assessed ratings, while the human scores provide an external reference for evaluating the reviewers’ assessments\. We apply 10 rhetorical conditions\. Six single\-dimension conditions alter novelty stance, scope framing, evidence framing, contribution salience, technical register, or linguistic complexity\. Four complex conditions apply a joint rewrite across dimensions, two or three recursive rewrite rounds \(R2 and R3\), or a reviewer\-guided rewrite based on model feedback\. Each condition is independently instantiated by GPT\-5\.5\([OpenAI, 2026](https://arxiv.org/html/2609.39027#bib.bib37)\)and Claude Opus 4\.8\([Anthropic, 2026a](https://arxiv.org/html/2609.39027#bib.bib1)\), producing two variants per paper and condition\. The resulting corpus contains 60 original manuscripts and 1,200 rhetorical variants, or 1,260 full manuscripts in total\. All rewrites operate on completeLaTeXprojects under content\-preservation and structural controls\. Automated and human audits indicate that core technical content is largely preserved across the five assessed dimensions \(Appendix[E\.1](https://arxiv.org/html/2609.39027#A5.SS1)\)\. Appendices[C\.1](https://arxiv.org/html/2609.39027#A3.SS1)and[C\.2](https://arxiv.org/html/2609.39027#A3.SS2)give the sampling procedure, complete intervention definitions, rewrite procedure, and construction checks\.
### 2\.3Reviewer Configurations
We evaluate 30 reviewer configurations across three system families\. The general\-purpose family comprises GPT\-5\.5\([OpenAI, 2026](https://arxiv.org/html/2609.39027#bib.bib37)\), GPT\-5\-mini\([OpenAI, 2025](https://arxiv.org/html/2609.39027#bib.bib36)\), Claude Sonnet 5\([Anthropic, 2026b](https://arxiv.org/html/2609.39027#bib.bib2)\), GLM\-5\.2\([GLM\-5\-Team et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib16)\), Kimi\-K2\.6\([Moonshot AI, 2026](https://arxiv.org/html/2609.39027#bib.bib34)\), GPT\-OSS\-120B\([OpenAI, 2025](https://arxiv.org/html/2609.39027#bib.bib35)\), Gemini\-3\.5\-Flash\-Lite\([Google, 2026](https://arxiv.org/html/2609.39027#bib.bib17)\), and Qwen\-3\.5\-Flash\([Qwen Team, 2026](https://arxiv.org/html/2609.39027#bib.bib39)\)\. Each is prompted under three protocols:Standard,Strict, andPersistent\.Standardfollows the ICLR review criteria and scoring scheme\.Strictraises the evidentiary threshold through a more conservative rubric\.Persistentretains the standard criteria while repeatedly instructing the reviewer to base scientific judgments on substantive content rather than rhetorical presentation\. It therefore directly tests whether content\-only instructions are sufficient to separate scientific judgment from rhetorical presentation\. We additionally evaluate the specialized models OpenReviewer, CycleReviewer, and DeepReviewer\([Idahl and Ahmadi, 2025](https://arxiv.org/html/2609.39027#bib.bib18);[Weng et al\., 2025](https://arxiv.org/html/2609.39027#bib.bib47);[Zhu et al\., 2025b](https://arxiv.org/html/2609.39027#bib.bib55)\), and the agentic systems AI Scientist, OpenJudge, and ProReviewer\([Lu et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib33);[The OpenJudge Team, 2025](https://arxiv.org/html/2609.39027#bib.bib42);[Fang et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib14)\), using their native review procedures\. For every configuration, an original manuscript and all of its rhetorical variants are evaluated with the same procedure and scoring criteria\. Appendix[C\.3](https://arxiv.org/html/2609.39027#A3.SS3)reports the system configurations and execution settings\.
### 2\.4Evaluation Metrics
Following the definition above, we evaluate every reviewer configuration with seven metrics\. MAD and Drift SD directly measure within\-paper stability \(lower is better\)\. Because low drift alone can result from score collapse, ICC, SPR, and discriminability jointly evaluate within\-paper consistency relative to cross\-paper variation or separation \(higher is better\); we refer to these as*joint stability\-discrimination metrics*\. Human MAE measures absolute agreement with mean human overall\-assessment scores, and Spearman correlation measures agreement with the human ranking of the original papers \(lower and higher are better, respectively\)\. On the score\-stratified benchmark, these human\-alignment metrics complement the robustness metrics by assessing whether the reviewers’ scores and rankings agree with human evaluations\. Together, the seven metrics characterize stability under rhetorical rewriting, discrimination among papers, and alignment with human judgments\. Appendix[C\.4](https://arxiv.org/html/2609.39027#A3.SS4)gives the complete definitions and equations\.
## 3Benchmark Findings: Limits of Current AI Reviewers
Table[1](https://arxiv.org/html/2609.39027#S3.T1)reports all seven metrics for 30 existing reviewer configurations andSciCoreunder the same matched\-manuscript design onRobustReview\. The within\-paper metrics quantify absolute score movement, the joint metrics test whether stability coexists with paper\-level discrimination, and the human\-alignment metrics provide a separate external comparison\. The analysis here focuses on the existing reviewers, withSciCoreincluded as a common\-scale reference\. Because the metrics capture different behaviors, no single column is sufficient for identifying a robust reviewer\.
Table 1:Main comparison of rhetorical robustness and human alignment\.MAD and Drift SD directly measure rewrite\-induced within\-paper instability\. ICC, SPR, and discriminability are joint stability\-discrimination metrics: each evaluates within\-paper consistency relative to cross\-paper variation or separation and should not be interpreted as a between\-paper\-only measure\. Arrows indicate the preferred direction\. The best point estimate in each column is shown in bold, the second\-best is underlined, and the third\-best is italicized\.SystemProtocolHuman alignmentWithin\-paper stabilityJoint stability\-discriminationH\-MAE↓\\downarrowSpearman↑\\uparrowMAD↓\\downarrowDrift SD↓\\downarrowICC↑\\uparrowSPR↑\\uparrowDiscrim\.↑\\uparrowGeneral\-purpose LLM reviewersGPT\-5\.5\([OpenAI, 2026](https://arxiv.org/html/2609.39027#bib.bib37)\)Standard1\.2940\.4480\.5980\.9560\.6150\.5550\.649Strict1\.0780\.5290\.7661\.2020\.6320\.5070\.672Persistent1\.2610\.4230\.4760\.8750\.6970\.5860\.682GPT\-5\-mini\([OpenAI, 2025](https://arxiv.org/html/2609.39027#bib.bib36)\)Standard1\.7110\.4170\.5430\.9450\.4010\.3900\.586Strict1\.2400\.4030\.8951\.2210\.3930\.4240\.598Persistent1\.6280\.3350\.5740\.9820\.4760\.4750\.589Claude Sonnet 5\([Anthropic, 2026b](https://arxiv.org/html/2609.39027#bib.bib2)\)Standard1\.1890\.4140\.4680\.8630\.5790\.5550\.646Strict1\.1510\.4750\.6121\.0500\.5460\.5500\.637Persistent1\.2390\.3600\.3070\.7100\.4890\.4910\.615GLM\-5\.2\([GLM\-5\-Team et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib16)\)Standard2\.044−0\.034\-0\.0341\.5002\.3410\.2260\.3220\.558Strict2\.047−0\.038\-0\.0381\.5482\.0940\.2120\.3890\.556Persistent2\.403−0\.190\-0\.1901\.9322\.6820\.2320\.3810\.558Kimi\-K2\.6\([Moonshot AI, 2026](https://arxiv.org/html/2609.39027#bib.bib34)\)Standard1\.6890\.2801\.3531\.9950\.2110\.4280\.555Strict1\.7370\.0560\.7531\.2610\.2870\.3770\.544Persistent1\.5780\.0501\.1391\.6450\.2830\.3920\.568GPT\-OSS\-120B\([OpenAI, 2025](https://arxiv.org/html/2609.39027#bib.bib35)\)Standard1\.5440\.0720\.7171\.1290\.1380\.3840\.532Strict1\.6920\.1190\.3850\.8880\.0910\.3530\.508Persistent1\.528−0\.008\-0\.0080\.8581\.3170\.1610\.3510\.536Gemini 3\.5 Flash\-Lite\([Google, 2026](https://arxiv.org/html/2609.39027#bib.bib17)\)Standard3\.1780\.4080\.2050\.6310\.1990\.4680\.511Strict1\.8780\.4650\.8401\.1930\.4240\.5200\.599Persistent2\.8280\.4870\.4430\.9680\.3470\.5020\.539Qwen 3\.5 Flash\([Qwen Team, 2026](https://arxiv.org/html/2609.39027#bib.bib39)\)Standard2\.3500\.3790\.8091\.3390\.2910\.4720\.561Strict1\.2390\.4090\.8661\.3400\.2260\.4040\.541Persistent2\.0370\.4610\.8631\.2810\.5970\.5190\.597Specialized review modelsOpenReviewer\([Idahl and Ahmadi, 2025](https://arxiv.org/html/2609.39027#bib.bib18)\)–1\.4060\.3391\.0801\.5240\.2150\.3740\.542CycleReviewer\([Weng et al\., 2025](https://arxiv.org/html/2609.39027#bib.bib47)\)–1\.4030\.0740\.7791\.0470\.1280\.3500\.533DeepReviewer\([Zhu et al\., 2025b](https://arxiv.org/html/2609.39027#bib.bib55)\)–1\.4710\.4060\.4540\.6890\.1610\.3810\.537Agentic review systemsAI Scientist\([Lu et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib33)\)–1\.5130\.0940\.8381\.2530\.2130\.3820\.544OpenJudge\([The OpenJudge Team, 2025](https://arxiv.org/html/2609.39027#bib.bib42)\)–1\.3060\.1200\.3350\.5970\.2390\.4440\.526ProReviewer\([Fang et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib14)\)–1\.4780\.2300\.7090\.9780\.0940\.3420\.516Our methodSciCore–1\.0720\.4880\.4760\.6870\.7750\.6520\.726
### 3\.1Content\-Only Review Is Not Reliably Achieved by Prompting Alone
We evaluatePersistent, a full\-manuscript review protocol that repeatedly emphasizes scientific content and provides explicit evidence\-based scoring guidance\. Relative toStandard, its effects vary across backbones: only GPT\-5\.5 improves on all five robustness metrics\. For Claude Sonnet 5,Persistentreduces MAD and Drift SD but lowers ICC, SPR, and discriminability, while the remaining backbones also show mixed changes\. These results show that the evaluated content\-focused prompting protocol does not consistently improve rhetorical robustness across backbones, motivating our investigation of an explicit content\-normalized branch\.
### 3\.2False Robustness: Within\-Paper Stability without Discrimination
Gemini\-3\.5\-Flash\-Lite underStandardachieves the lowest MAD among the existing configurations, at 0\.205, but its ICC is only 0\.199 and its discriminability is 0\.511, close to chance\. The same pattern appears among specialized and agentic reviewers: DeepReviewer and OpenJudge show relatively low drift, yet attain ICC values of only 0\.161 and 0\.239 and discriminability values of 0\.537 and 0\.526\. By contrast, GPT\-5\.5 underPersistenthas a higher MAD of 0\.476 but the strongest ICC, SPR, and discriminability among the 30 configurations\. We call low rewrite\-induced drift accompanied by weak paper differentiation*false robustness*: invariance alone can create the appearance of robustness without the discrimination that robustness is meant to preserve\.
### 3\.3Human Alignment and Rhetorical Robustness Are Distinct
Within GPT\-5\.5,Strictprovides the strongest human\-alignment profile among the existing configurations, with a Human MAE of 1\.078 and a Spearman correlation of 0\.529, whereasPersistenthas weaker human alignment but higher ICC, SPR, and discriminability\. Human alignment asks whether a reviewer reproduces human judgments on observed manuscripts; rhetorical robustness asks whether judgments remain stable across matched rhetorical variants while preserving differences among papers\. The two evaluations therefore favor different protocols, so human agreement cannot substitute for matched robustness evaluation\.
## 4SciCore: Dual\-Branch Review for Rhetorical Robustness
The central idea ofSciCoreis to augment manuscript\-based review with a scientific judgment that is less coupled to rhetorical presentation\. UnlikePersistent, which asks a reviewer to disregard rhetoric while still reading the complete manuscript, the science\-core branch first transforms the input into a content\-normalized scientific record\.SciCoreuses two complementary branches\. The*manuscript branch*reviews the complete paper under theStrictprotocol, retaining the full manuscript context\. The*science\-core branch*extracts a structured*science core*containing the reported scientific record and reviews that record with an adapted protocol\. The final overall assessment is the arithmetic mean of the two branch scores\. This design preserves a conventional judgment of the paper while giving equal weight to a content\-normalized judgment intended to vary less with rhetorical framing, organization, and linguistic expression\.
### 4\.1Design Principle: A Stable Scientific Branch without Discarding the Manuscript
The science\-core branch is designed to provide a scientific judgment that is less sensitive to rhetorical realization while retaining distinctions among papers\. Conceptually, it maps different presentations of the same reported science to similar structured records without collapsing scientifically distinct manuscripts\. This invariant\-representation view motivates extracting and reviewing a science core, but the implemented extractor is only an approximation: its stability and paper separation must be established empirically rather than assumed\.
The science\-core branch complements rather than replaces manuscript review\. When the science\-core branch varies less across matched rhetorical realizations than the manuscript branch, averaging their scores can reduce rhetorical sensitivity while retaining the full paper context\. This intuition addresses direct within\-paper stability only; the final reviewer must still be evaluated using the joint stability\-discrimination metrics to ensure that lower score drift does not come from cross\-paper collapse\.
We review the extracted record directly instead of first reconstructing another manuscript from it\. Reconstruction cannot create decision\-relevant information absent from its inputs and may introduce another rhetorical realization, but this information\-theoretic argument does not guarantee better performance for a fixed LLM or order our implemented pipelines\. We therefore treat direct review as a design choice and compare it empirically with ReconstructReview\. Appendix[A](https://arxiv.org/html/2609.39027#A1)gives the formal invariant\-representation, conditional\-stability, fusion, and Blackwell arguments\.
### 4\.2Science\-Core Extraction: Isolating the Reported Scientific Record
The science\-core branch first uses an LLM to extract a structured science core from the complete manuscript\. The extraction focuses on information directly relevant to scientific evaluation, including the central research idea and claims, problem formulation, mathematical formulations and derivations, methods and assumptions, experimental or theoretical evidence, reported results, author\-stated contributions, reproducibility information, and stated limitations\.
The LLM is instructed to extract this information in objective, third\-person language while preserving its attribution to the original manuscript\. This is particularly important for scientific claims and author\-stated contributions, where rhetorical framing could otherwise carry into the science core\. The extractor is instructed to attribute claims and contributions to the authors, preserve reported results and numerical evidence, and avoid subjective assessments, evaluative language, emotional framing, and unsupported interpretations\. It does not determine whether a claim is convincing, whether a contribution is significant, or how the paper should be scored\.
The science core is produced as a text\-based scientific record rather than a shortened rewrite of the manuscript\. Figures are not retained directly\. Tables are instead explicitly transcribed so that their reported values, comparisons, and other scientific information remain available to the downstream reviewer\. The prompt asks the extractor to retain equations, mathematical arguments, reported numbers, and other details whenever they can be reliably recovered\. When information cannot be located or read reliably, the extractor records this uncertainty rather than inferring or reconstructing the missing content\.
### 4\.3Science\-Core Branch: Evaluating the Extracted Record
After extraction, the science\-core branch reviews only the extracted record\. The original manuscript is not provided to this branch at the review stage, so its judgment is based on the retained scientific record rather than its original rhetorical presentation\.
We build the review protocol on theStandardprotocol used inRobustReview, while adapting its instructions to the structure of the science core\. The prompt explains the role of each extracted section and how information across sections should be combined when forming a scientific judgment\. In particular, the reviewer is instructed to connect the stated research problem and claims with the corresponding methods, assumptions, mathematical reasoning, and empirical or theoretical evidence; to assess whether the reported results support the author\-stated claims and contributions; and to use the reproducibility information and stated limitations when evaluating the strength and scope of the evidence\. Information that is explicitly absent from the manuscript can be treated as missing scientific support when relevant, while extraction uncertainty is handled separately\.
The review retains the ICLR\-style evaluation criteria and scoring scheme used byStandard\. The reviewer produces written feedback, including a summary of the work, strengths, weaknesses, and questions, and an overall assessment, together with supporting scientific\-evaluation scores\. The score anchors follow the corresponding ICLR\-style scales, with the review decision grounded in the scientific evidence contained in the science core\.
### 4\.4Dual\-Branch Fusion: Integrating Manuscript and Science\-Core Judgments
In parallel with the science\-core branch, the manuscript branch applies theStrictprotocol from Section[2\.3](https://arxiv.org/html/2609.39027#S2.SS3)directly to the complete manuscript PDF\. This branch retains the ordinary manuscript\-level review context, while its evidence\-focused rubric provides the conventional judgment used in the fusion\. Both branches produce an overall assessment on the same ICLR\-style scale\.
LetM\(x\)M\(x\)denote the manuscript\-branch overall\-assessment score and letB\(x\)B\(x\)denote the science\-core\-branch score for manuscript realizationxx\. The finalSciCorescore is their unweighted arithmetic mean:
SSciCore\(x\)=12M\(x\)\+12B\(x\)\.S\_\{\\mbox\{SciCore\}\}\(x\)=\\frac\{1\}\{2\}M\(x\)\+\\frac\{1\}\{2\}B\(x\)\.\(2\)Equal weighting provides a direct symmetric combination without introducing a tuned fusion parameter\. The manuscript branch and science\-core branch remain independently interpretable, which allows the evaluation to distinguish the behavior of each component from that of the fused reviewer\.
## 5SciCore: Results and Analysis
### 5\.1Main Results: Rhetorical Robustness onRobustReview
Table[1](https://arxiv.org/html/2609.39027#S3.T1)comparesSciCorewith the general\-purpose, specialized, and agentic AI reviewers onRobustReview\.SciCoreleads the primary comparison on ICC \(0\.775\), SPR \(0\.652\), and discriminability \(0\.726\), while attaining the lowest Human MAE \(1\.072\) and second\-highest human Spearman correlation \(0\.488\)\. However, it does not achieve the lowest MAD or Drift SD, and its human Spearman correlation remains below GPT\-5\.5 underStrict\.SciCoretherefore achieves a leading joint stability\-discrimination profile while maintaining competitive human alignment\. Paired paper\-family bootstrap analysis supports its improvements over Manuscript\-Strict across all five robustness metrics, alongside improved human alignment relative to Core\-Adapted \(Appendix[E\.4](https://arxiv.org/html/2609.39027#A5.SS4)\)\. These results reinforce the complementary roles of the two branches and demonstrate the effectiveness of a simple equal\-weight fusion\.
### 5\.2Dissecting the Gains: Architecture and Review\-Policy Ablations
Table[2](https://arxiv.org/html/2609.39027#S5.T2)isolates the design choices using a common GPT\-5\.5 backbone\. Manuscript\-Standard and Manuscript\-Strict review the PDF directly\. ReconstructReview reconstructs a paper from the science core, permitted figures, and bibliography before applyingStandardreview \(Appendix[C\.6](https://arxiv.org/html/2609.39027#A3.SS6)\)\. Four core\-only configurations share the same cached science core and differ only in review policy: Standard, Strict, Persistent, or the adapted protocol\. The finalSciCorerow averages Core\-Adapted with Manuscript\-Strict\.
Directly reviewing the science core provides a stronger robustness branch than reconstruction in this pipeline\.Relative to Manuscript\-Standard, ReconstructReview worsens MAD from 0\.598 to 0\.651 and Drift SD from 0\.956 to 1\.101, with only modest gains on the joint metrics\. Core\-Standard then improves all seven metrics over ReconstructReview\. The result shows that reconstruction does not recover the robustness of direct science\-core review here\. Proposition[A\.2](https://arxiv.org/html/2609.39027#A1.Thmproposition2)provides an information\-theoretic rationale for direct review but does not predict the ordering of the implemented pipelines\.
Content\-focused review remains nontrivial after extraction\.Core\-Persistent worsens all five robustness metrics relative to Core\-Standard, whereas Core\-Adapted achieves the lowest MAD and Drift SD and the highest ICC and SPR among the core\-only configurations\. The review policy thus affects the balance between score stability and paper discrimination even when the extracted representation is held fixed\.
The two branches provide complementary judgments\.Core\-Adapted is more stable than Manuscript\-Strict but aligns less closely with human scores\. Fusion reduces MAD from 0\.766 to 0\.476 and increases ICC from 0\.632 to 0\.775, improving all five robustness metrics over Manuscript\-Strict\. It also improves human alignment, ICC, and discriminability over Core\-Adapted, which retains lower MAD and Drift SD and higher SPR\. Paired bootstrap intervals support these directions \(Appendix[E\.4](https://arxiv.org/html/2609.39027#A5.SS4)\)\.
With Manuscript\-Strict and equal weighting fixed, the science\-core branch also improves all five robustness metrics over adding Standard or Persistent manuscript review, supported by paired bootstrap intervals\. These controls share its two\-score averaging rule and half\-point resolution\. Appendix[E\.2](https://arxiv.org/html/2609.39027#A5.SS2)discusses the mechanical effects of averaging and a three\-review control matchingSciCore’s nominal call count\. That control has lower Drift SD and higher ICC, whileSciCorehas better point estimates on the other five metrics\.
Table 2:Within\-backbone analysis ofSciCoreand its components\.Every configuration uses GPT\-5\.5\. MAD and Drift SD are direct within\-paper metrics, whereas ICC, SPR, and discriminability jointly relate within\-paper consistency to cross\-paper variation or separation\. The finalSciCorerow averages the Manuscript\-Strict and Core\-Adapted scores\. The best point estimate in each column is shown in bold and the second\-best is underlined\.ConfigurationHuman alignmentWithin\-paper stabilityJoint stability\-discriminationH\-MAE↓\\downarrowSpearman↑\\uparrowMAD↓\\downarrowDrift SD↓\\downarrowICC↑\\uparrowSPR↑\\uparrowDiscrim\.↑\\uparrowManuscript\-Based ReviewManuscript\-Standard1\.2940\.4480\.5980\.9560\.6150\.5550\.649Manuscript\-Strict1\.0780\.5290\.7661\.2020\.6320\.5070\.672ReconstructReview1\.4900\.1330\.6511\.1010\.6230\.5660\.664Science\-Core BranchCore\-Standard1\.3940\.2160\.3540\.7370\.7340\.6770\.697Core\-Strict1\.3500\.2640\.4030\.8700\.6080\.5840\.654Core\-Persistent1\.3460\.2890\.4580\.8210\.7080\.5860\.681Core\-Adapted1\.4610\.2850\.2920\.6500\.7510\.7100\.695Dual\-Branch FusionSciCore\(ours\)1\.0720\.4880\.4760\.6870\.7750\.6520\.726
### 5\.3Representation Analysis: Science\-Core Stability and Paper Separation
A useful science\-core representation should remain similar across rhetorical variants of the same paper while remaining distinct across different papers\. Table[3](https://arxiv.org/html/2609.39027#S5.T3)compares each variant core with its matched original and with nonmatching originals across three extractors and two rewrite producers with cosine similarity of text embeddings\.
Matched cores are consistently much more similar than cores from different papers\. For GPT\-5\.5, matched similarity is 0\.978, while cross\-paper similarity ranges from 0\.616 to 0\.619\. This gap provides embedding\-level evidence that the extracted science cores remain stable across rhetorical rewrites while retaining paper\-specific differences\. Condition\-level comparisons are reported in Appendix[E\.5](https://arxiv.org/html/2609.39027#A5.SS5)\.
Table 3:Science\-core cosine similarity across extractor models and rewrite producers\.All matched comparisons pair the science core extracted from each rewritten manuscript with the core extracted from the original manuscript of the same paper; different\-paper controls pair it with the 59 nonmatching original cores\. Each matched cell contains 600 pairs and each control cell contains 35,400 comparisons per producer\. P05 denotes the fifth percentile of pair\-level similarities\.
### 5\.4Criterion\-Level Behavior of the Science\-Core Branch
Table[4](https://arxiv.org/html/2609.39027#S5.T4)shows that branch behavior is criterion\-dependent\. Science\-core review generally improves soundness stability, but the contribution is less consistent, and human alignment does not improve uniformly\. Presentation is diagnostic because manuscript\-based review evaluates the paper’s presentation, whereas science\-core review evaluates the clarity and completeness of the extracted scientific record\.
Confidence exposes a concrete false\-robustness failure mode\. Core\-Standard and Core\-Strict assign a constant confidence score, and Core\-Persistent produces almost no variation, so their near\-zero drift reflects output collapse and coincides with zero or near\-zero SPR\. Core\-Adapted restores paper\-level variation and raises confidence SPR to 0\.394\. This result again shows why within\-paper stability cannot be interpreted without discrimination\.
Table 4:Secondary\-score behavior across GPT\-5\.5 branch configurations\.For each secondary score, we report direct within\-paper stability \(MAD\), joint stability\-discrimination \(SPR\), and Spearman correlation with the corresponding mean human score\. Presentation is marked as diagnostic because its meaning differs between manuscript\-based and direct science\-core review\. A dash denotes an undefined Spearman correlation due to constant scores; SPR is defined as zero for fully constant outputs\.
### 5\.5Cross\-Backbone Generalization of the Science\-Core Branch
Table[5](https://arxiv.org/html/2609.39027#S5.T5)compares Core\-Adapted with Manuscript\-Standard under the same backbone\. We useStandardas a common baseline rather than selecting the strongest direct\-review protocol separately for each model\. Core\-Adapted improves ICC for all three backbones\. GPT\-5\.5 improves on all five robustness metrics, GPT\-5\-mini improves MAD, Drift SD, and ICC but weakens SPR and discriminability, and GLM\-5\.2 improves all seven reported metrics\. Human alignment does not improve consistently\.
Table 5:Cross\-backbone evaluation of the science\-core branch\.For each backbone, Manuscript\-Standard applies theStandardprotocol to the full paper, whereas Core\-Adapted uses the same backbone for science\-core extraction and review\. Bold marks the better result within each backbone pair\.The manuscript\-side Strict evaluations for these backbones are already included in Table[1](https://arxiv.org/html/2609.39027#S3.T1)\. All reviewer configurations are evaluated on the same matched corpus, which includes rewrites produced independently by GPT\-5\.5 and Claude Opus 4\.8\. Because each experiment changes the backbone for both extraction and core review, it measures branch\-level transfer without isolating which stage limits performance and does not test the final dual\-branch fusion across backbones\. The evidence therefore supports partial, backbone\-dependent transfer of the science\-core branch rather than a universal gain\.
### 5\.6Science\-Core Weight Sensitivity
As a sensitivity analysis ofSciCore, we alter the science\-core weightα\\alphato test how the fusion balance affects performance\. We examine the science\-core branch paired with Manuscript\-Standard, Manuscript\-Strict, and Manuscript\-Persistent reviews:
Sα,p\(x\)=αB\(x\)\+\(1−α\)Mp\(x\),S\_\{\\alpha,p\}\(x\)=\\alpha B\(x\)\+\(1\-\\alpha\)M\_\{p\}\(x\),\(3\)
whereB\(x\)B\(x\)is the science\-core\-branch score andMp\(x\)M\_\{p\}\(x\)is the manuscript\-branch score under protocolpp\. Our prespecified configuration uses Manuscript\-Strict withα=0\.5\\alpha=0\.5\. The sweep characterizes sensitivity around this fixed design\. For descriptive visualization, each metric is direction\-aligned and min\-max normalized over the fullα∈\[0,1\]\\alpha\\in\[0,1\]sweep\.
Figure 2:Science\-core weight sensitivity\.Normalized performance across science\-core weights under three manuscript protocols\. Higher values indicate better performance for each metric\. The dashed line marks the prespecified equal\-weight fusion \(α=0\.5\\alpha=0\.5\)\. Complete numerical results are reported in Tables[17](https://arxiv.org/html/2609.39027#A5.T17)–[19](https://arxiv.org/html/2609.39027#A5.T19)\.
Larger science\-core weights tend to improve robustness while weakening human alignment overall\. Across the three manuscript protocols, Manuscript\-Persistent favors robustness most strongly, while Manuscript\-Strict retains the best human alignment\. Under Manuscript\-Strict,α=0\.5\\alpha=0\.5lies near the transition between these two objectives and maintains strong ICC, SPR, and discriminability without a large loss in human alignment\.
Equal weighting is a simple and effective default that requires no search and naturally balances the science\-core and manuscript branches\. The prespecified equal\-weight fusion balances rhetorical robustness and human alignment and is Pareto non\-dominated among the evaluated protocol\-weight combinations on the seven reported point estimates\.
## 6Related Work
LLM\-as\-a\-judge and AI peer\-review research evaluates human agreement, review quality, score prediction, and workflow capability, while documenting sensitivity to order, length, prompts, and evaluator identity\([Liu et al\., 2023](https://arxiv.org/html/2609.39027#bib.bib32);[Zheng et al\., 2023](https://arxiv.org/html/2609.39027#bib.bib52);[Wang et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib45);[Dubois et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib11);[Liang et al\., 2023](https://arxiv.org/html/2609.39027#bib.bib30);[Zhou et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib53)\)\. Studies of automated review further reveal vulnerabilities to hidden instructions, artificial perturbations, and visible content\-preserving revisions\([Ye et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib50);[Lin et al\., 2025](https://arxiv.org/html/2609.39027#bib.bib31);[Kaneko, 2026](https://arxiv.org/html/2609.39027#bib.bib20);[Li et al\., 2026b](https://arxiv.org/html/2609.39027#bib.bib25);[Baumann et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib3)\)\.[Li et al\. \(2026c\)](https://arxiv.org/html/2609.39027#bib.bib26)characterize rhetorical reward hacking through controlled full\-manuscript rewrites\. Our work formulates rhetorical robustness as a joint requirement of within\-paper stability and cross\-paper discrimination, and introducesSciCoreto improve this balance while maintaining competitive human alignment\.
Existing interventions control particular biases or judge intermediate representations\([Dubois et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib11);[Li et al\., 2026d](https://arxiv.org/html/2609.39027#bib.bib29)\)\. Rather than relying only on instructions to discount rhetoric,SciCoreaugments full\-manuscript review with a science\-core branch that extracts and evaluates structured scientific content, then averages the two judgments\. This retains manuscript\-level assessment while adding a content\-normalized view, motivated by invariant representation and Blackwell comparison\([Eaton, 1989](https://arxiv.org/html/2609.39027#bib.bib13);[Dubois et al\., 2021](https://arxiv.org/html/2609.39027#bib.bib10);[Blackwell, 1953](https://arxiv.org/html/2609.39027#bib.bib6)\)\. Appendix[B](https://arxiv.org/html/2609.39027#A2)provides the extended discussion\.
## 7Conclusion
We identify rhetorical robustness as an important but comparatively neglected requirement for trustworthy AI reviewers\. Within\-paper stability is necessary but not sufficient: low drift can be obtained by collapsing scores across papers, so robustness must also preserve between\-paper discrimination\.RobustReviewoperationalizes this requirement with direct within\-paper metrics and joint stability\-discrimination metrics that relate same\-paper consistency to cross\-paper variation or separation\. It shows that the evaluated content\-focused prompting protocol produces model\-dependent rather than consistent robustness gains, low score drift can conceal weak discrimination, and human alignment ranks reviewers differently from matched robustness evaluation\. Content\-focused review is therefore a nontrivial design problem that is not resolved by score\-sensitivity or human\-agreement measures alone\.
SciCoreprovides one method for incorporating this requirement into reviewer design\. It augments a conventional full\-manuscript judgment with an equally weighted content\-normalized judgment from a science\-core branch\. The extracted science cores are stable and paper\-specific under our representation diagnostics, while the fused reviewer achieves a leading joint stability\-discrimination profile and competitive human alignment in the primary GPT\-5\.5 comparison\. Branch\-level behavior remains criterion\- and backbone\-dependent\. Together, the benchmark and method position rhetorical robustness as a distinct evaluation and design objective for more reliable AI reviewers\.
## 8Limitations
RobustReviewcontains 60 ICLR 2026 submissions sampled evenly across score strata from papers with matched arXiv sources, so its results may not generalize to other venues, fields, or manuscript formats\. The rhetorical rewrites are designed to preserve reported scientific content but cannot guarantee exact equivalence, particularly under complex conditions; the automated fidelity audit identifies a nonzero mismatch rate\. Mean human scores provide only a limited external reference and do not capture disagreement among reviewers\. The finalSciCorefusion is evaluated with GPT\-5\.5, while the criterion\-level and cross\-backbone analyses characterize the science\-core branch rather than the fused reviewer\. Those branch experiments vary extraction and review together rather than isolating the two stages\. Science\-core extraction can omit or misread scientific details and requires additional inference, and equal\-weight score fusion assumes that the two branch assessments are comparable on the shared review scale\.
## References
- Anthropic \(2026a\)Anthropic\.Introducing Claude Opus 4\.8\.[https://www\.anthropic\.com/news/claude\-opus\-4\-8](https://www.anthropic.com/news/claude-opus-4-8), 2026a\.
- Anthropic \(2026b\)Anthropic\.Introducing Claude Sonnet 5\.[https://www\.anthropic\.com/news/claude\-sonnet\-5](https://www.anthropic.com/news/claude-sonnet-5), 2026b\.
- Baumann et al\. \(2026\)Joachim Baumann, Jiaxin Pei, Sanmi Koyejo, and Dirk Hovy\.Stop automating peer review without rigorous evaluation\.*arXiv preprint arXiv:2605\.03202*, 2026\.
- Bavaresco et al\. \(2025\)Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, Andre Martins, Philipp Mondorf, Vera Neplenbroek, Sandro Pezzelle, Barbara Plank, David Schlangen, Alessandro Suglia, Aditya K Surikuchi, Ece Takmaz, and Alberto Testoni\.LLMs instead of human judges? a large scale empirical study across 20 NLP evaluation tasks\.In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors,*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\)*, pages 238–255, Vienna, Austria, July 2025\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-252\-7\.[10\.18653/v1/2025\.acl\-short\.20](https://doi.org/10.18653/v1/2025.acl-short.20)\.[https://aclanthology\.org/2025\.acl\-short\.20/](https://aclanthology.org/2025.acl-short.20/)\.
- Bhat and Varma \(2026\)Savita Bhat and Vasudeva Varma\.All prompts are created equal? evaluating robustness of llm judges against non\-adversarial prompt variations\.In*Findings of the Association for Computational Linguistics: ACL 2026*, pages 38730–38745, 2026\.
- Blackwell \(1953\)David Blackwell\.Equivalent comparisons of experiments\.*The Annals of Mathematical Statistics*, 24\(2\):265–272, 1953\.[10\.1214/aoms/1177729032](https://doi.org/10.1214/aoms/1177729032)\.
- Chen et al\. \(2026\)Zeyuan Chen, Ziqing Yang, Yihan Ma, Michael Backes, and Yang Zhang\.Peercheck: Enhancing llm\-generated academic reviews towards human\-level quality\.In*Findings of the Association for Computational Linguistics: ACL 2026*, pages 23362–23386, 2026\.
- Collu et al\. \(2026\)Matteo Gioele Collu, Umberto Salviati, Roberto Confalonieri, Mauro Conti, and Giovanni Apruzzese\.Misleading large language models used \(or misused\) in scientific peer\-reviewing via hidden prompt\-injection attacks\.*ACM Transactions on AI Security and Privacy*, 2026\.
- Du \(2025\)Shurui Du\.TitleTrap: Probing presentation bias in LLM\-based scientific reviewing\.In Mousumi Akter, Tahiya Chowdhury, Steffen Eger, Christoph Leiter, Juri Opitz, and Erion Çano, editors,*Proceedings of the 5th Workshop on Evaluation and Comparison of NLP Systems*, pages 119–125, Mumbai, India, December 2025\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-305\-0\.[10\.18653/v1/2025\.eval4nlp\-1\.10](https://doi.org/10.18653/v1/2025.eval4nlp-1.10)\.[https://aclanthology\.org/2025\.eval4nlp\-1\.10/](https://aclanthology.org/2025.eval4nlp-1.10/)\.
- Dubois et al\. \(2021\)Yann Dubois, Benjamin Bloem\-Reddy, Karen Ullrich, and Chris J\. Maddison\.Lossy compression for lossless prediction\.In*Advances in Neural Information Processing Systems*, volume 34, 2021\.[https://proceedings\.neurips\.cc/paper/2021/hash/7535bbb91c8fde347ad861f293126633\-Abstract\.html](https://proceedings.neurips.cc/paper/2021/hash/7535bbb91c8fde347ad861f293126633-Abstract.html)\.
- Dubois et al\. \(2024\)Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B Hashimoto\.Length\-controlled alpacaeval: A simple way to debias automatic evaluators\.*arXiv preprint arXiv:2404\.04475*, 2024\.
- Dycke and Gurevych \(2026\)Nils Dycke and Iryna Gurevych\.Automatic reviewers fail to detect faulty reasoning in research papers: A new counterfactual evaluation framework\.*Transactions of the Association for Computational Linguistics*, 14:465–488, 2026\.
- Eaton \(1989\)Morris L\. Eaton\.*Group Invariance in Applications in Statistics*, volume 1 of*NSF\-CBMS Regional Conference Series in Probability and Statistics*\.Institute of Mathematical Statistics and American Statistical Association, 1989\.[10\.1214/cbms/1462061029](https://doi.org/10.1214/cbms/1462061029)\.
- Fang et al\. \(2026\)Haishuo Fang, Yue Feng, and Iryna Gurevych\.From passive generation to investigation: A proactive scientific peer review agent, 2026\.[https://arxiv\.org/abs/2606\.13349](https://arxiv.org/abs/2606.13349)\.
- Fytas et al\. \(2021\)Panagiotis Fytas, Georgios Rizos, and Lucia Specia\.What makes a scientific paper be accepted for publication?In*Proceedings of the first workshop on causal inference and NLP*, pages 44–60, 2021\.
- GLM\-5\-Team et al\. \(2026\)GLM\-5\-Team, :, Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chenghua Huang, Chengxing Xie, Chenzheng Zhu, Congfeng Yin, Cunxiang Wang, Gengzheng Pan, Hao Zeng, Haoke Zhang, Haoran Wang, Huilong Chen, Jiajie Zhang, Jian Jiao, Jiaqi Guo, Jingsen Wang, Jingzhao Du, Jinzhu Wu, Kedong Wang, Lei Li, Lin Fan, Lucen Zhong, Mingdao Liu, Mingming Zhao, Pengfan Du, Qian Dong, Rui Lu, Shuang\-Li, Shulin Cao, Song Liu, Ting Jiang, Xiaodong Chen, Xiaohan Zhang, Xuancheng Huang, Xuezhen Dong, Yabo Xu, Yao Wei, Yifan An, Yilin Niu, Yitong Zhu, Yuanhao Wen, Yukuo Cen, Yushi Bai, Zhongpei Qiao, Zihan Wang, Zikang Wang, Zilin Zhu, Ziqiang Liu, Zixuan Li, Bojie Wang, Bosi Wen, Can Huang, Changpeng Cai, Chao Yu, Chen Li, Chengwei Hu, Chenhui Zhang, Dan Zhang, Daoyan Lin, Dayong Yang, Di Wang, Ding Ai, Erle Zhu, Fangzhou Yi, Feiyu Chen, Guohong Wen, Hailong Sun, Haisha Zhao, Haiyi Hu, Hanchen Zhang, Hanrui Liu, Hanyu Zhang, Hao Peng, Hao Tai, Haobo Zhang, He Liu, Hongwei Wang, Hongxi Yan, Hongyu Ge, Huan Liu, Huanpeng Chu, Jia’ni Zhao, Jiachen Wang, Jiajing Zhao, Jiamin Ren, Jiapeng Wang, Jiaxin Zhang, Jiayi Gui, Jiayue Zhao, Jijie Li, Jing An, Jing Li, Jingwei Yuan, Jinhua Du, Jinxin Liu, Junkai Zhi, Junwen Duan, Kaiyue Zhou, Kangjian Wei, Ke Wang, Keyun Luo, Laiqiang Zhang, Leigang Sha, Liang Xu, Lindong Wu, Lintao Ding, Lu Chen, Minghao Li, Nianyi Lin, Pan Ta, Qiang Zou, Rongjun Song, Ruiqi Yang, Shangqing Tu, Shangtong Yang, Shaoxiang Wu, Shengyan Zhang, Shijie Li, Shuang Li, Shuyi Fan, Wei Qin, Wei Tian, Weining Zhang, Wenbo Yu, Wenjie Liang, Xiang Kuang, Xiangmeng Cheng, Xiangyang Li, Xiaoquan Yan, Xiaowei Hu, Xiaoying Ling, Xing Fan, Xingye Xia, Xinyuan Zhang, Xinze Zhang, Xirui Pan, Xu Zou, Xunkai Zhang, Yadi Liu, Yandong Wu, Yanfu Li, Yidong Wang, Yifan Zhu, Yijun Tan, Yilin Zhou, Yiming Pan, Ying Zhang, Yinpei Su, Yipeng Geng, Yong Yan, Yonglin Tan, Yuean Bi, Yuhan Shen, Yuhao Yang, Yujiang Li, Yunan Liu, Yunqing Wang, Yuntao Li, Yurong Wu, Yutao Zhang, Yuxi Duan, Yuxuan Zhang, Zezhen Liu, Zhengtao Jiang, Zhenhe Yan, Zheyu Zhang, Zhixiang Wei, Zhuo Chen, Zhuoer Feng, Zijun Yao, Ziwei Chai, Ziyuan Wang, Zuzhou Zhang, Bin Xu, Minlie Huang, Hongning Wang, Juanzi Li, Yuxiao Dong, and Jie Tang\.Glm\-5: from vibe coding to agentic engineering, 2026\.[https://arxiv\.org/abs/2602\.15763](https://arxiv.org/abs/2602.15763)\.
- Google \(2026\)Google\.Introducing Gemini 3\.6 Flash, 3\.5 Flash\-Lite, and 3\.5 Flash Cyber\.[https://blog\.google/innovation\-and\-ai/models\-and\-research/gemini\-models/gemini\-3\-6\-flash\-3\-5\-flash\-lite\-3\-5\-flash\-cyber/](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/), 2026\.
- Idahl and Ahmadi \(2025\)Maximilian Idahl and Zahra Ahmadi\.OpenReviewer: A specialized large language model for generating critical scientific paper reviews\.In Nouha Dziri, Sean \(Xiang\) Ren, and Shizhe Diao, editors,*Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(System Demonstrations\)*, pages 550–562, Albuquerque, New Mexico, April 2025\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-191\-9\.[10\.18653/v1/2025\.naacl\-demo\.44](https://doi.org/10.18653/v1/2025.naacl-demo.44)\.[https://aclanthology\.org/2025\.naacl\-demo\.44/](https://aclanthology.org/2025.naacl-demo.44/)\.
- James et al\. \(2024\)Joseph James, Chenghao Xiao, Yucheng Li, and Chenghua Lin\.On the rigour of scientific writing: Criteria, analysis, and insights\.In Yaser Al\-Onaizan, Mohit Bansal, and Yun\-Nung Chen, editors,*Findings of the Association for Computational Linguistics: EMNLP 2024*, pages 6523–6538, Miami, Florida, USA, November 2024\. Association for Computational Linguistics\.[10\.18653/v1/2024\.findings\-emnlp\.380](https://doi.org/10.18653/v1/2024.findings-emnlp.380)\.[https://aclanthology\.org/2024\.findings\-emnlp\.380/](https://aclanthology.org/2024.findings-emnlp.380/)\.
- Kaneko \(2026\)Masahiro Kaneko\.Paraphrasing adversarial attack on llm\-as\-a\-reviewer\.*arXiv preprint arXiv:2601\.06884*, 2026\.
- Kang et al\. \(2018\)Dongyeop Kang, Waleed Ammar, Bhavana Dalvi, Madeleine Van Zuylen, Sebastian Kohlmeier, Eduard Hovy, and Roy Schwartz\.A dataset of peer reviews \(peerread\): Collection, insights and nlp applications\.In*Proceedings of the 2018 conference of the North American chapter of the Association for Computational Linguistics: Human language technologies, volume 1 \(long papers\)*, pages 1647–1661, 2018\.
- Kim et al\. \(2024\)Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, and Minjoon Seo\.Prometheus: Inducing fine\-grained evaluation capability in language models\.In*The Twelfth International Conference on Learning Representations*, 2024\.[https://openreview\.net/forum?id=8euJaTveKw](https://openreview.net/forum?id=8euJaTveKw)\.
- Lee et al\. \(2025\)Dongryeol Lee, Yerin Hwang, Yongil Kim, Joonsuk Park, and Kyomin Jung\.Are llm\-judges robust to expressions of uncertainty? investigating the effect of epistemic markers on llm\-based evaluation\.In*Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\)*, pages 8962–8984, 2025\.
- Li et al\. \(2026a\)Haowen Li, Yoichi Ishibashi, and Masafumi Oyamada\.Evaluating the impact of reviewer guideline design on llm\-based automated peer review\.In*Findings of the Association for Computational Linguistics: ACL 2026*, pages 30223–30240, 2026a\.
- Li et al\. \(2026b\)Lin Li, Qi Zhang, Xander Davies, Jianing Qiu, and Yarin Gal\.Gaming ai\-assisted peer reviews poses new risks to the scientific community\.*arXiv preprint arXiv:2606\.10159*, 2026b\.
- Li et al\. \(2026c\)Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, and Tianyi Zhou\.How can rhetoric reward\-hack ai reviewers? dissecting rhetorical sensitivity in ai\-based peer review\.*arXiv preprint arXiv:2608\.08975*, 2026c\.
- Li et al\. \(2025a\)Ruochi Li, Haoxuan Zhang, Edward Gehringer, Ting Xiao, Junhua Ding, and Haihua Chen\.Unveiling the merits and defects of llms in automatic review generation for scientific papers\.In*2025 IEEE International Conference on Data Mining \(ICDM\)*, pages 1370–1379\. IEEE, 2025a\.
- Li et al\. \(2025b\)Yifei Li, Xiaoting Xu, Dongqing Lyu, Zhen Zhang, Juan Xie, and Ying Cheng\.Developing a criteria framework for peer review: a critical interpretive synthesis\.*Learned Publishing*, 38\(3\):e2016, 2025b\.
- Li et al\. \(2026d\)Zhuochun Li, Yong Zhang, Ming Li, Yuelyu Ji, Yiming Zeng, Ning Cheng, Yun Zhu, Yanmeng Wang, Shaojun Wang, Jing Xiao, and Daqing He\.Rethinking LLM\-as\-a\-judge: Representation\-as\-a\-judge with small language models via semantic capacity asymmetry\.In*The Fourteenth International Conference on Learning Representations*, 2026d\.[https://openreview\.net/forum?id=VAISvCsrvG](https://openreview.net/forum?id=VAISvCsrvG)\.
- Liang et al\. \(2023\)Weixin Liang, Yuhui Zhang, Hancheng Cao, Binglu Wang, Daisy Ding, Xinyu Yang, Kailas Vodrahalli, Siyu He, Daniel Smith, Yian Yin, Daniel McFarland, and James Zou\.Can large language models provide useful feedback on research papers? a large\-scale empirical analysis, 2023\.[https://arxiv\.org/abs/2310\.01783](https://arxiv.org/abs/2310.01783)\.
- Lin et al\. \(2025\)Tzu\-Ling Lin, Wei\-Chih Chen, Teng\-Fang Hsiao, Hou\-I Liu, Ya\-Hsin Yeh, Yu\-Kai Chan, Wen\-Sheng Lien, Po\-Yen Kuo, Philip S\. Yu, and Hong\-Han Shuai\.Breaking the reviewer: Assessing the vulnerability of large language models in automated peer review under textual adversarial attacks\.In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng, editors,*Findings of the Association for Computational Linguistics: EMNLP 2025*, pages 4819–4839, Suzhou, China, November 2025\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-335\-7\.[10\.18653/v1/2025\.findings\-emnlp\.259](https://doi.org/10.18653/v1/2025.findings-emnlp.259)\.[https://aclanthology\.org/2025\.findings\-emnlp\.259/](https://aclanthology.org/2025.findings-emnlp.259/)\.
- Liu et al\. \(2023\)Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu\.G\-eval: Nlg evaluation using gpt\-4 with better human alignment\.In*Proceedings of the 2023 conference on empirical methods in natural language processing*, pages 2511–2522, 2023\.
- Lu et al\. \(2024\)Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha\.The ai scientist: Towards fully automated open\-ended scientific discovery, 2024\.[https://arxiv\.org/abs/2408\.06292](https://arxiv.org/abs/2408.06292)\.
- Moonshot AI \(2026\)Moonshot AI\.Kimi K2\.6: Advancing Open\-Source Coding\.[https://www\.kimi\.ai/blog/kimi\-k2\-6/](https://www.kimi.ai/blog/kimi-k2-6/), 2026\.
- OpenAI \(2025\)OpenAI\.gpt\-oss\-120b & gpt\-oss\-20b model card, 2025\.[https://arxiv\.org/abs/2508\.10925](https://arxiv.org/abs/2508.10925)\.
- OpenAI \(2025\)OpenAI\.GPT\-5 mini\.[https://developers\.openai\.com/api/docs/models/gpt\-5\-mini](https://developers.openai.com/api/docs/models/gpt-5-mini), 2025\.
- OpenAI \(2026\)OpenAI\.GPT\-5\.5\.[https://openai\.com/index/introducing\-gpt\-5\-5/](https://openai.com/index/introducing-gpt-5-5/), 2026\.
- Panickssery et al\. \(2024\)Arjun Panickssery, Samuel R Bowman, and Shi Feng\.Llm evaluators recognize and favor their own generations\.*Advances in Neural Information Processing Systems*, 37:68772–68802, 2024\.
- Qwen Team \(2026\)Qwen Team\.Qwen3\.5: Towards native multimodal agents, February 2026\.[https://qwen\.ai/blog?id=qwen3\.5](https://qwen.ai/blog?id=qwen3.5)\.
- Stureborg et al\. \(2024\)Rickard Stureborg, Dimitris Alikaniotis, and Yoshi Suhara\.Large language models are inconsistent and biased evaluators\.*arXiv preprint arXiv:2405\.01724*, 2024\.
- Thakkar et al\. \(2026\)Nitya Thakkar, Mert Yuksekgonul, Jake Silberg, Animesh Garg, Nanyun Peng, Fei Sha, Rose Yu, Carl Vondrick, and James Zou\.A large\-scale randomized study of large language model feedback in peer review\.*Nature Machine Intelligence*, 8\(3\):326–336, 2026\.
- The OpenJudge Team \(2025\)The OpenJudge Team\.Openjudge: A unified framework for holistic evaluation and quality rewards, 07 2025\.[https://github\.com/agentscope\-ai/OpenJudge](https://github.com/agentscope-ai/OpenJudge)\.
- Vasu et al\. \(2026\)Sai Suresh Macharla Vasu, Ivaxi Sheth, Hui\-Po Wang, Ruta Binkyte, and Mario Fritz\.Justice in judgment: Unveiling \(hidden\) bias in llm\-assisted peer reviews\.In*Findings of the Association for Computational Linguistics: ACL 2026*, pages 307–330, 2026\.
- Wang et al\. \(2026\)Chenguang Wang, Ming Li, Adebayo Braimah, Chenrui Fan, Tuo Wang, Weijie Guan, Ruiyi Zhang, Tianyi Zhou, and Dawei Zhou\.The emerging ai paper\-review arms race: Adversarial co\-evolution in scholarly publishing\.*arXiv preprint arXiv:2609\.07713*, 2026\.
- Wang et al\. \(2024\)Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, et al\.Large language models are not fair evaluators\.In*Proceedings of the 62nd annual meeting of the association for computational linguistics \(volume 1: Long papers\)*, pages 9440–9450, 2024\.
- Wang et al\. \(2020\)Qingyun Wang, Qi Zeng, Lifu Huang, Kevin Knight, Heng Ji, and Nazneen Fatema Rajani\.Reviewrobot: Explainable paper review generation based on knowledge synthesis\.In*Proceedings of the 13th International Conference on Natural Language Generation*, pages 384–397, 2020\.
- Weng et al\. \(2025\)Yixuan Weng, Minjun Zhu, Guangsheng Bao, Hongbo Zhang, Jindong Wang, Yue Zhang, and Linyi Yang\.Cycleresearcher: Improving automated research via automated review, 2025\.[https://arxiv\.org/abs/2411\.00816](https://arxiv.org/abs/2411.00816)\.
- Yang et al\. \(2026a\)Xianglin Yang, Bryan Hooi, Gelei Deng, Tianwei Zhang, and Jin Song Dong\.Turning bias into bugs: Bandit\-guided style manipulation attacks on llm judges\.*arXiv preprint arXiv:2605\.26156*, 2026a\.
- Yang et al\. \(2026b\)Xu Yang, Zhizhou Sha, Junbo Li, Jian Yu, Yifan Sun, Matthew Zhao, Jinrui Fang, Xinyue Guo, Yining Wu, Xu Hu, Yifu Luo, Qiang Liu, and Zhangyang Wang\.No hidden prompts needed\! you can game ai peer review with presentation\-only revisions, 2026b\.[https://arxiv\.org/abs/2606\.13044](https://arxiv.org/abs/2606.13044)\.
- Ye et al\. \(2024\)Rui Ye, Xianghe Pang, Jingyi Chai, Jiaao Chen, Zhenfei Yin, Zhen Xiang, Xiaowen Dong, Jing Shao, and Siheng Chen\.Are we there yet? revealing the risks of utilizing large language models in scholarly peer review, 2024\.[https://arxiv\.org/abs/2412\.01708](https://arxiv.org/abs/2412.01708)\.
- Yuan et al\. \(2022\)Weizhe Yuan, Pengfei Liu, and Graham Neubig\.Can we automate scientific reviewing?*Journal of Artificial Intelligence Research*, 75:171–212, 2022\.
- Zheng et al\. \(2023\)Lianmin Zheng, Wei\-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al\.Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.*Advances in neural information processing systems*, 36:46595–46623, 2023\.
- Zhou et al\. \(2024\)Ruiyang Zhou, Lu Chen, and Kai Yu\.Is llm a reliable reviewer? a comprehensive evaluation of llm on automatic paper reviewing tasks\.In*Proceedings of the 2024 joint international conference on computational linguistics, language resources and evaluation \(LREC\-COLING 2024\)*, pages 9340–9351, 2024\.
- Zhu et al\. \(2025a\)Changjia Zhu, Junjie Xiong, Renkai Ma, Zhicong Lu, Yao Liu, and Lingyao Li\.When your reviewer is an llm: Biases, divergence, and prompt injection risks in peer review\.*arXiv preprint arXiv:2509\.09912*, 2025a\.
- Zhu et al\. \(2025b\)Minjun Zhu, Yixuan Weng, Linyi Yang, and Yue Zhang\.DeepReview: Improving LLM\-based paper review with human\-like deep thinking process\.In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors,*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 29330–29355, Vienna, Austria, July 2025b\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-251\-0\.[10\.18653/v1/2025\.acl\-long\.1420](https://doi.org/10.18653/v1/2025.acl-long.1420)\.[https://aclanthology\.org/2025\.acl\-long\.1420/](https://aclanthology.org/2025.acl-long.1420/)\.
## Appendix ATheoretical Motivation forSciCore
This section formalizes the representation and reconstruction arguments motivating the science\-core branch\. These results provide design rationales rather than performance guarantees for a fixed LLM reviewer\. Throughout,x∈𝒳x\\in\\mathcal\{X\}denotes a generic manuscript realization\. When we specialize the discussion toRobustReview,xi0x\_\{i0\}andxikx\_\{ik\}retain their definitions from Section[2](https://arxiv.org/html/2609.39027#S2), and we suppress the reviewer\-configuration index when a branch is fixed\.
### A\.1The Maximal\-Invariant Ideal
We writex∼scix′x\\sim\_\{\\mathrm\{sci\}\}x^\{\\prime\}when two realizations express the same reported scientific record, including claims, methods, assumptions, evidence, results, and conclusions, while differing rhetorically\. This relation defines the ideal invariance target\. Our generated rewrites are designed to approximate it, and their preservation is assessed empirically rather than assumed to be exact\.
An ideal representationC⋆:𝒳→𝒵C^\{\\star\}:\\mathcal\{X\}\\rightarrow\\mathcal\{Z\}is a maximal invariant to rhetorical realization when
C⋆\(x\)=C⋆\(x′\)⟺x∼scix′\.C^\{\\star\}\(x\)=C^\{\\star\}\(x^\{\\prime\}\)\\quad\\Longleftrightarrow\\quad x\\sim\_\{\\mathrm\{sci\}\}x^\{\\prime\}\.\(4\)The right\-to\-left implication expresses rhetorical invariance; the left\-to\-right implication prevents different scientific records from being collapsed into a single representation\. We use maximal invariant with respect to the equivalence relation∼sci\\sim\_\{\\mathrm\{sci\}\}\. This equivalence\-class view follows the standard maximal\-invariant principle and its use for invariant downstream prediction\([Eaton, 1989](https://arxiv.org/html/2609.39027#bib.bib13);[Dubois et al\., 2021](https://arxiv.org/html/2609.39027#bib.bib10)\)\.
###### Proposition A\.1\(Factorization through an ideal science core\)\.
IfC⋆C^\{\\star\}satisfies Equation[4](https://arxiv.org/html/2609.39027#A1.E4), every rhetorically invariant judgmentF:𝒳→𝒜F:\\mathcal\{X\}\\rightarrow\\mathcal\{A\}factors throughC⋆C^\{\\star\}: there exists a unique mapR:C⋆\(𝒳\)→𝒜R:C^\{\\star\}\(\\mathcal\{X\}\)\\rightarrow\\mathcal\{A\}such thatF=R∘C⋆F=R\\circ C^\{\\star\}\. Conversely, every judgment of the formR∘C⋆R\\circ C^\{\\star\}is rhetorically invariant\.
###### Proof\.
Forz=C⋆\(x\)z=C^\{\\star\}\(x\), defineR\(z\)=F\(x\)R\(z\)=F\(x\)\. IfC⋆\(x\)=C⋆\(x′\)C^\{\\star\}\(x\)=C^\{\\star\}\(x^\{\\prime\}\), maximality givesx∼scix′x\\sim\_\{\\mathrm\{sci\}\}x^\{\\prime\}, and invariance ofFFgivesF\(x\)=F\(x′\)F\(x\)=F\(x^\{\\prime\}\); thereforeRRis well defined\. The converse follows from the invariance ofC⋆C^\{\\star\}\. ∎
Proposition[A\.1](https://arxiv.org/html/2609.39027#A1.Thmproposition1)motivates the science\-core branch: an invariant judgment can, in principle, operate on a representation of the scientific equivalence class rather than on a particular manuscript realization\. Our implemented extractorC^\\widehat\{C\}is not claimed to be maximal or exactly invariant\. It instead approximates the two properties in Equation[4](https://arxiv.org/html/2609.39027#A1.E4): stability across matched variants designed to preserve the reported science and separation across different papers\. The completeSciCorereviewer does not replace manuscript\-based judgment with this representation\. It uses the resulting content\-normalized judgment as one branch of the final assessment\.
### A\.2Conditional Stability of the Science\-Core Branch and Fusion
Specializing the downstream judgment to a numerical overall\-assessment score, letr:𝒵→ℝr:\\mathcal\{Z\}\\rightarrow\\mathbb\{R\}denote the score map applied to the implemented science\-core representation, and letd𝒵d\_\{\\mathcal\{Z\}\}be a task\-relevant distance between extracted representations\. IfrrisLL\-Lipschitz on the observed representation range, then
\|r\(C^\(xik\)\)−r\(C^\(xi0\)\)\|≤Ld𝒵\(C^\(xik\),C^\(xi0\)\)\.\\left\|r\\\!\\left\(\\widehat\{C\}\(x\_\{ik\}\)\\right\)\-r\\\!\\left\(\\widehat\{C\}\(x\_\{i0\}\)\\right\)\\right\|\\leq L\\,d\_\{\\mathcal\{Z\}\}\\\!\\left\(\\widehat\{C\}\(x\_\{ik\}\),\\widehat\{C\}\(x\_\{i0\}\)\\right\)\.\(5\)Averaging over the controlled comparisons gives the corresponding deterministic\-score MAD bound
MADr∘C^≤LNK∑i=1N∑k=1Kd𝒵\(C^\(xik\),C^\(xi0\)\)\.\\mathrm\{MAD\}\_\{r\\circ\\widehat\{C\}\}\\leq\\frac\{L\}\{NK\}\\sum\_\{i=1\}^\{N\}\\sum\_\{k=1\}^\{K\}d\_\{\\mathcal\{Z\}\}\\\!\\left\(\\widehat\{C\}\(x\_\{ik\}\),\\widehat\{C\}\(x\_\{i0\}\)\\right\)\.\(6\)For stochastic reviewing, the same inequality applies to anLL\-Lipschitz conditional mean score map, while realized single\-run scores include additional decoding variation\. Equation[6](https://arxiv.org/html/2609.39027#A1.E6)should therefore be interpreted as a deterministic or conditional\-mean analogue of the empirical MAD in Appendix[C\.4](https://arxiv.org/html/2609.39027#A3.SS4), not as a bound on every realized review\. The relation is not an empirical guarantee because we do not establish exact rewrite preservation, identify the reviewer’s task\-relevant metric, or estimateLL\. It addresses only direct within\-paper stability; whether the reviewer retains distinctions among papers must still be assessed empirically through the joint stability\-discrimination metrics\.
LetM\(x\)M\(x\)denote the manuscript\-branch score and letB\(x\)=r\(C^\(x\)\)B\(x\)=r\(\\widehat\{C\}\(x\)\)denote the science\-core\-branch score\. Applied toRobustReview,M\(xik\)M\(x\_\{ik\}\)andB\(xik\)B\(x\_\{ik\}\)are the branch\-specific instances ofymiky\_\{mik\}in Section[2](https://arxiv.org/html/2609.39027#S2), with the fixed configuration index suppressed\. Under the fusion rule in Equation[2](https://arxiv.org/html/2609.39027#S4.E2), the triangle inequality and Equation[6](https://arxiv.org/html/2609.39027#A1.E6)give
MADSSciCore≤12MADM\+12MADB≤12MADM\+L2NK∑i=1N∑k=1Kd𝒵\(C^\(xik\),C^\(xi0\)\)\.\\mathrm\{MAD\}\_\{S\_\{\\mbox\{SciCore\}\}\}\\leq\\frac\{1\}\{2\}\\mathrm\{MAD\}\_\{M\}\+\\frac\{1\}\{2\}\\mathrm\{MAD\}\_\{B\}\\leq\\frac\{1\}\{2\}\\mathrm\{MAD\}\_\{M\}\+\\frac\{L\}\{2NK\}\\sum\_\{i=1\}^\{N\}\\sum\_\{k=1\}^\{K\}d\_\{\\mathcal\{Z\}\}\\\!\\left\(\\widehat\{C\}\(x\_\{ik\}\),\\widehat\{C\}\(x\_\{i0\}\)\\right\)\.\(7\)This decomposition provides a route to attenuating rhetorical variation: when the science\-core branch has lower rewrite\-induced MAD than the manuscript branch, their average is bounded by a value below the manuscript branch’s MAD\. Stable extracted representations can contribute to this condition when the downstream score map is sufficiently regular, while the manuscript branch retains the complete paper context\. The bound remains explanatory rather than a performance guarantee, and the fused reviewer must still be evaluated using the joint stability\-discrimination metrics\.
### A\.3Reconstruction as a Blackwell Garbling
The invariant\-representation view explains why a science core can support rhetorically stable judgment, but it does not by itself distinguish direct review from first realizing the record as another manuscript\. Blackwell’s comparison of statistical experiments supplies an information\-theoretic rationale for avoiding an unnecessary reconstruction layer, rather than a performance guarantee for a fixed reviewer\([Blackwell, 1953](https://arxiv.org/html/2609.39027#bib.bib6)\)\. LetZ=C^\(X\)Z=\\widehat\{C\}\(X\)be the extracted record, letUUcollect any auxiliary assets supplied unchanged to reconstruction, and writeW=\(Z,U\)W=\(Z,U\)for the complete reconstruction input\. If a reconstructed manuscript is sampled through a channelX~∼Q\(⋅∣W\)\\widetilde\{X\}\\sim Q\(\\,\\cdot\\mid W\), then for any review\-relevant stateΘ\\Theta,
Θ⟶W⟶X~\\Theta\\longrightarrow W\\longrightarrow\\widetilde\{X\}\(8\)forms a Markov chain\.
###### Proposition A\.2\(Reconstruction is a Blackwell garbling\)\.
Observing the reconstruction inputWWBlackwell\-dominates observing the reconstructed manuscriptX~\\widetilde\{X\}\. Consequently, across decision problems, the best attainable risk usingWWis no worse than the best attainable risk usingX~\\widetilde\{X\}\.
###### Proof\.
For any reconstruction\-based decision ruleδX~\(a∣x~\)\\delta\_\{\\widetilde\{X\}\}\(a\\mid\\widetilde\{x\}\), an observer ofWWcan first sampleX~\\widetilde\{X\}fromQQand then apply that rule:
δW\(a∣w\)=∫δX~\(a∣x~\)Q\(𝑑x~∣w\)\.\\delta\_\{W\}\(a\\mid w\)=\\int\\delta\_\{\\widetilde\{X\}\}\(a\\mid\\widetilde\{x\}\)Q\(d\\widetilde\{x\}\\mid w\)\.The two procedures induce the same conditional distribution of decisions givenΘ\\Theta, so every risk attainable fromX~\\widetilde\{X\}is attainable fromWW\. ∎
Thus reconstruction cannot add decision\-relevant information absent from its inputs, although it can discard information or introduce another rhetorical realization\. This proposition establishes only the information ordering between the complete reconstruction inputWWand its reconstructed outputX~\\widetilde\{X\}\. It does not imply that a fixed LLM will reviewWWmore effectively thanX~\\widetilde\{X\}\. Moreover, the implemented reconstruction arm receives permitted figures and bibliography in addition toZZ, whereas the science\-core branch reviews onlyZZ\. The proposition therefore motivates the ReconstructReview ablation but does not order the two implemented pipelines; their comparison is empirical\.
## Appendix BExtended Related Work
### B\.1Reliability and Robustness of LLM\-as\-a\-Judge Systems
LLM\-as\-a\-judge methods operationalize evaluation through natural\-language rubrics, pairwise decisions, and task\-specific evaluator models\. Systems such as G\-Eval, MT\-Bench, and Prometheus show that these designs can reproduce human preferences and provide criterion\-specific feedback\([Liu et al\., 2023](https://arxiv.org/html/2609.39027#bib.bib32);[Zheng et al\., 2023](https://arxiv.org/html/2609.39027#bib.bib52);[Kim et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib22)\)\. Their judgments are nevertheless affected by candidate order, response length, prompt wording, and evaluator identity\([Wang et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib45);[Dubois et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib11);[Stureborg et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib40);[Panickssery et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib38)\)\. Further variation arises from epistemic markers and semantically equivalent evaluation instructions\([Lee et al\., 2025](https://arxiv.org/html/2609.39027#bib.bib23);[Bhat and Varma, 2026](https://arxiv.org/html/2609.39027#bib.bib5)\), with meta\-evaluation revealing substantial dependence on the judge, task, property, and data source\([Bavaresco et al\., 2025](https://arxiv.org/html/2609.39027#bib.bib4)\)\.
Prior work has also begun to intervene on the evaluation process itself\. Length\-controlled evaluation reduces a known presentation bias\([Dubois et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib11)\), while representation\-based judging replaces direct generation with an intermediate semantic representation\([Li et al\., 2026d](https://arxiv.org/html/2609.39027#bib.bib29)\)\. These approaches suggest that robustness can depend on what information reaches the judge, not only on which rubric the judge receives\. Existing studies, however, largely address generic evaluation tasks or isolated biases\. We instead test scientific reviewers for both stability across rewrites designed to preserve reported scientific content and discrimination across papers, and test whether a science\-core branch improves this balance when combined with manuscript\-based review\.
Our theoretical view draws on two complementary traditions\. Maximal invariants represent equivalence classes while retaining the information needed by invariant downstream tasks\([Eaton, 1989](https://arxiv.org/html/2609.39027#bib.bib13);[Dubois et al\., 2021](https://arxiv.org/html/2609.39027#bib.bib10)\)\. Blackwell’s comparison of experiments orders observations by the decision risks they make attainable and identifies stochastic post\-processing as a garbling of its source\([Blackwell, 1953](https://arxiv.org/html/2609.39027#bib.bib6)\)\. We use these results as design rationales for the science\-core branch inSciCoreand for avoiding manuscript reconstruction within that branch, not as performance guarantees for a fixed LLM reviewer\.
### B\.2LLMs for Scientific Peer Review
Before modern LLMs, computational peer\-review research focused on data collection, decision prediction, and structured feedback generation\. PeerRead linked manuscripts with expert reports, aspect scores, and publication decisions for predictive and analytical tasks\([Kang et al\., 2018](https://arxiv.org/html/2609.39027#bib.bib21)\)\. ReviewRobot and ReviewAdvisor then moved toward evidence\-linked and aspect\-specific critique\([Wang et al\., 2020](https://arxiv.org/html/2609.39027#bib.bib46);[Yuan et al\., 2022](https://arxiv.org/html/2609.39027#bib.bib51)\)\. This progression expanded the modeled review workflow beyond acceptance prediction, while leaving reliable, decision\-relevant criticism as a central challenge\.
Modern LLMs extend this progression from modeling individual review components to generating complete natural\-language reviews at scale\. In a large\-scale study, GPT\-4 feedback overlaps with human review comments at rates comparable to the overlap between human reviews, and many authors report finding such feedback useful\([Liang et al\., 2023](https://arxiv.org/html/2609.39027#bib.bib30)\)\. More direct evaluations remain qualified: LLMs struggle with long\-paper processing, zero\-shot scoring, and consistently correct criticism\([Zhou et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib53)\); they reproduce summaries and stated strengths more readily than substantive weaknesses, discriminating questions, or differences in paper quality\([Li et al\., 2025a](https://arxiv.org/html/2609.39027#bib.bib27)\)\. Review quality also depends on prompting, retrieval, and reviewer\-guideline design\([Chen et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib7);[Li et al\., 2026a](https://arxiv.org/html/2609.39027#bib.bib24)\)\. At the same time, a randomized deployment shows that AI feedback can improve the specificity and actionability of human reviews\([Thakkar et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib41)\)\. Together, these results support AI as a review aid but leave open whether its scientific judgments are stable under content\-preserving rhetorical changes\.
Recent specialized models and agentic systems seek stronger reviewing through fine\-tuning, multi\-stage reasoning, retrieval, verification, reviewer ensembles, or proactive evidence gathering\([Idahl and Ahmadi, 2025](https://arxiv.org/html/2609.39027#bib.bib18);[Weng et al\., 2025](https://arxiv.org/html/2609.39027#bib.bib47);[Zhu et al\., 2025b](https://arxiv.org/html/2609.39027#bib.bib55);[Lu et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib33);[The OpenJudge Team, 2025](https://arxiv.org/html/2609.39027#bib.bib42);[Fang et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib14)\)\. Their evaluations primarily emphasize review quality, similarity to human feedback, score prediction, or workflow capability\. These goals are complementary to rhetorical robustness: a review may be detailed and plausible yet still assign different scientific scores to different presentations of the same work\. Such sensitivity makes merit judgments depend on rhetorical choices and undermines comparability across submissions\. We therefore evaluate general\-purpose LLM reviewers, specialized review models, and agentic review systems under the same controlled matched\-family design\.
### B\.3Manipulation and Presentation Sensitivity in AI Scientific Review
Scientific evaluation legitimately responds to clarity and exposition, but these criteria should not be conflated with methodological soundness or scientific contribution\([James et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib19);[Li et al\., 2025b](https://arxiv.org/html/2609.39027#bib.bib28)\)\. This distinction is difficult to study with naturally occurring papers because scientific merit and writing quality are entangled\. Controlled interventions provide a more direct test by varying presentation through revisions designed to preserve the reported scientific content\.
One line of work studies overtly adversarial manuscript interventions\. Hidden instructions can inflate ratings, suppress criticism, or redirect generated reviews\([Ye et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib50);[Collu et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib8)\), while character\-, word\-, and sentence\-level perturbations can distort review judgments\([Lin et al\., 2025](https://arxiv.org/html/2609.39027#bib.bib31)\)\. Such attacks establish serious vulnerabilities but differ from ordinary academic rewriting because they rely on concealed instructions or artificial perturbations\. Related evidence also shows that metadata and other manuscript\-side factors can introduce bias into automated review\([Vasu et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib43);[Zhu et al\., 2025a](https://arxiv.org/html/2609.39027#bib.bib54)\)\.
A more closely related line examines visible revisions intended to preserve reported scientific content\. Even title\-level stylistic variants can alter AI\-review scores\([Du, 2025](https://arxiv.org/html/2609.39027#bib.bib9)\)\. At the abstract and full\-paper levels, paraphrasing, rewriting, and overclaiming can be optimized to improve automated ratings\([Kaneko, 2026](https://arxiv.org/html/2609.39027#bib.bib20);[Li et al\., 2026b](https://arxiv.org/html/2609.39027#bib.bib25);[Wang et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib44)\), and recent work describes full\-manuscript*paper laundering*or*adversarial repackaging*that raises review scores without new experiments\([Baumann et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib3);[Yang et al\., 2026b](https://arxiv.org/html/2609.39027#bib.bib49)\)\. More general attacks similarly learn meaning\-preserving stylistic edits that exploit a particular LLM judge\([Yang et al\., 2026a](https://arxiv.org/html/2609.39027#bib.bib48)\)\.
[Li et al\. \(2026c\)](https://arxiv.org/html/2609.39027#bib.bib26)construct controlled full\-manuscript rewrites to characterize how rhetorical choices reward\-hack AI reviewers\. Using complete manuscripts, our investigation evaluates within\-paper stability, cross\-paper discrimination, and human alignment\. We introduceSciCore, which combines manuscript\-level assessment with a content\-normalized scientific judgment, and examine direct science\-core review against manuscript reconstruction\. Complementarily,[Dycke and Gurevych \(2026\)](https://arxiv.org/html/2609.39027#bib.bib12)changes substantive reasoning to test defect detection, whereas our counterfactuals are constructed to preserve reported reasoning and evidence when testing rhetorical invariance\. Sensitivity to substantive defects and invariance to content\-preserving rhetorical changes are complementary requirements for robust review\.
## Appendix CExperimental Details
This section provides additional details onRobustReviewconstruction and experiments,SciCore, and its ablation studies\. Our rhetorical intervention design, rewrite construction procedures, and manuscript\-review protocols follow the methodology of[Li et al\. \(2026c\)](https://arxiv.org/html/2609.39027#bib.bib26), with prompts instantiated for the present benchmark\. All manuscript rewrites and model\-review scores reported here were generated for this study\.
### C\.1RobustReviewComposition and Sampling
We constructRobustReviewfrom ICLR 2026 submissions using metadata downloaded through the OpenReview API\. We first identify submissions for which title matching yields a corresponding arXiv paper with an available source package\. To ensure that the matched arXiv manuscript closely corresponds to the ICLR submission, we require the text similarity between the ICLR submission and at least one current or historical arXiv version to exceed 0\.8\. Among submissions satisfying these criteria, we stratify papers by their mean human overall\-assessment rating\. Using a fixed random seed of 42, we randomly sample 10 papers from each of the six rating intervals\[1,3\)\[1,3\),\[3,4\)\[3,4\),\[4,5\)\[4,5\),\[5,6\)\[5,6\),\[6,7\)\[6,7\), and\[7,8\.5\]\[7,8\.5\], yielding 60 source papers\. Equal allocation across these intervals ensures coverage of lower\-, middle\-, and higher\-rated submissions\. We retain the mean human ratings as the reference for the human\-alignment evaluation\. The benchmark focuses on ICLR 2026 submissions with matched arXiv sources\. Generalization to other venues, fields, and manuscript formats remains to be evaluated\.
We construct 10 rhetorical rewrite conditions for each source paper\. Six conditions each apply a positive\-direction transformation to a single rhetorical dimension\.*Novelty stance*changes how strongly the novelty of the work is presented\.*Scope framing*changes how broadly or narrowly the scope and generalizability of the work are presented\.*Evidence framing*changes how the reported evidence and quantitative results are characterized\.*Contribution salience*changes the prominence and organization of the stated contributions\.*Technical register*changes the style and degree of technical and formal presentation\.*Linguistic complexity*changes lexical and syntactic realization\. The remaining 4 conditions introduce broader transformations\.*Joint rewrite*varies multiple rhetorical dimensions together\.*Recursive rewrite \(R2\)*and*recursive rewrite \(R3\)*correspond to 2 and 3 successive rounds of rewriting, respectively\.*Reviewer\-guided rewrite*uses reviewer feedback to guide the rhetorical transformation\.
Each condition contains 2 independently generated variants for each of the 60 source papers, yielding 120 variants per condition and 1,200 rhetorical variants across all 10 conditions\. Together with the 60 original manuscripts,RobustReviewcontains 1,260 full\-manuscript PDFs\. During evaluation, each variant is paired with the corresponding original manuscript from the same source\-paper family, yielding 120 matched original\-variant pairs per condition and 1,200 matched pairs overall\.
### C\.2Rewrite Construction and Structural Controls
Each rhetorical condition is instantiated independently by GPT\-5\.5 through Codex CLI and Claude Opus 4\.8\([Anthropic, 2026a](https://arxiv.org/html/2609.39027#bib.bib1)\)through Claude Code\. Rewriting is performed on the completeLaTeXproject associated with each source paper rather than on isolated sections or extracted text\. The transformation instructions require the models to preserve the reported scientific content, including claims, methods, evidence, numerical results, findings, and conclusions, while allowing the targeted changes in rhetorical framing, organization, wording, and presentation, including how quantitative evidence or tables are described and presented\.
Citations, cross\-references, figures, bibliography files, and other project dependencies are protected during rewriting\. Each transformed project is subsequently recompiled to verify structural validity\. Outputs that fail structural or compilation checks are repaired when possible and otherwise excluded from the benchmark\. These controls are designed to permit substantial changes in presentation while minimizing unintended changes to the reported scientific content\.
### C\.3Reviewer Execution Details
We evaluate eight general\-purpose LLMs prompted as scientific reviewers: GPT\-5\.5\([OpenAI, 2026](https://arxiv.org/html/2609.39027#bib.bib37)\), GPT\-5\-mini\([OpenAI, 2025](https://arxiv.org/html/2609.39027#bib.bib36)\), Claude Sonnet 5\([Anthropic, 2026b](https://arxiv.org/html/2609.39027#bib.bib2)\), GLM\-5\.2\([GLM\-5\-Team et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib16)\), Kimi\-K2\.6\([Moonshot AI, 2026](https://arxiv.org/html/2609.39027#bib.bib34)\), GPT\-OSS\-120B\([OpenAI, 2025](https://arxiv.org/html/2609.39027#bib.bib35)\), Gemini\-3\.5\-Flash\-Lite\([Google, 2026](https://arxiv.org/html/2609.39027#bib.bib17)\), and Qwen\-3\.5\-Flash\([Qwen Team, 2026](https://arxiv.org/html/2609.39027#bib.bib39)\)\. This reflects a common use case in which a general\-purpose LLM is directly instructed to review a scientific manuscript\. Each model is evaluated under three protocols\.Standardfollows the standard ICLR review criteria and scoring scheme\.StrictandPersistentare derived fromStandard:Strictapplies a more demanding, evidence\-focused rubric, andPersistentrepeatedly emphasizes that scientific judgments should depend on substantive scientific content rather than rhetorical presentation\. We additionally evaluate three specialized scientific\-review models: OpenReviewer\([Idahl and Ahmadi, 2025](https://arxiv.org/html/2609.39027#bib.bib18)\), CycleReviewer\([Weng et al\., 2025](https://arxiv.org/html/2609.39027#bib.bib47)\), and DeepReviewer\([Zhu et al\., 2025b](https://arxiv.org/html/2609.39027#bib.bib55)\)\. We also evaluate three agentic review systems: AI Scientist\([Lu et al\., 2024](https://arxiv.org/html/2609.39027#bib.bib33)\), OpenJudge\([The OpenJudge Team, 2025](https://arxiv.org/html/2609.39027#bib.bib42)\), and ProReviewer\([Fang et al\., 2026](https://arxiv.org/html/2609.39027#bib.bib14)\)\. The specialized and agentic systems use their native review procedures\.
All prompted\-review experiments operate directly on the manuscript PDFs\. GPT\-5\.5 and GPT\-5\-mini are accessed through the OpenAI Responses API,111[https://platform\.openai\.com/docs/api\-reference/responses](https://platform.openai.com/docs/api-reference/responses)using the default API settings without overriding sampling or generation parameters\. Claude Sonnet 5, GLM\-5\.2, Kimi\-K2\.6, GPT\-OSS\-120B, Gemini\-3\.5\-Flash\-Lite, and Qwen\-3\.5\-Flash are accessed through the OpenRouter API\.222[https://openrouter\.ai/docs/features/multimodal/pdfs](https://openrouter.ai/docs/features/multimodal/pdfs)For GLM\-5\.2, Kimi\-K2\.6, and GPT\-OSS\-120B, the reasoning effort is set tohigh\. All other generation and inference parameters are left at their OpenRouter defaults\. Each review request consists of the complete manuscript PDF together with the corresponding review prompt\. For models accessed through OpenRouter, PDF ingestion and processing are handled by OpenRouter’s file\-processing pipeline\.
The specialized reviewer systems are evaluated using their officially released checkpoints and inference procedures\. Specifically, we use Llama\-OpenReviewer\-8B for OpenReviewer, CycleReviewer\-ML\-Llama\-3\.1\-8B for CycleReviewer, and DeepReviewer\-14B for DeepReviewer\. CycleReviewer and DeepReviewer follow their released multi\-reviewer evaluation procedures, with the resulting reviewer scores aggregated according to their native implementations\. These models are served locally on a server equipped with 8 NVIDIA A100 GPUs\.
For specialized reviewers whose released interfaces do not accept PDF input, we first convert each manuscript to Markdown with the Datalab OCR API333[https://documentation\.datalab\.to/](https://documentation.datalab.to/)and pass the resulting Markdown to the model’s native inference pipeline\.
The agentic review systems, AI Scientist, OpenJudge, and ProReviewer, are evaluated using their officially released implementations and configurations\. We retain the model backbones, inference parameters, review workflows, and score extraction procedures specified by their respective implementations without additional modification\.
Single\-review execution\.We use one independent review per manuscript version in the main experiments, as repeated reviewing would substantially increase the computational cost ofRobustReview\. We conduct an auxiliary audit with GPT\-5\-mini under the Standard protocol, covering all 60 source papers, both rewrite producers, and the six single\-dimension rewrite conditions, for a total of 720 rewritten manuscripts\. For each manuscript, we compare a single review with the mean of three independent reviews of the same PDF\. Averaging three reviews changes the mean OA from 6\.493 to 6\.457\. The 12 condition–producer mean effect estimates are closely aligned between the two settings, with Pearson and Spearman correlations of 0\.975 and 0\.944, respectively\. The mean absolute change in these effects is 0\.056 OA points, with a maximum change of 0\.117\. This audit supports the consistency of condition\-level mean effects under the evaluated configuration\. Our robustness metrics characterize observed score variation under single\-review execution, including stochastic generation variability\.
Appendix[E\.3](https://arxiv.org/html/2609.39027#A5.SS3)reports the same\-PDF repeatability baseline for the manuscript protocols andSciCore\.
### C\.4Full Evaluation Metric Definitions
All configurations are evaluated on the same 60 complete paper families \(1,260 manuscript versions\), with no missing overall\-assessment scores\.
For reviewer configurationmmand a given score dimension, we define the rewrite\-induced score change as
Δmik=ymik−ymi0,k=1,…,K\.\\Delta\_\{mik\}=y\_\{mik\}\-y\_\{mi0\},\\qquad k=1,\\ldots,K\.\(9\)
Direct within\-paper stability\.MAD measures the average magnitude of score changes induced by rhetorical rewriting:
MADm=1NK∑i=1N∑k=1K\|Δmik\|\.\\mathrm\{MAD\}\_\{m\}=\\frac\{1\}\{NK\}\\sum\_\{i=1\}^\{N\}\\sum\_\{k=1\}^\{K\}\\left\|\\Delta\_\{mik\}\\right\|\.\(10\)Drift SD measures the variability of the signed score changes:
DriftSDm=SDi,k\(Δmik\)\.\\mathrm\{DriftSD\}\_\{m\}=\\operatorname\{SD\}\_\{i,k\}\\left\(\\Delta\_\{mik\}\\right\)\.\(11\)Lower values indicate greater rhetorical stability\. MAD captures the overall magnitude of score drift, whereas Drift SD captures its heterogeneity across papers and rewrite conditions\.
Joint stability\-discrimination\.Direct within\-paper stability is insufficient when low drift is produced by score collapse\. We therefore use three complementary metrics that relate within\-paper consistency to cross\-paper variation or separation\. These are joint metrics rather than between\-paper\-only measures\. ICC measures whether different presentations of the same paper receive similar scores relative to differences across papers\. We use the two\-way mixed\-effects, absolute\-agreement, single\-measure form, ICC\(A,1\):
ICCm\(A,1\)=MSR−MSEMSR\+\(P−1\)MSE\+PN\(MSC−MSE\),\\mathrm\{ICC\}\_\{m\}\(A,1\)=\\frac\{MS\_\{R\}\-MS\_\{E\}\}\{MS\_\{R\}\+\(P\-1\)MS\_\{E\}\+\\frac\{P\}\{N\}\(MS\_\{C\}\-MS\_\{E\}\)\},\(12\)whereMSRMS\_\{R\},MSCMS\_\{C\}, andMSEMS\_\{E\}denote the paper, presentation, and residual mean squares, respectively;NNis the number of papers; andP=K\+1P=K\+1is the number of presentations including the original manuscript\. Higher ICC indicates that between\-paper differences are large relative to presentation\- induced and residual variation\.
We further define the signal preservation ratio \(SPR\) as
SPRm=Vari\(ymi0\)Vari\(ymi0\)\+𝔼i,k\[Δmik2\]\.\\mathrm\{SPR\}\_\{m\}=\\frac\{\\operatorname\{Var\}\_\{i\}\\left\(y\_\{mi0\}\\right\)\}\{\\operatorname\{Var\}\_\{i\}\\left\(y\_\{mi0\}\\right\)\+\\mathbb\{E\}\_\{i,k\}\\left\[\\Delta\_\{mik\}^\{2\}\\right\]\}\.\(13\)SPR compares the cross\-paper score variation in the original manuscripts with the magnitude of rewrite\-induced movement\. Higher values indicate that cross\-paper variation remains large relative to within\-paper rhetorical perturbation\. When both the original\-paper score variance and the rewrite\-induced mean squared drift are zero, we define SPR as zero, reflecting the absence of between\-paper discrimination\.
Finally, discriminability measures whether presentations of the same paper are closer in score than presentations of different papers:
Discm=𝔼i≠j,a≠ba,b,c∈\{0,…,K\}\[\\displaystyle\\mathrm\{Disc\}\_\{m\}=\\mathbb\{E\}\_\{\\begin\{subarray\}\{c\}i\\neq j,\\;a\\neq b\\\\ a,b,c\\in\\\{0,\\ldots,K\\\}\\end\{subarray\}\}\\Big\[𝟏\(\|ymia−ymib\|<\|ymia−ymjc\|\)\\displaystyle\\mathbf\{1\}\\left\(\|y\_\{mia\}\-y\_\{mib\}\|<\|y\_\{mia\}\-y\_\{mjc\}\|\\right\)\(14\)\+12𝟏\(\|ymia−ymib\|=\|ymia−ymjc\|\)\]\.\\displaystyle\+\\frac\{1\}\{2\}\\mathbf\{1\}\\left\(\|y\_\{mia\}\-y\_\{mib\}\|=\|y\_\{mia\}\-y\_\{mjc\}\|\\right\)\\Big\]\.Here,aaandbbindex two presentations of paperii, whileccindexes a presentation of a different paperjj\. A value of0\.50\.5is chance\-like, while higher values indicate stronger separation between same\-paper and different\- paper presentations\.
ICC, SPR, and discriminability operationalize the same joint requirement in different ways: each rewards consistency across presentations of the same paper only when distinctions across papers remain detectable\. None treats cross\-paper score variation as ground\-truth scientific merit\.
Human judgment alignment\.Lethih\_\{i\}denote the mean human OA score for the original version of paperii\. Human MAE measures absolute agreement between AI and human scores:
HMAEm=1N∑i=1N\|ymi0−hi\|\.\\mathrm\{HMAE\}\_\{m\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\left\|y\_\{mi0\}\-h\_\{i\}\\right\|\.\(15\)We additionally measure rank agreement using Spearman rank correlation:
ρmhuman=ρS\(\(ymi0\)i=1N,\(hi\)i=1N\)\.\\rho^\{\\mathrm\{human\}\}\_\{m\}=\\rho\_\{\\mathrm\{S\}\}\\left\(\\left\(y\_\{mi0\}\\right\)\_\{i=1\}^\{N\},\\left\(h\_\{i\}\\right\)\_\{i=1\}^\{N\}\\right\)\.\(16\)Lower Human MAE and higher Spearman rank correlation indicate stronger agreement with human reviewer judgments\. These metrics assess agreement with mean human ratings and do not characterize disagreement among individual reviewers\. Spearman correlation is undefined when either input is constant; ICC is undefined when its denominator is zero\. We report these cases as dashes\.
### C\.5SciCoreExecution Protocol
The Core\-Adapted prompt, the Manuscript\-Strict branch, and the equal\-weight fusion rule were fixed before examining results on the 60\-paper evaluation panel\.
TheSciCoreexperiments follow the same model\-access and inference settings as the corresponding direct\-review experiments described above\. The science\-core branch first uses GPT\-5\.5 to extract the science core from the manuscript and then reviews the extracted record with the adapted protocol\. The manuscript branch is the GPT\-5\.5Strictreview of the complete PDF reported in the main benchmark\. The final overall assessment is computed as the unweighted arithmetic mean of the manuscript\-branch and science\-core\-branch scores\. Both branches use the same ICLR overall\-assessment scale, and their arithmetic fusion assumes that scores are comparable across branches\. No additional model call is required for fusion\. In the branch\-policy ablation study, each extracted science core is reused across the evaluated policies; the reviewer backbone and input representation remain the same while the review instructions vary\.
### C\.6ReconstructReview Ablation
ReconstructReview evaluates whether the extracted record should be reviewed directly within the science\-core branch or first realized as another manuscript\. For each manuscript presentation, GPT\-5\.5 first extracts the science core using the same extraction procedure asSciCore\. The extracted content is then provided to GPT\-5\.5 through Codex, which reconstructs it into a complete anonymous ICLR 2026 submission using the official conference template\. The reconstruction workspace contains the extracted science core as the authoritative scientific record, together with the permitted bibliography, available scientific figures, and the ICLR 2026 template\. The original manuscript prose is not provided to the reconstruction model\.
The reconstruction model is given substantial authorial freedom over the title, framing, organization, mathematical presentation, allocation between the main text and appendix, and the selection and placement of figures and tables\. At the same time, it is explicitly instructed to preserve the scientific record, retain nonredundant methods and experimental evidence, preserve reported numerical results and qualifications, and avoid introducing unsupported experiments, claims, citations, or implementation details\. The resultingLaTeXproject is compiled into a full\-manuscript PDF and reviewed using the Standard protocol\.
We apply this reconstruction procedure independently to every presentation inRobustReview, including the original baseline and the variants produced by both rewrite models under all 10 rhetorical conditions\. This yields 21 reconstructed manuscripts per source paper and 1,260 reconstructed manuscripts in total\. The resulting scores are then analyzed using the same matched original\-variant evaluation procedure as inRobustReview\.
## Appendix DCompleteRobustReviewResults
This section reports the complete baseline evidence supporting the compact comparisons in the main paper\. Baseline reviewer models and protocols are kept separate from the backbone\-transfer experiments forSciCore\.
### D\.1Full Score\-Dimension Results
Complete OA results are already reported in Table[1](https://arxiv.org/html/2609.39027#S3.T1)and are not repeated here\. The tables below use the same seven metrics and column order as the main table, but restrict the evaluated target to soundness, presentation, contribution, or confidence\. Human alignment is computed against the mean score from the official reviews of eachRobustReviewsource paper\. Presentation scores were recovered directly from the OpenReview API; the other three secondary\-score means exactly match the archived metadata\. MAD and Drift SD directly measure within\-paper stability, whereas ICC, SPR, and discriminability are joint stability\-discrimination metrics\. The GPT\-5\-mini entries are recomputed from their complete paper\-by\-rewrite score matrices rather than from the earlier aggregate\-only export\. External reviewers and the finalSciCorefusion expose only OA\-compatible outputs and are therefore not included in these secondary\-score tables\.
Table 6:Full Soundness results for secondary scores\.Metrics and column order match[Table1](https://arxiv.org/html/2609.39027#S3.T1)\. The best result in each column is shown in bold, the second\-best is underlined, and the third\-best is italicized\.Table 7:Full Presentation results for secondary scores\.Metrics and column order match[Table1](https://arxiv.org/html/2609.39027#S3.T1)\. The best result in each column is shown in bold, the second\-best is underlined, and the third\-best is italicized\. A dash denotes an undefined Spearman correlation\.Table 8:Full Contribution results for secondary scores\.Metrics and column order match[Table1](https://arxiv.org/html/2609.39027#S3.T1)\. The best result in each column is shown in bold, the second\-best is underlined, and the third\-best is italicized\.Table 9:Full Confidence results for secondary scores\.Metrics and column order match[Table1](https://arxiv.org/html/2609.39027#S3.T1)\. The best result in each column is shown in bold, the second\-best is underlined, and the third\-best is italicized\. A dash denotes an undefined Spearman correlation or ICC\.
## Appendix EAdditional Validation
### E\.1Core Technical\-Content Fidelity Audit
We use DeepSeek V4 Flash as an independent auditor to assess preservation of core technical content between each rhetorical rewrite and its matched original\. Rhetorical emphasis and evaluative framing are allowed to vary as part of the intended interventions\. The auditor is not used as a rewrite producer, science\-core extractor, or reviewer elsewhere in our experiments\. The audit examines five dimensions:*table fidelity*, covering table structure, labels, values, and method–dataset–metric associations;*numerical fidelity*, covering numerical values, statistics, confidence intervals, and p\-values in the manuscript body;*experimental fidelity*, covering datasets, evaluated subsets, splits, sample sizes, metrics, baselines, and experimental settings;*method fidelity*, covering method components, algorithmic operations, equations, variables, assumptions, and technical dependencies; and*result fidelity*, covering the direction and ordering of empirical results and positive versus negative observed effects\.
The audit uses a binary decision rule\. For paperii, rhetorical variantkk, and applicable fidelity dimensiondd, letzik\(d\)=1z\_\{ik\}^\{\(d\)\}=1when no concrete technical\-content mismatch is identified andzik\(d\)=0z\_\{ik\}^\{\(d\)\}=0when the auditor identifies a mismatch supported by evidence from the matched manuscripts\. The fidelity rate for dimensionddis
Fidelityd=1Nd∑\(i,k\)∈𝒜dzik\(d\),Nd=\|𝒜d\|,\\operatorname\{Fidelity\}\_\{d\}=\\frac\{1\}\{N\_\{d\}\}\\sum\_\{\(i,k\)\\in\\mathcal\{A\}\_\{d\}\}z\_\{ik\}^\{\(d\)\},\\qquad N\_\{d\}=\\lvert\\mathcal\{A\}\_\{d\}\\rvert,where𝒜d\\mathcal\{A\}\_\{d\}contains the comparisons to which dimensionddapplies\.
To complement the automated audit, we conduct a human validation on 150 comparisons between originals and variants\. We first randomly sample 120 of the 1,200 rhetorical variants \(10%\), each evaluated against its matched original, and supplement them with 30 additional comparisons for which the automated auditor identifies at least one fidelity mismatch\. Three graduate\-level annotators independently assess all 150 comparisons using the same five dimensions: table, numerical, experimental, method, and result fidelity\. For each comparison and dimension, the human verdict is determined by majority vote among the three annotators\. Dimension\-level human pass rates summarize these consensus verdicts across the combined 150\-comparison audit sample\. Overall inter\-annotator agreement, measured as the mean pairwise agreement across all 750 dimension\-level judgments, is 94\.6%\.
Table 10:Core technical\-content fidelity of rhetorical rewrites\.Dimension\-level rates report the proportion of applicable comparisons that pass each check\. Human rates summarize the combined audit sample of 120 randomly selected comparisons and 30 additional comparisons flagged by the automated auditor\.
Across the five assessed dimensions, the automated and human checks indicate that core technical content is largely preserved, with occasional discrepancies\. The audit serves as a diagnostic assessment, and all comparisons are retained in the reported evaluation\. Residual content differences may contribute to observed score variation\.
### E\.2Complementary Review Branches under Equal\-Weight Fusion
Table[11](https://arxiv.org/html/2609.39027#A5.T11)comparesSciCorewith full\-manuscript score averages\. We reuse cached scores, average them without rounding, and recompute all seven metrics\. The Strict\-anchored comparisons reuse the same manuscript\-branch scores\.
Some apparent robustness gains may arise from averaging itself: averaging integer scores allows half points, can reduce MAD and Drift SD through drift cancellation, and changes distance ties in discriminability\. To assess whetherSciCoreretains an advantage under the same averaging operation, the two\-score manuscript controls match its averaging rule and score resolution\. The three\-protocol average additionally matches its nominal three\-call budget: three reviews versus one extraction and two reviews, with potentially different token costs\.
Table 11:Full\-manuscript averaging controls onRobustReview\.All configurations use GPT\-5\.5 and the same 60 papers with 21 versions each\. Protocol names denote full\-manuscript reviews\. Scores are averaged separately for each original and rewrite before computing the metrics, with no rounding\.SciCoreaverages Manuscript\-Strict and Core\-Adapted\. The three\-protocol control averages three review scores\. All entries are point estimates\.ConfigurationHuman alignmentWithin\-paper stabilityJoint stability\-discriminationH\-MAE↓\\downarrowSpearman↑\\uparrowMAD↓\\downarrowDrift SD↓\\downarrowICC↑\\uparrowSPR↑\\uparrowDiscrim\.↑\\uparrowSingle full\-manuscript reviewsManuscript\-Standard1\.2940\.4480\.5980\.9560\.6150\.5550\.649Manuscript\-Strict1\.0780\.5290\.7661\.2020\.6320\.5070\.672Manuscript\-Persistent1\.2610\.4230\.4760\.8750\.6970\.5860\.682Two\-score averagesStandard \+ Strict1\.1300\.4950\.5600\.8200\.7500\.6200\.705Standard \+ Persistent1\.1800\.4550\.4700\.6810\.7550\.6400\.712Strict \+ Persistent1\.0660\.4940\.5500\.8300\.7580\.6350\.715SciCore1\.0720\.4880\.4760\.6870\.7750\.6520\.726Three\-score averageStandard \+ Strict \+ Persistent1\.1100\.4660\.5200\.6810\.7810\.6350\.715
With Manuscript\-Strict and equal weighting fixed, Core\-Adapted improves all five robustness metrics over either manuscript alternative, supported by paired bootstrap intervals \(Appendix[E\.4](https://arxiv.org/html/2609.39027#A5.SS4)\)\. Thus, the robustness advantage over these two controls persists when the manuscript branch, averaging rule, and score resolution are matched\. Both alternatives have higher Spearman correlation, and Strict \+ Persistent also has lower H\-MAE\. Standard \+ Persistent outperformsSciCoreon MAD and Drift SD, while the three\-protocol average outperforms it on Drift SD and ICC;SciCorehas better point estimates on the other five metrics in each comparison and remains Pareto non\-dominated\.
### E\.3Repeated\-Review Noise Baseline
Following the repeated\-review procedure in Appendix[C\.3](https://arxiv.org/html/2609.39027#A3.SS3), we obtain three independent reviews of each identical PDF, holding the review prompts and inference settings fixed\. We apply this procedure to the three manuscript protocols andSciCorewith GPT\-5\.5, following each configuration’s evaluation pipeline\. MAD and ICC quantify variation across repeated reviews of the same manuscript using the definitions in Appendix[C\.4](https://arxiv.org/html/2609.39027#A3.SS4)\.
Repeated reviews of identical PDFs yield MAD of 0\.294–0\.329 and ICC of 0\.846–0\.862 across the three manuscript protocols andSciCore\.SciCoreand Manuscript\-Strict have similar repeatability \(MAD: 0\.320 versus 0\.329; ICC: 0\.858 versus 0\.851\), but differ more under rewriting \(MAD: 0\.476 versus 0\.766; ICC: 0\.775 versus 0\.632\), supporting a robustness gain beyond the small difference in same\-PDF repeatability\.
### E\.4Paper\-Family Bootstrap Analysis
Tables[12](https://arxiv.org/html/2609.39027#A5.T12)and[13](https://arxiv.org/html/2609.39027#A5.T13)use 5,000 paired bootstrap resamples of the 60 paper families, keeping each original and its 20 variants together\. The same resamples are used across configurations to recompute metrics and paired differences\. Marginal 95% intervals use the 2\.5th and 97\.5th percentiles, with the recorded reviews held fixed\.
The paired intervals in Table[12](https://arxiv.org/html/2609.39027#A5.T12)supportSciCore’s gains on all five robustness metrics over Manuscript\-Strict and both two\-score manuscript ensembles, with tradeoffs in human alignment\. Relative to Core\-Adapted, fusion improves human alignment, ICC, and discriminability, while increasing MAD and Drift SD and reducing SPR\. Table[13](https://arxiv.org/html/2609.39027#A5.T13)gives metric intervals for the primary comparison\.
Table 12:Paired bootstrap comparisons forSciCore\.All configurations use GPT\-5\.5 and the same 60 complete paper families\. Each cell shows the difference \(SciCoreminus the column configuration\), followed by its marginal 95% percentile interval from 5,000 paired paper\-family resamples\. Negative differences favorSciCorefor H\-MAE, MAD, and Drift SD; positive differences favorSciCorefor the other metrics\.Table 13:Paper\-family bootstrap intervals for the primary comparison\.Entries are marginal 95% percentile intervals from 5,000 paired resamples of the 60 complete paper families; the corresponding point estimates appear in Table[1](https://arxiv.org/html/2609.39027#S3.T1)\. All original and rewritten versions stay together within each resample\.SystemProtocolHuman alignmentWithin\-paper stabilityJoint stability\-discriminationH\-MAE↓\\downarrowSpearman↑\\uparrowMAD↓\\downarrowDrift SD↓\\downarrowICC↑\\uparrowSPR↑\\uparrowDiscrim\.↑\\uparrowGPT\-5\.5Standard\[1\.067,1\.531\]\[1\.067,\\,1\.531\]\[0\.213,0\.640\]\[0\.213,\\,0\.640\]\[0\.476,0\.717\]\[0\.476,\\,0\.717\]\[0\.809,1\.075\]\[0\.809,\\,1\.075\]\[0\.492,0\.698\]\[0\.492,\\,0\.698\]\[0\.441,0\.630\]\[0\.441,\\,0\.630\]\[0\.614,0\.679\]\[0\.614,\\,0\.679\]Strict\[0\.884,1\.287\]\[0\.884,\\,1\.287\]\[0\.305,0\.700\]\[0\.305,\\,0\.700\]\[0\.603,0\.927\]\[0\.603,\\,0\.927\]\[1\.038,1\.338\]\[1\.038,\\,1\.338\]\[0\.520,0\.709\]\[0\.520,\\,0\.709\]\[0\.416,0\.593\]\[0\.416,\\,0\.593\]\[0\.632,0\.702\]\[0\.632,\\,0\.702\]Persistent\[1\.037,1\.491\]\[1\.037,\\,1\.491\]\[0\.161,0\.638\]\[0\.161,\\,0\.638\]\[0\.364,0\.593\]\[0\.364,\\,0\.593\]\[0\.715,1\.017\]\[0\.715,\\,1\.017\]\[0\.573,0\.770\]\[0\.573,\\,0\.770\]\[0\.430,0\.696\]\[0\.430,\\,0\.696\]\[0\.640,0\.716\]\[0\.640,\\,0\.716\]GPT\-5\-miniStandard\[1\.387,2\.048\]\[1\.387,\\,2\.048\]\[0\.148,0\.634\]\[0\.148,\\,0\.634\]\[0\.431,0\.664\]\[0\.431,\\,0\.664\]\[0\.821,1\.057\]\[0\.821,\\,1\.057\]\[0\.276,0\.504\]\[0\.276,\\,0\.504\]\[0\.263,0\.487\]\[0\.263,\\,0\.487\]\[0\.553,0\.616\]\[0\.553,\\,0\.616\]Strict\[1\.014,1\.471\]\[1\.014,\\,1\.471\]\[0\.142,0\.616\]\[0\.142,\\,0\.616\]\[0\.749,1\.046\]\[0\.749,\\,1\.046\]\[1\.111,1\.306\]\[1\.111,\\,1\.306\]\[0\.283,0\.481\]\[0\.283,\\,0\.481\]\[0\.352,0\.487\]\[0\.352,\\,0\.487\]\[0\.567,0\.625\]\[0\.567,\\,0\.625\]Persistent\[1\.344,1\.919\]\[1\.344,\\,1\.919\]\[0\.098,0\.542\]\[0\.098,\\,0\.542\]\[0\.426,0\.749\]\[0\.426,\\,0\.749\]\[0\.792,1\.152\]\[0\.792,\\,1\.152\]\[0\.326,0\.581\]\[0\.326,\\,0\.581\]\[0\.357,0\.555\]\[0\.357,\\,0\.555\]\[0\.551,0\.623\]\[0\.551,\\,0\.623\]Claude Sonnet 5Standard\[0\.962,1\.432\]\[0\.962,\\,1\.432\]\[0\.137,0\.649\]\[0\.137,\\,0\.649\]\[0\.364,0\.577\]\[0\.364,\\,0\.577\]\[0\.730,0\.972\]\[0\.730,\\,0\.972\]\[0\.467,0\.671\]\[0\.467,\\,0\.671\]\[0\.439,0\.636\]\[0\.439,\\,0\.636\]\[0\.607,0\.678\]\[0\.607,\\,0\.678\]Strict\[0\.944,1\.372\]\[0\.944,\\,1\.372\]\[0\.214,0\.679\]\[0\.214,\\,0\.679\]\[0\.491,0\.740\]\[0\.491,\\,0\.740\]\[0\.919,1\.157\]\[0\.919,\\,1\.157\]\[0\.422,0\.641\]\[0\.422,\\,0\.641\]\[0\.457,0\.627\]\[0\.457,\\,0\.627\]\[0\.601,0\.668\]\[0\.601,\\,0\.668\]Persistent\[1\.027,1\.473\]\[1\.027,\\,1\.473\]\[0\.053,0\.602\]\[0\.053,\\,0\.602\]\[0\.218,0\.414\]\[0\.218,\\,0\.414\]\[0\.571,0\.834\]\[0\.571,\\,0\.834\]\[0\.378,0\.568\]\[0\.378,\\,0\.568\]\[0\.350,0\.575\]\[0\.350,\\,0\.575\]\[0\.573,0\.648\]\[0\.573,\\,0\.648\]GLM\-5\.2Standard\[1\.706,2\.394\]\[1\.706,\\,2\.394\]\[−0\.296,0\.241\]\[\-0\.296,\\,0\.241\]\[1\.264,1\.744\]\[1\.264,\\,1\.744\]\[2\.071,2\.572\]\[2\.071,\\,2\.572\]\[0\.143,0\.305\]\[0\.143,\\,0\.305\]\[0\.196,0\.411\]\[0\.196,\\,0\.411\]\[0\.539,0\.576\]\[0\.539,\\,0\.576\]Strict\[1\.687,2\.393\]\[1\.687,\\,2\.393\]\[−0\.294,0\.222\]\[\-0\.294,\\,0\.222\]\[1\.391,1\.709\]\[1\.391,\\,1\.709\]\[1\.898,2\.250\]\[1\.898,\\,2\.250\]\[0\.139,0\.283\]\[0\.139,\\,0\.283\]\[0\.337,0\.429\]\[0\.337,\\,0\.429\]\[0\.538,0\.572\]\[0\.538,\\,0\.572\]Persistent\[2\.002,2\.822\]\[2\.002,\\,2\.822\]\[−0\.451,0\.094\]\[\-0\.451,\\,0\.094\]\[1\.663,2\.218\]\[1\.663,\\,2\.218\]\[2\.407,2\.902\]\[2\.407,\\,2\.902\]\[0\.147,0\.308\]\[0\.147,\\,0\.308\]\[0\.328,0\.432\]\[0\.328,\\,0\.432\]\[0\.536,0\.579\]\[0\.536,\\,0\.579\]Kimi\-K2\.6Standard\[1\.374,2\.025\]\[1\.374,\\,2\.025\]\[0\.027,0\.504\]\[0\.027,\\,0\.504\]\[1\.128,1\.595\]\[1\.128,\\,1\.595\]\[1\.698,2\.237\]\[1\.698,\\,2\.237\]\[0\.131,0\.291\]\[0\.131,\\,0\.291\]\[0\.357,0\.474\]\[0\.357,\\,0\.474\]\[0\.535,0\.577\]\[0\.535,\\,0\.577\]Strict\[1\.419,2\.072\]\[1\.419,\\,2\.072\]\[−0\.200,0\.300\]\[\-0\.200,\\,0\.300\]\[0\.610,0\.912\]\[0\.610,\\,0\.912\]\[1\.093,1\.424\]\[1\.093,\\,1\.424\]\[0\.135,0\.431\]\[0\.135,\\,0\.431\]\[0\.280,0\.447\]\[0\.280,\\,0\.447\]\[0\.522,0\.569\]\[0\.522,\\,0\.569\]Persistent\[1\.261,1\.923\]\[1\.261,\\,1\.923\]\[−0\.233,0\.329\]\[\-0\.233,\\,0\.329\]\[0\.987,1\.319\]\[0\.987,\\,1\.319\]\[1\.453,1\.824\]\[1\.453,\\,1\.824\]\[0\.161,0\.410\]\[0\.161,\\,0\.410\]\[0\.347,0\.426\]\[0\.347,\\,0\.426\]\[0\.543,0\.596\]\[0\.543,\\,0\.596\]GPT\-OSS\-120BStandard\[1\.298,1\.798\]\[1\.298,\\,1\.798\]\[−0\.195,0\.336\]\[\-0\.195,\\,0\.336\]\[0\.586,0\.863\]\[0\.586,\\,0\.863\]\[0\.940,1\.297\]\[0\.940,\\,1\.297\]\[0\.089,0\.186\]\[0\.089,\\,0\.186\]\[0\.277,0\.443\]\[0\.277,\\,0\.443\]\[0\.521,0\.541\]\[0\.521,\\,0\.541\]Strict\[1\.379,2\.024\]\[1\.379,\\,2\.024\]\[−0\.163,0\.375\]\[\-0\.163,\\,0\.375\]\[0\.276,0\.511\]\[0\.276,\\,0\.511\]\[0\.733,1\.013\]\[0\.733,\\,1\.013\]\[0\.049,0\.132\]\[0\.049,\\,0\.132\]\[0\.176,0\.428\]\[0\.176,\\,0\.428\]\[0\.504,0\.513\]\[0\.504,\\,0\.513\]Persistent\[1\.289,1\.774\]\[1\.289,\\,1\.774\]\[−0\.250,0\.237\]\[\-0\.250,\\,0\.237\]\[0\.707,1\.043\]\[0\.707,\\,1\.043\]\[1\.096,1\.545\]\[1\.096,\\,1\.545\]\[0\.112,0\.207\]\[0\.112,\\,0\.207\]\[0\.266,0\.392\]\[0\.266,\\,0\.392\]\[0\.526,0\.544\]\[0\.526,\\,0\.544\]Gemini 3\.5 Flash\-LiteStandard\[2\.828,3\.532\]\[2\.828,\\,3\.532\]\[0\.208,0\.556\]\[0\.208,\\,0\.556\]\[0\.105,0\.323\]\[0\.105,\\,0\.323\]\[0\.450,0\.773\]\[0\.450,\\,0\.773\]\[0\.074,0\.311\]\[0\.074,\\,0\.311\]\[0\.340,0\.521\]\[0\.340,\\,0\.521\]\[0\.503,0\.523\]\[0\.503,\\,0\.523\]Strict\[1\.553,2\.209\]\[1\.553,\\,2\.209\]\[0\.242,0\.654\]\[0\.242,\\,0\.654\]\[0\.691,0\.998\]\[0\.691,\\,0\.998\]\[1\.071,1\.296\]\[1\.071,\\,1\.296\]\[0\.316,0\.512\]\[0\.316,\\,0\.512\]\[0\.431,0\.587\]\[0\.431,\\,0\.587\]\[0\.573,0\.622\]\[0\.573,\\,0\.622\]Persistent\[2\.491,3\.176\]\[2\.491,\\,3\.176\]\[0\.271,0\.649\]\[0\.271,\\,0\.649\]\[0\.268,0\.651\]\[0\.268,\\,0\.651\]\[0\.715,1\.218\]\[0\.715,\\,1\.218\]\[0\.217,0\.449\]\[0\.217,\\,0\.449\]\[0\.446,0\.540\]\[0\.446,\\,0\.540\]\[0\.518,0\.561\]\[0\.518,\\,0\.561\]Qwen 3\.5 FlashStandard\[1\.988,2\.722\]\[1\.988,\\,2\.722\]\[0\.129,0\.595\]\[0\.129,\\,0\.595\]\[0\.645,1\.019\]\[0\.645,\\,1\.019\]\[1\.105,1\.628\]\[1\.105,\\,1\.628\]\[0\.187,0\.388\]\[0\.187,\\,0\.388\]\[0\.396,0\.516\]\[0\.396,\\,0\.516\]\[0\.538,0\.584\]\[0\.538,\\,0\.584\]Strict\[1\.046,1\.433\]\[1\.046,\\,1\.433\]\[0\.180,0\.586\]\[0\.180,\\,0\.586\]\[0\.739,1\.009\]\[0\.739,\\,1\.009\]\[1\.177,1\.487\]\[1\.177,\\,1\.487\]\[0\.133,0\.311\]\[0\.133,\\,0\.311\]\[0\.304,0\.468\]\[0\.304,\\,0\.468\]\[0\.524,0\.558\]\[0\.524,\\,0\.558\]Persistent\[1\.699,2\.393\]\[1\.699,\\,2\.393\]\[0\.244,0\.635\]\[0\.244,\\,0\.635\]\[0\.752,0\.975\]\[0\.752,\\,0\.975\]\[1\.167,1\.372\]\[1\.167,\\,1\.372\]\[0\.228,0\.752\]\[0\.228,\\,0\.752\]\[0\.369,0\.640\]\[0\.369,\\,0\.640\]\[0\.553,0\.642\]\[0\.553,\\,0\.642\]OpenReviewer–\[1\.126,1\.708\]\[1\.126,\\,1\.708\]\[0\.096,0\.550\]\[0\.096,\\,0\.550\]\[0\.929,1\.251\]\[0\.929,\\,1\.251\]\[1\.328,1\.703\]\[1\.328,\\,1\.703\]\[0\.140,0\.279\]\[0\.140,\\,0\.279\]\[0\.300,0\.424\]\[0\.300,\\,0\.424\]\[0\.527,0\.556\]\[0\.527,\\,0\.556\]CycleReviewer–\[1\.159,1\.660\]\[1\.159,\\,1\.660\]\[−0\.172,0\.310\]\[\-0\.172,\\,0\.310\]\[0\.663,0\.913\]\[0\.663,\\,0\.913\]\[0\.895,1\.178\]\[0\.895,\\,1\.178\]\[0\.081,0\.176\]\[0\.081,\\,0\.176\]\[0\.292,0\.389\]\[0\.292,\\,0\.389\]\[0\.521,0\.545\]\[0\.521,\\,0\.545\]DeepReviewer–\[1\.204,1\.749\]\[1\.204,\\,1\.749\]\[0\.178,0\.587\]\[0\.178,\\,0\.587\]\[0\.362,0\.553\]\[0\.362,\\,0\.553\]\[0\.583,0\.784\]\[0\.583,\\,0\.784\]\[0\.087,0\.229\]\[0\.087,\\,0\.229\]\[0\.293,0\.433\]\[0\.293,\\,0\.433\]\[0\.523,0\.549\]\[0\.523,\\,0\.549\]AI Scientist–\[1\.219,1\.822\]\[1\.219,\\,1\.822\]\[−0\.171,0\.347\]\[\-0\.171,\\,0\.347\]\[0\.698,0\.989\]\[0\.698,\\,0\.989\]\[1\.027,1\.471\]\[1\.027,\\,1\.471\]\[0\.142,0\.273\]\[0\.142,\\,0\.273\]\[0\.240,0\.461\]\[0\.240,\\,0\.461\]\[0\.529,0\.559\]\[0\.529,\\,0\.559\]OpenJudge–\[1\.112,1\.504\]\[1\.112,\\,1\.504\]\[−0\.131,0\.362\]\[\-0\.131,\\,0\.362\]\[0\.244,0\.434\]\[0\.244,\\,0\.434\]\[0\.501,0\.678\]\[0\.501,\\,0\.678\]\[0\.107,0\.345\]\[0\.107,\\,0\.345\]\[0\.338,0\.512\]\[0\.338,\\,0\.512\]\[0\.509,0\.548\]\[0\.509,\\,0\.548\]ProReviewer–\[1\.260,1\.697\]\[1\.260,\\,1\.697\]\[−0\.017,0\.463\]\[\-0\.017,\\,0\.463\]\[0\.632,0\.796\]\[0\.632,\\,0\.796\]\[0\.882,1\.057\]\[0\.882,\\,1\.057\]\[0\.058,0\.128\]\[0\.058,\\,0\.128\]\[0\.275,0\.385\]\[0\.275,\\,0\.385\]\[0\.509,0\.524\]\[0\.509,\\,0\.524\]SciCore–\[0\.878,1\.278\]\[0\.878,\\,1\.278\]\[0\.248,0\.678\]\[0\.248,\\,0\.678\]\[0\.394,0\.558\]\[0\.394,\\,0\.558\]\[0\.600,0\.759\]\[0\.600,\\,0\.759\]\[0\.667,0\.835\]\[0\.667,\\,0\.835\]\[0\.545,0\.731\]\[0\.545,\\,0\.731\]\[0\.678,0\.760\]\[0\.678,\\,0\.760\]
### E\.5Science\-Core Preservation Diagnostics
The preservation analysis compares the embedding of a science\-core report extracted from an original manuscript with the embedding of the report extracted from its matched rhetorical variant:
cos\(E\(rioriginal\),E\(rikpvariant\)\)\.\\operatorname\{cos\}\\\!\\left\(E\(r\_\{i\}^\{\\mathrm\{original\}\}\),E\(r\_\{ikp\}^\{\\mathrm\{variant\}\}\)\\right\)\.Hereiiindexes papers,kkrewrite conditions, andpprewrite producers\. The condition\-level tables keep GPT\-5\.5 and Opus 4\.8 rewrite producers separate and report the mean, median, and fifth percentile\.
We embed reports withtext\-embedding\-3\-smallvia the OpenAI API using default API parameters\.
Table 14:Condition\-level science\-core similarity for GPT\-5\-mini\.P05 denotes the fifth percentile across comparisons\. The cross\-paper baseline compares variant cores with nonmatching original cores\.Table 15:Condition\-level science\-core similarity for GPT\-5\.5\.P05 denotes the fifth percentile across comparisons\. The cross\-paper baseline compares variant cores with nonmatching original cores\.Table 16:Condition\-level science\-core similarity for Gemini\-3\.5\-Flash\-Lite\.P05 denotes the fifth percentile across comparisons\. The cross\-paper baseline compares variant cores with nonmatching original cores\.
### E\.6Complete Science\-Core Weight Sensitivity Results
Figure[2](https://arxiv.org/html/2609.39027#S5.F2)provides a descriptive, direction\-normalized visualization of the science\-core weight sweep\. The complete unnormalized numerical results are reported in Tables[17](https://arxiv.org/html/2609.39027#A5.T17)–[19](https://arxiv.org/html/2609.39027#A5.T19)\. Each table fixes the manuscript protocol and varies the science\-core weightα\\alphafrom 0 \(manuscript only\) to 1 \(science core only\)\. Intermediate rows use the continuous post\-hoc fusion in Equation[3](https://arxiv.org/html/2609.39027#S5.E3)\.
Table 17:Complete science\-core weight sensitivity with Manuscript\-Standard\.All entries are unnormalized numerical results\. Arrows indicate the preferred direction\.𝜶\\bm\{\\alpha\}Within\-paper stabilityJoint stability\-discriminationHuman alignmentMAD↓\\downarrowDrift SD↓\\downarrowICC↑\\uparrowSPR↑\\uparrowDiscrim\.↑\\uparrowH\-MAE↓\\downarrowSpearman↑\\uparrow0\.000\.5980\.9560\.6150\.5550\.6491\.2940\.4480\.050\.5770\.9110\.6330\.5660\.6991\.2940\.4410\.100\.5560\.8670\.6520\.5770\.6991\.2930\.4410\.150\.5340\.8240\.6710\.5900\.6991\.2920\.4410\.200\.5130\.7840\.6900\.6030\.7011\.2910\.4410\.250\.4920\.7450\.7080\.6180\.7051\.2900\.4340\.300\.4700\.7080\.7260\.6330\.7111\.2930\.4250\.350\.4500\.6750\.7430\.6480\.7161\.2960\.4060\.400\.4300\.6440\.7580\.6640\.7171\.3000\.4060\.450\.4100\.6170\.7720\.6790\.7191\.3040\.4060\.500\.3900\.5940\.7830\.6930\.7231\.3080\.3960\.550\.3800\.5760\.7910\.7050\.7391\.3210\.3750\.600\.3690\.5620\.7970\.7160\.7381\.3330\.3690\.650\.3590\.5550\.8000\.7240\.7381\.3460\.3690\.700\.3500\.5520\.8000\.7300\.7421\.3580\.3690\.750\.3400\.5560\.7970\.7320\.7391\.3710\.3690\.800\.3300\.5650\.7920\.7320\.7371\.3870\.3690\.850\.3210\.5790\.7840\.7290\.7361\.4040\.3690\.900\.3110\.5990\.7750\.7240\.7351\.4230\.3690\.950\.3010\.6220\.7630\.7180\.7351\.4420\.3691\.000\.2920\.6500\.7510\.7100\.6951\.4610\.285Table 18:Complete science\-core weight sensitivity with Manuscript\-Strict\.All entries are unnormalized numerical results\. Arrows indicate the preferred direction\.𝜶\\bm\{\\alpha\}Within\-paper stabilityJoint stability\-discriminationHuman alignmentMAD↓\\downarrowDrift SD↓\\downarrowICC↑\\uparrowSPR↑\\uparrowDiscrim\.↑\\uparrowH\-MAE↓\\downarrowSpearman↑\\uparrow0\.000\.7661\.2020\.6320\.5070\.6721\.0780\.5290\.050\.7361\.1430\.6450\.5160\.7121\.0610\.5130\.100\.7071\.0850\.6590\.5260\.7121\.0440\.5130\.150\.6771\.0280\.6740\.5380\.7121\.0320\.5130\.200\.6480\.9720\.6890\.5500\.7131\.0220\.5130\.250\.6190\.9180\.7040\.5650\.7151\.0140\.5100\.300\.5890\.8660\.7200\.5800\.7191\.0160\.5040\.350\.5600\.8170\.7350\.5970\.7221\.0220\.4970\.400\.5320\.7700\.7490\.6150\.7231\.0390\.4970\.450\.5040\.7270\.7630\.6330\.7251\.0560\.4970\.500\.4760\.6870\.7750\.6520\.7261\.0720\.4880\.550\.4560\.6530\.7860\.6710\.7351\.1070\.4740\.600\.4350\.6230\.7940\.6880\.7341\.1420\.4510\.650\.4140\.6000\.8000\.7030\.7331\.1770\.4510\.700\.3960\.5850\.8020\.7160\.7451\.2120\.4150\.750\.3780\.5760\.8010\.7240\.7411\.2470\.4060\.800\.3610\.5760\.7970\.7290\.7421\.2860\.3970\.850\.3440\.5840\.7900\.7290\.7431\.3260\.3970\.900\.3260\.5990\.7790\.7260\.7421\.3710\.3970\.950\.3090\.6210\.7660\.7190\.7421\.4160\.3971\.000\.2920\.6500\.7510\.7100\.6951\.4610\.285Table 19:Complete science\-core weight sensitivity with Manuscript\-Persistent\.All entries are unnormalized numerical results\. Arrows indicate the preferred direction\.𝜶\\bm\{\\alpha\}Within\-paper stabilityJoint stability\-discriminationHuman alignmentMAD↓\\downarrowDrift SD↓\\downarrowICC↑\\uparrowSPR↑\\uparrowDiscrim\.↑\\uparrowH\-MAE↓\\downarrowSpearman↑\\uparrow0\.000\.4760\.8750\.6970\.5860\.6821\.2610\.4230\.050\.4610\.8300\.7120\.6010\.7271\.2690\.4080\.100\.4470\.7870\.7260\.6170\.7271\.2760\.4080\.150\.4320\.7460\.7410\.6340\.7271\.2840\.4080\.200\.4180\.7070\.7550\.6520\.7291\.2910\.4080\.250\.4030\.6700\.7690\.6690\.7321\.2990\.4080\.300\.3890\.6360\.7820\.6870\.7361\.3060\.4080\.350\.3740\.6040\.7940\.7040\.7361\.3140\.4080\.400\.3600\.5770\.8040\.7200\.7381\.3210\.4080\.450\.3470\.5540\.8120\.7340\.7391\.3290\.4080\.500\.3330\.5350\.8190\.7470\.7381\.3360\.4020\.550\.3290\.5220\.8230\.7560\.7521\.3490\.3780\.600\.3250\.5140\.8250\.7630\.7511\.3610\.3780\.650\.3210\.5120\.8240\.7660\.7511\.3740\.3780\.700\.3160\.5170\.8200\.7660\.7541\.3860\.3780\.750\.3120\.5270\.8140\.7620\.7501\.3990\.3780\.800\.3080\.5420\.8060\.7560\.7471\.4110\.3780\.850\.3040\.5630\.7950\.7470\.7451\.4240\.3780\.900\.3000\.5880\.7820\.7360\.7441\.4360\.3780\.950\.2960\.6180\.7670\.7230\.7441\.4490\.3781\.000\.2920\.6500\.7510\.7100\.6951\.4610\.285
## Appendix FCompute, Cost, and Efficiency
We report API expenditure by experimental task\.RobustReviewconstruction covers the full\-manuscript rewrites and reviewer\-guided feedback used to construct the controlled corpus\. Review\-only covers the general\-purpose reviewer grid and the repeated\-review audit\. TheSciCoretask includes science\-core extraction, the branch\-policy evaluations, and embedding; its manuscript branch reuses the GPT\-5\.5Strictreviews counted under Review\-only, and score fusion introduces no additional API call\. ReconstructReview includes manuscript reconstruction and final review while reusing cached science cores\. The AI Scientist and OpenJudge entries cover their released agentic review workflows\. The task\-level costs in Table[20](https://arxiv.org/html/2609.39027#A6.T20)are the recorded expenditures from the provider API consoles and include billable retries and failed attempts\.
Table 20:API expenditure by experimental task\.Review\-only costs are aggregated by billing route rather than by model\. Costs are the recorded expenditures from the provider API consoles\.OpenReviewer, CycleReviewer, DeepReviewer, and the trained ProReviewer backbone are excluded from the API expenditure total because they were served locally on a server equipped with 8 NVIDIA A100 GPUs\.Similar Articles
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
This paper investigates how rhetorical framing biases AI-based peer review scores, finding that evidence framing and novelty stance have the largest effects and that score movements depend on the reviewer's initial score and evaluation strictness.
When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation
This paper studies the feedback loop where AI-generated reviews influence future training of AI reviewers, leading to reduced judgment diversity called 'scientific-judgment collapse,' and introduces TrustReviewer, an open-source system to mitigate this through curated training and activation steering.
Does AI Reviewer See the Full Picture? Attacking and Defending Multimodal Peer Review
This paper introduces PaperGuard, a benchmark for evaluating and defending against adversarial attacks on multimodal AI peer review systems, covering both text and figure-based attacks across multiple scientific domains.
Benchmarking Agentic Review Systems
This paper benchmarks agentic review systems for peer review, evaluating open-source and proprietary systems on research papers. The best configuration achieves 83.0% pairwise accuracy and catches 71.6% of injected errors, but user feedback highlights issues with false positives and nitpicks.
Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review
This paper proposes a benchmarking protocol using automated multi-model LLM review to evaluate AI Scientist systems, comparing frameworks like Sakana AI, CycleResearcher, and Data-to-Paper, and finds that FARS benchmark papers significantly outperform other systems.