AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification

arXiv cs.CL 论文

摘要

This paper introduces AutoSupervision, a benchmark and method for verifying whether manuscript revisions actually address reviewer concerns using grounded evidence from peer-review records. Experiments on 56,000 Nature Communications articles show LLMs can summarize reviewer concerns but still struggle with evidence-based verification.

arXiv:2607.27845v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have enabled AI systems to assist scientific research and peer review. However, an essential capability for reliable AI-assisted scientific workflows remains underexplored: verifying whether reviewer feedback leads to meaningful and evidence-supported manuscript improvements. We introduce AutoSupervision, which evaluates whether scientific manuscript revisions genuinely address reviewer concerns through grounded evidence. AutoSupervision leverages transparent peer-review records as a natural source of supervision, where reviewer comments specify scientific concerns, author responses describe claimed resolutions, and revised manuscripts provide evidence of changes. Given reviewer comments, author responses, and revised manuscripts, models must characterize reviewer concerns, determine whether concerns have been addressed, and identify supporting manuscript evidence. We construct AutoSupervision from 56,000 Nature Communications articles and corresponding review records. Then we conducted experiments on LLMs, the ablation study, and the case study. Our results show that while LLMs perform well in characterizing reviewer concerns, with GPT-5.5 achieving a score of 0.754, evidence-based verification remains the primary bottleneck, with the best-performing model reaching only 0.501.
查看原文
查看缓存全文

缓存时间: 2026/07/31 10:03

# AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification
Source: [https://arxiv.org/html/2607.27845](https://arxiv.org/html/2607.27845)
Haobo Li1,Eunseo Jung3,Wenxiao Zhao2,Feng Liu1,Jiong Wang1,Kaiyi Xu1, Zijie Guo1,Zixin Chen3,Ben Fei1,Fenghua Ling1,Lei Bai1 1Shanghai AI Laboratory2University of California, Los Angeles 3The Hong Kong University of Science and Technology

###### Abstract

Recent advances in large language models \(LLMs\) have enabled AI systems to assist scientific research and peer review\. However, an essential capability for reliable AI\-assisted scientific workflows remains underexplored: verifying whether reviewer feedback leads to meaningful and evidence\-supported manuscript improvements\. We introduceAutoSupervision, which evaluates whether scientific manuscript revisions genuinely address reviewer concerns through grounded evidence\.AutoSupervisionleverages transparent peer\-review records as a natural source of supervision, where reviewer comments specify scientific concerns, author responses describe claimed resolutions, and revised manuscripts provide evidence of changes\. Given reviewer comments, author responses, and revised manuscripts, models must characterize reviewer concerns, determine whether concerns have been addressed, and identify supporting manuscript evidence\. We constructAutoSupervisionfrom 56,000*Nature Communications*articles and corresponding review records\. Then we conducted experiments on LLMs, the ablation study, and the case study\. Our results show that while LLMs perform well in characterizing reviewer concerns, with GPT\-5\.5 achieving a score of 0\.754, evidence\-based verification remains the primary bottleneck, with the best\-performing model reaching only 0\.501\.

## 1Introduction

Scientific discovery is not a one\-shot generation process, but an iterative cycle in which claims are proposed, challenged, revised, and re\-evaluated in light of evidence and critique\(Wanget al\.,[2023](https://arxiv.org/html/2607.27845#bib.bib20)\)\. In scholarly publishing, peer review institutionalizes this cycle: reviewers identify weaknesses, request additional evidence or clarification, authors revise their manuscripts, and reviewers or editors then assess whether the new version has actually addressed the raised concerns\(Mulliganet al\.,[2013](https://arxiv.org/html/2607.27845#bib.bib29)\)\. Recent advances in LLMs have motivated a vision of AI\-assisted scientific discovery, where AI systems participate not only in generating research artifacts but also in evaluating and revising them\. This raises a central question for iterative AI scientific workflows: can AI systems verify, with grounded manuscript evidence, whether feedback has actually been resolved in a subsequent revision?

Current AI\-driven scientific workflows \(Figure[1](https://arxiv.org/html/2607.27845#S1.F1)\) have begun to address the first two parts of this cycle: generating scientific artifacts and producing feedback on them\. AutoResearch systems aim to automate or assist activities such as literature analysis, hypothesis generation, experiment design, coding, and manuscript preparation\(Luet al\.,[2024](https://arxiv.org/html/2607.27845#bib.bib22); Schmidgallet al\.,[2025](https://arxiv.org/html/2607.27845#bib.bib23); Ghareebet al\.,[2026](https://arxiv.org/html/2607.27845#bib.bib21)\)\. In parallel, AutoReview systems investigate whether LLM\-based agents can evaluate manuscripts, identify weaknesses, generate reviewer comments, and assist human reviewers\(Liu and Shah,[2023](https://arxiv.org/html/2607.27845#bib.bib11); Zhouet al\.,[2024](https://arxiv.org/html/2607.27845#bib.bib12); Latonaet al\.,[2024](https://arxiv.org/html/2607.27845#bib.bib14); Caoet al\.,[2025](https://arxiv.org/html/2607.27845#bib.bib26); Jiang and Ng,[2025](https://arxiv.org/html/2607.27845#bib.bib28); PapeReview,[2026](https://arxiv.org/html/2607.27845#bib.bib27)\)\. However, these systems mostly stop at generating artifacts or critiques\. They do not directly verify whether a subsequent revision substantively addresses the critique, whether the author’s response is reflected in the revised manuscript, or whether the revised text provides sufficient evidence for the paper’s scientific claims\.

![Refer to caption](https://arxiv.org/html/2607.27845v1/x1.png)Figure 1:AutoSupervisionfor scientific revision verification\. Black loop: workflow withoutAutoSupervision\. Orange loop: workflow withAutoSupervision: verify whether reviewer concerns are truly addressed\.Revision verification poses challenges beyond lexical or semantic matching\. Reviewer concerns may target experimental design, claim\-evidence alignment, baseline coverage, figure clarity, statistical reporting, or the scope of conclusions, while corresponding revisions may appear as local clarifications, new analyses, additional experiments, revised visualizations, added limitations, or distributed changes across sections\. A verifier must therefore recover the intent of the original concern, interpret the author’s claimed resolution, and assess whether the revised manuscript provides evidence that substantively addresses the underlying scientific issue\. This assessment must also be grounded in the manuscript rather than in the response letter alone\. Author responses describe claimed resolutions, but they do not by themselves establish that the manuscript has changed in the relevant way\. A response may overstate the revision, address only part of the concern, or describe a change whose evidential support remains insufficient\. Reliable verification must therefore jointly reason over the reviewer concern, the author response, and concrete manuscript evidence\.

Existing peer\-review datasets provide valuable resources for modeling review artifacts, but they do not directly support grounded revision verification\. PeerRead, NLPeer, and MOPRD include combinations of papers, reviews, rebuttals, manuscript versions, meta\-reviews, and decisions\(Kanget al\.,[2018](https://arxiv.org/html/2607.27845#bib.bib1); Dyckeet al\.,[2023](https://arxiv.org/html/2607.27845#bib.bib4); Linet al\.,[2023](https://arxiv.org/html/2607.27845#bib.bib5)\)\. Related resources study meta\-review generation, review discussion modeling, review\-rebuttal argument relations, review quality assessment, and substantiation of review statements\(Chenget al\.,[2020](https://arxiv.org/html/2607.27845#bib.bib8); Shenet al\.,[2022](https://arxiv.org/html/2607.27845#bib.bib6); Kennardet al\.,[2022](https://arxiv.org/html/2607.27845#bib.bib7); Ghosalet al\.,[2022](https://arxiv.org/html/2607.27845#bib.bib9); Guoet al\.,[2023](https://arxiv.org/html/2607.27845#bib.bib10)\)\. These datasets have enabled substantial progress in computational peer\-review research, but they are primarily organized around reviews, rebuttals, decisions, or discussion structures\. They do not define concern\-level revision verification as a benchmark task, where a model must jointly read the reviewer concern, the author response, and the revised manuscript to decide whether the concern was resolved and to identify the manuscript evidence supporting that judgment\. This leaves the following task largely unaddressed:

> Given a reviewer’s concern, an author’s response, and a revised manuscript, can a model characterize the concerns and whether the concerns have been addressed and identify the exact manuscript evidence supporting this judgment?

To make this question benchmarkable, we introduceAutoSupervision, a benchmark for AutoSupervision that evaluates concern\-level, evidence\-grounded verification of scientific manuscript revisions\. Each instance is centered on a reviewer concern within a revision episode\. Given the concern, the corresponding author response, and the revised manuscript, a model must characterize the concern, determine whether it has been resolved, and identify the manuscript evidence supporting its judgment\. This formulation isolates revision assessment from concern discovery, so model performance reflects the ability to verify revisions rather than the ability to extract reviewer comments\.

We constructAutoSupervisionfrom transparent peer\-review records of*Nature Communications*, which provide reviewer comments and author responses alongside published articles\(Anonymous,[2022](https://arxiv.org/html/2607.27845#bib.bib16)\)\. From 56,000 article\-level records, we extract revision episodes and concern\-level verification instances by aligning reviewer concerns, author responses, and manuscript evidence\. We then evaluate LLMs, supervised models, and agentic verification systems\. The results show that modern models can often characterize reviewer concerns, but still struggle to determine whether revisions genuinely resolve them and to ground their judgments in manuscript evidence\.

Our contributions are summarized as follows:

- •We define AutoSupervision as a new evaluation problem for AI\-assisted scientific workflows and introduceAutoSupervision, a benchmark for concern\-level, evidence\-grounded verification of scientific manuscript revisions\.
- •We construct a large\-scale dataset from 56,000 transparent peer\-review article records from*Nature Communications*and develop a systematic pipeline to align reviewer concerns, author responses, and manuscript revisions, yielding 8,790 episode\-level instances with resolution labels and evidence annotations\.
- •We provide comprehensive evaluations across LLMs, supervised models, agentic systems, and ablation studies, revealing the remaining challenges in evidence\-grounded revision assessment\.

## 2Related Work

#### Resources for computational peer review\.

Computational studies of peer review have produced several valuable resources for modeling scientific evaluation\. PeerRead pairs paper drafts with reviews and acceptance decisions, enabling research on review generation, review analysis, and acceptance prediction\(Kanget al\.,[2018](https://arxiv.org/html/2607.27845#bib.bib1)\)\. NLPeer unifies peer\-review datasets with structured representations, licensing information, and reviewing\-assistance tasks\(Dyckeet al\.,[2023](https://arxiv.org/html/2607.27845#bib.bib4)\)\. MOPRD introduces a multidisciplinary open peer\-review dataset containing review comments, rebuttal letters, manuscript versions, meta\-reviews, and decisions\(Linet al\.,[2023](https://arxiv.org/html/2607.27845#bib.bib5)\)\. Peer Review Analyze further annotates ICLR reviews with paper\-section correspondence, aspect categories, statement purposes, and significance labels\(Ghosalet al\.,[2022](https://arxiv.org/html/2607.27845#bib.bib9)\)\. These datasets enable computational studies of peer review, but primarily focus on review artifacts, paper\-level decisions, or properties of reviewer feedback rather than whether feedback is successfully incorporated into revised manuscripts\.

#### Modeling scientific feedback and revision processes\.

Prior work has studied how reviewers and authors interact through review comments and response documents\. APE introduced argument pair extraction between peer reviews and rebuttals, enabling analysis of reviewer arguments and author responses\(Chenget al\.,[2020](https://arxiv.org/html/2607.27845#bib.bib8)\)\. DISAPERE annotated discourse structures in review\-rebuttal pairs and characterized author stances toward reviewer arguments\(Kennardet al\.,[2022](https://arxiv.org/html/2607.27845#bib.bib7)\)\. MReD studied structured meta\-review generation by modeling review discussions and decision rationales\(Shenet al\.,[2022](https://arxiv.org/html/2607.27845#bib.bib6)\)\. These efforts provide important foundations for understanding scientific feedback exchange\. However, they mainly treat rebuttals as the endpoint of the interaction\. In contrast,AutoSupervisionincorporates the revised manuscript as an additional component and evaluates the complete feedback chain from reviewer concern to author response, to evidence of manuscript change\.

![Refer to caption](https://arxiv.org/html/2607.27845v1/x2.png)Figure 2:AutoSupervisionconstruction and evaluation pipeline\. First, we pair articles and peer reviews \(A1\), parse them into blocks with identifiers \(A2\), extract concerns \(A3\), and export revision episodes for concern characterization, resolution verification, and evidence grounding \(A4\)\. After buildingAutoSupervision, we benchmark existing LLMs on this task \(B\)\.
#### LLMs for scientific research and review assistance\.

Recent advances in LLMs have motivated AutoResearch systems that automate or assist scientific research such as literature analysis, hypothesis generation, experimentation, coding, and manuscript preparation\(Luet al\.,[2024](https://arxiv.org/html/2607.27845#bib.bib22); Schmidgallet al\.,[2025](https://arxiv.org/html/2607.27845#bib.bib23); Ghareebet al\.,[2026](https://arxiv.org/html/2607.27845#bib.bib21)\)\. In parallel, AutoReview systems explore the use of LLMs to evaluate manuscripts and provide scientific feedback\. ReviewerGPT investigates LLM\-based generation of reviewer comments and manuscript critiques\(Liu and Shah,[2023](https://arxiv.org/html/2607.27845#bib.bib11)\)\. Other studies evaluate the quality and usefulness of LLM\-generated reviews and research\-paper feedback\(Zhouet al\.,[2024](https://arxiv.org/html/2607.27845#bib.bib12); Lianget al\.,[2024](https://arxiv.org/html/2607.27845#bib.bib13)\)\. CSPaper Review provides fast, rubric\-faithful feedback aligned with conference review criteria\(Caoet al\.,[2025](https://arxiv.org/html/2607.27845#bib.bib26)\)\. PapeReview offers an agentic pre\-submission review workflow that produces structured feedback on manuscript clarity, methodology, and contribution\(PapeReview,[2026](https://arxiv.org/html/2607.27845#bib.bib27)\)\. The Stanford Agentic Reviewer uses an agentic workflow that retrieves relevant work from arXiv and synthesizes it with manuscript content to generate actionable feedback\(Jiang and Ng,[2025](https://arxiv.org/html/2607.27845#bib.bib28)\)\. REVAS has also been deployed in ACL Rolling Review to provide reviewers with feedback on the quality of their reviews\(ACL Rolling Review,[2026](https://arxiv.org/html/2607.27845#bib.bib24)\)\.

These systems automate or assist the generation and improvement of scientific feedback\. However, they do not evaluate whether reviewer feedback is subsequently incorporated into a revised manuscript\.AutoSupervisionaddresses this complementary stage by characterizing reviewer concerns, verifying their resolution, and grounding the resulting judgments in manuscript evidence\.

## 3Task Definition

### 3\.1Scientific Revision Loop

We formalize scientific revision as an interaction among three roles: an author process, a reviewer process, and an AutoSupervision model\. Let𝒜\\mathcal\{A\}denote the author,ℛ\\mathcal\{R\}denote the reviewer process, and𝒮θ\\mathcal\{S\}\_\{\\theta\}denote the AutoSupervision model to be evaluated\. At revision roundtt, the manuscript before revision isMt−1M\_\{t\-1\}\. A reviewer observes the manuscript and produces a set of reviewer concerns:

Ct=ℛ​\(Mt−1\)=\{ct,1,…,ct,nt\},C\_\{t\}=\\mathcal\{R\}\(M\_\{t\-1\}\)=\\\{c\_\{t,1\},\\ldots,c\_\{t,n\_\{t\}\}\\\},\(1\)wherennis the number of concerns at roundtt\. The author then responds to these concerns and revises the manuscript:

\(Mt,At\)=𝒜​\(Mt−1,Ct\),\(M\_\{t\},A\_\{t\}\)=\\mathcal\{A\}\(M\_\{t\-1\},C\_\{t\}\),\(2\)whereMtM\_\{t\}is the post\-revision manuscript andAt=\{at,1,…,at,nt\}A\_\{t\}=\\\{a\_\{t,1\},\\ldots,a\_\{t,n\_\{t\}\}\\\}denotes the author responses or rebuttals corresponding to the reviewer concerns\.

For each concernct,ic\_\{t,i\}, we define a review\-response record

rt,i=\(ct,i,at,i\),r\_\{t,i\}=\(c\_\{t,i\},a\_\{t,i\}\),\(3\)and letRt=\{rt,1,…,rt,nt\}R\_\{t\}=\\\{r\_\{t,1\},\\ldots,r\_\{t,n\_\{t\}\}\\\}be the set of records associated with revision roundtt\. In the benchmark,𝒜\\mathcal\{A\}andℛ\\mathcal\{R\}are not models to be trained\. They are observed through public transparent peer\-review records: reviewer comments instantiateCtC\_\{t\}, author responses instantiateAtA\_\{t\}, and published manuscript revisions instantiateMtM\_\{t\}\. The task is to evaluate whether𝒮θ\\mathcal\{S\}\_\{\\theta\}can audit the result of this interaction\.

### 3\.2Round\-Level Episodes

Peer\-review data contains multiple manuscript versions across revision rounds\. We therefore define a*revision episode*as a transition between two consecutive manuscript versions, such asv0\_to\_v1orv1\_to\_v2\. Formally, an episode is

et=\(Mt,Rt\)\.e\_\{t\}=\(M\_\{t\},R\_\{t\}\)\.\(4\)This formulation evaluates each concern together with the author response and manuscript state available at the corresponding revision stage, rather than aggregating information across the entire review history\.

### 3\.3AutoSupervision Prediction Task

Given an episodeete\_\{t\}, the AutoSupervision model reads the post\-revision manuscript together with the full set of review\-response records from that round and predicts a set of structured outputsY^t\\hat\{Y\}\_\{t\}:

Y^t=𝒮θ​\(et\)=\{y^t,i\}i=1nt\.\\hat\{Y\}\_\{t\}=\\mathcal\{S\}\_\{\\theta\}\(e\_\{t\}\)=\\\{\\hat\{y\}\_\{t,i\}\\\}\_\{i=1\}^\{n\_\{t\}\}\.\(5\)Each predicted element corresponds to one recordrt,ir\_\{t,i\}\. It contains concern\-characterization fields, revision\-verification fields, review\-record evidence block IDs, and revised\-manuscript grounding block IDs\.

The gold output for the same episodeYt∗Y^\{\*\}\_\{t\}is

Yt∗=\{yt,i∗\}i=1nt\.Y^\{\*\}\_\{t\}=\\\{y^\{\*\}\_\{t,i\}\\\}\_\{i=1\}^\{n\_\{t\}\}\.\(6\)Evaluation comparesY^t\\hat\{Y\}\_\{t\}withYt∗Y^\{\*\}\_\{t\}for every episode, aligning predictions by the fixed review\-response record identifiers\.

This task is different from automatic review generation and automatic manuscript revision\. The model is not asked to generate reviewer concernsCtC\_\{t\}, write author responsesAtA\_\{t\}, or produce a revised manuscriptMtM\_\{t\}\. Instead,𝒮θ\\mathcal\{S\}\_\{\\theta\}audits the closed feedback loop: given what the reviewer requested, what the author claimed, and what appears in the revised manuscript, it must determine whether the concern was addressed and cite the exact evidence supporting that judgment\.

The full prediction schema and label definitions are provided in Appendix[A](https://arxiv.org/html/2607.27845#A1)\.

## 4Dataset Construction

### 4\.1Pipeline

Figure[2](https://arxiv.org/html/2607.27845#S2.F2)\(A\) summarizes the dataset construction pipeline\.

#### Source Corpus and Splits

We collect 56,000 published articles and corresponding transparent peer\-review files from*Nature Communications*\. Within each publication year, papers are assigned to training, validation, and test partitions using an approximately8:1:18\{:\}1\{:\}1ratio\. All review rounds and manuscript states from the same paper remain in the same partition, preventing cross\-round information leakage while preserving the publication\-year distribution across splits\.

StatisticCountSource\-corpus papers56,000▶\\scriptscriptstyle\\blacktrianglerightTraining split44,796▶\\scriptscriptstyle\\blacktrianglerightValidation split5,594▼\\scriptstyle\\blacktriangledownTest split5,610▼\\scriptstyle\\blacktriangledownSampled test papers557Revision episodes847Unique reviewer concerns6,543Episode\-level concern instances8,790Table 1:Source\-corpus and benchmark statistics\.
#### Document Parsing

Both the published articles and their corresponding peer\-review files are downloaded in PDF format\. Because PDFs primarily encode visual layout rather than structured document content, we first parse them into ordered, typed blocks using MinerU\(Wanget al\.,[2024](https://arxiv.org/html/2607.27845#bib.bib18)\)\. The resulting blocks include paragraphs, section headings, figures, tables, equations, references, captions, and other scientific content\. Each manuscript block receives a stable identifier with the prefixp, while each peer\-review block receives an identifier with the prefixr\.

#### Concern Extraction

For each sampled paper, we use GPT\-5\.5\(OpenAI,[2026](https://arxiv.org/html/2607.27845#bib.bib25)\)to process the published article and its peer review\. A peer review may contain multiple rounds of reviewer comments, author responses, and reviewer follow\-up messages\. The model extracts candidate concern threads, links each reviewer concern to its corresponding author response and follow\-up discussion, and assigns the resulting review\-response record to the revision episode in which the response was made\.

During benchmark construction, the model has access to the complete public review history\. This retrospective context provides useful evidence\. For example, a later reviewer follow\-up may explicitly confirm that a concern was addressed, restate an unresolved issue, or narrow the remaining objection after examining the authors’ revision\. Reviewer follow\-up information is used as gold information only during benchmark construction for concern\-response alignment and quality filtering\. It is not provided as input during evaluation\.

Each extracted record preserves the original reviewer concern, the corresponding author response, the normalized reviewer identifier, the assigned revision round, and the supporting review\-block identifiers\. It also includes the concern\-characterization, resolution, and target\-scope annotations defined by the benchmark schema\.

#### Episode Export

We apply conservative filtering to ensure that retained instances represent actionable and auditable scientific revision requests\. We exclude:

- •praise\-only comments and broad paper\-level assessments without a specific revision request;
- •concerns without verifiable reviewer\-comment evidence;
- •concerns whose targets exist only in the rebuttal letter, an external repository, or other unavailable artifacts\.

When a concern recurs across multiple revision rounds, we retain the reviewer comment and corresponding author response for each relevant episode while storing later follow\-up discussion\. This prevents short follow\-up statements, such as “OK” or “Addressed,” from replacing the original scientific concern and response context\.

For each valid transition from manuscript versiont−1t\-1to versiontt, we export one revision episodeete\_\{t\}\. Each episode includes the complete set of manuscript blocks available after the revision and only the review information associated with the corresponding round\. Information from later rounds may support annotation and filtering but is never exposed as input for an earlier evaluation episode\.

### 4\.2Dataset Statistics

For the benchmark evaluation in this paper, we sample 10% of the test papers from each year using a fixed random seed, yielding 557 candidate papers\. After review\-response extraction, revision alignment, and quality filtering, the resulting benchmark contains 847 revision episodes\. Across these episodes,AutoSupervisioncontains 6,543 unique point\-level reviewer concerns\. Because a concern may remain active or recur across multiple revision rounds, the unique concerns correspond to 8,790 episode\-level concern instances\. These episode\-level instances are the primary units used for benchmark evaluation\. More details about dataset distribution are provided in Appendix[B\.1](https://arxiv.org/html/2607.27845#A2.SS1)\. To evaluate the robustness of our annotations to the choice of model, we measure inter\-model agreement among the annotation outputs, as reported in Appendix[B\.2](https://arxiv.org/html/2607.27845#A2.SS2)\.

Table 2:MainAutoSupervisionperformance results\. Coverage is concern\-level prediction coverage, and Overall is the equal\-weight average of characterization, verification, and manuscript grounding\. Bold and underlined values indicate the best and second\-best performance per backbone model, respectively\.

## 5Evaluation

### 5\.1Evaluation Pipeline

Figure[2](https://arxiv.org/html/2607.27845#S2.F2)\(B\) summarizes the Evaluation pipeline\. We evaluate models under a unified structured prediction setting\. All models must output JSON conforming to the benchmark schema\. Missing concern IDs are scored as missing predictions, and exact concern\-set match is reported separately from label and grounding scores\. The evaluator aligns predictions by the fixed input concern IDs, so systems do not receive credit for identifying concerns outside the provided evaluation set\.

The benchmark evaluates three complementary capabilities: Characterization, Verification, and Grounding\. The Characterization score averages comment\-kind macro\-F1, point\-type macro\-F1, and a hierarchical target\-scope score over target domain, section, and object\. The Verification score averages resolution\-label macro\-F1, paper\-status macro\-F1, and micro\-F1 over evidence block IDs supporting the resolution judgment\. The Grounding score is computed as micro\-F1 over target paper block IDs, which identify the revised\-manuscript evidence blocks where a reviewer concern is addressed\. Please refer to Appendix[A](https://arxiv.org/html/2607.27845#A1)for details\. The overall score assigns equal weights to the three components:

Soverall=13​\(Schar\+Sver\+Sground\)\.S\_\{\\mathrm\{overall\}\}=\\frac\{1\}\{3\}\\left\(S\_\{\\mathrm\{char\}\}\+S\_\{\\mathrm\{ver\}\}\+S\_\{\\mathrm\{ground\}\}\\right\)\.

### 5\.2Evaluated Models

We evaluateAutoSupervisionon a diverse set of recent LLMs, covering both closed\-source and open\-source models\. The evaluated models span different model families and scales, enabling a comprehensive assessment of current capabilities for reviewer concern analysis\.

#### Closed\-source models\.

We include several proprietary frontier models, including Claude\-Opus\-4\.8\(Anthropic,[2026](https://arxiv.org/html/2607.27845#bib.bib35)\), GPT\-5\.5\(OpenAI,[2026](https://arxiv.org/html/2607.27845#bib.bib25)\), and Gemini\-3\.1\-Pro\(Google DeepMind,[2026](https://arxiv.org/html/2607.27845#bib.bib37)\)\. These models represent state\-of\-the\-art general\-purpose language models with strong reasoning and instruction\-following capabilities\. We additionally evaluate GPT\-4o\-mini\(OpenAI,[2024](https://arxiv.org/html/2607.27845#bib.bib45)\)and its agentic variant to examine the performance gap between earlier\-generation compact models and recent frontier systems\.

#### Open\-source models\.

We evaluate a broad range of open\-source models, including GLM\-5\.2\(Z\.ai,[2026](https://arxiv.org/html/2607.27845#bib.bib36)\), Qwen3\.7\-max\(Qwen Team,[2026b](https://arxiv.org/html/2607.27845#bib.bib38)\), DeepSeek\-V4\-Pro\(Xuet al\.,[2026](https://arxiv.org/html/2607.27845#bib.bib40)\), Kimi\-K2\.6\(Moonshot AI,[2026](https://arxiv.org/html/2607.27845#bib.bib41)\), DeepSeek\-V3\.2\(Liuet al\.,[2025](https://arxiv.org/html/2607.27845#bib.bib42)\), MiniMax\-2\.7\(MiniMax,[2026](https://arxiv.org/html/2607.27845#bib.bib43)\), DeepSeek\-V4\-Flash\(Xuet al\.,[2026](https://arxiv.org/html/2607.27845#bib.bib40)\), Intern\-S2\-35B\(InternLM Team,[2026](https://arxiv.org/html/2607.27845#bib.bib44)\), and Qwen3\.5\-9B\(Qwen Team,[2026a](https://arxiv.org/html/2607.27845#bib.bib39)\)\. We also include an instruction\-tuned Qwen3\.5\-9B model to investigate the impact of supervised fine\-tuning on reviewer concern understanding\.

### 5\.3Main Results

Table[2](https://arxiv.org/html/2607.27845#S4.T2)and Figure[3](https://arxiv.org/html/2607.27845#S5.F3)summarize benchmark performance across closed\-source and open\-source models\. Overall, the results reveal a substantial capability gap between identifying reviewer concerns and performing evidence\-based assessment of revisions\. While modern models can almost always extract and characterize reviewer concerns, accurately verifying whether a revision resolves the concern remains a significant challenge\.

Among all evaluated systems, Claude\-Opus\-4\.8 and GPT\-5\.5 achieve the highest overall performance \(both 0\.637\), followed by GLM\-5\.2 \(0\.628\) and Gemini\-3\.1\-Pro \(0\.624\)\. Interestingly, the performance difference between the strongest closed\-source and open\-source models is relatively small, with GLM\-5\.2 approaching frontier proprietary systems\. However, their remaining limitations are primarily concentrated in deeper reasoning tasks that require evaluating the relationship between a claimed revision and the underlying manuscript evidence\.

#### Coverage

Coverage is concern\-level prediction coverage\. Nearly all models achieve perfect or near\-perfect coverage, indicating that modern LLMs rarely fail to identify the presence of reviewer concerns\. The only notable degradation occurs for smaller or earlier\-generation models, such as GPT\-4o\-mini\. This suggests that concern extraction is largely a solved capability and is no longer the primary bottleneck of reviewer\-assistance systems\.

#### Characterization

Characterization measures whether a model correctly identifies the type and scope of a reviewer concern\. GPT\-5\.5 achieves the strongest performance \(0\.754\), while most frontier models exceed 0\.65\. The high characterization scores, combined with near\-perfect coverage, indicate that models possess strong semantic understanding of reviewer feedback\.

#### Verification

Verification is consistently the most challenging dimension\. Claude\-Opus\-4\.8 obtains the highest verification score \(0\.501\), followed by the Qwen3\.5\-9B SFT model \(0\.490\) and GLM\-5\.2 \(0\.475\)\. This pattern is notable because verification does not simply require language understanding; it requires multi\-step reasoning over the reviewer concern, the claimed revision, and the manuscript content\. The results suggest that current models often recognize the intent of a revision but struggle to determine whether the provided evidence is sufficient, complete, and scientifically relevant\.

#### Grounding

Manuscript grounding evaluates whether models can identify supporting evidence\. Claude\-Opus\-4\.8 achieves the highest grounding score \(0\.719\), followed closely by GPT\-5\.5 \(0\.702\), GLM\-5\.2 \(0\.711\), and Gemini\-3\.1\-Pro \(0\.700\)\. In contrast, smaller models show substantial degradation, with GPT\-4o\-mini achieving only 0\.085\. The large performance gap highlights that grounding requires more than generating plausible explanations: models must accurately retrieve and connect manuscript\-level evidence to reviewer concerns\.

Overall, these results demonstrate thatAutoSupervisionevaluates capabilities beyond reviewer feedback understanding\. Existing LLMs already achieve strong performance in detecting and characterizing concerns, but they remain limited in evidence\-based verification and grounding\.

![Refer to caption](https://arxiv.org/html/2607.27845v1/x3.png)Figure 3:Model performance across characterization, verification, and grounding\. Gray segments show the three component scores, and the black diamond denotes the equal\-weight overall score\.Table 3:Ablation study\. GPT\-5\.5 performance under different input configurations\.

### 5\.4Impact of Agentic Inference

Beyond scaling model size, we evaluate whether a fixed structured inference pipeline can improve performance\. The agentic baseline uses GPT\-4o\-mini in a three\-stage pipeline for concern characterization, manuscript grounding, and resolution verification\. Between characterization and grounding, we add a deterministic manuscript\-block retriever that narrows the full manuscript to a high\-recall candidate evidence set for each concern\. The model then performs grounding and verification over this smaller candidate context and outputs the same schema as all other systems\. Appendix[C](https://arxiv.org/html/2607.27845#A3)provides the full pipeline details\.

This pipeline improves GPT\-4o\-mini substantially over the single\-shot setting: overall score increases from 0\.266 to 0\.403, coverage from 0\.933 to 0\.998, and manuscript grounding from 0\.085 to 0\.336\. The improvement shows that explicit retrieval, decomposition, and schema normalization help models produce more complete and better\-grounded predictions\.

### 5\.5Impact of Supervised Fine\-tuning

We further examine the effect of task\-specific supervised fine\-tuning by comparing Qwen3\.5\-9B with its supervised variant\. 500 papers are randomly sampled from the training dataset split for training\. Appendix[D](https://arxiv.org/html/2607.27845#A4)gives reproduce details, including the data construction, model, training, and inference details\.

SFT Qwen3\.5\-9B improves overall performance from 0\.463 to 0\.614, with gains across all three components: characterization improves from 0\.563 to 0\.708, verification from 0\.376 to 0\.490, and grounding from 0\.451 to 0\.643\. The improvements indicate that supervised adaptation helps models better capture the knowledge ofAutoSupervision, particularly the alignment between reviewer concerns, author responses, and manuscript evidence\. The substantial gain in grounding further suggests that task\-specific supervision improves evidence localization, enabling models to better identify relevant manuscript support for reviewer feedback\.

### 5\.6Ablation Study

Table[3](https://arxiv.org/html/2607.27845#S5.T3)analyzes how different sources of information contribute toAutoSupervisionperformance\. We progressively remove components from the full\-context setting to understand which capabilities depend on reviewer information, author responses, and manuscript evidence\.

When only reviewer concerns are provided, GPT\-5\.5 achieves a relatively strong characterization score \(0\.676\), but verification performance drops substantially \(0\.042\), and grounding is not valid by design because no manuscript evidence is available\. This indicates that reviewer text contains sufficient information to identify the type and scope of a concern, but is insufficient for determining whether the concern has been addressed\.

Adding author responses improves verification to 0\.236 while maintaining strong characterization \(0\.696\)\. However, grounding remains unavailable without manuscript evidence\. This suggests that response letters provide useful information about the claimed resolution, but these claims cannot be validated without examining the actual revision\.

Providing concerns and manuscript inference substantially improves grounding, achieving 0\.633 grounding and 0\.582 overall performance\. This demonstrates that manuscript evidence is the primary source for locating revision outcomes\. Nevertheless, manuscript\-only inference remains below full\-context performance, particularly on verification \(0\.394 vs\. 0\.455\), because the model lacks the original reviewer intent and the author’s explanation of the change\.

### 5\.7Closed\-Loop Revision Case Study

We also test whether the benchmark is useful for choosing a verifier that improves a downstream revision loop\. This case study asks whether systems that are stronger under the benchmark\-style supervisor task provide more useful feedback when inserted between two author\-revision calls\. The pipeline starts from a initial submission draft, uses fixed human reviewer concerns from the review record, asks GPT\-5\.5 to produce a first revision, applies a ReviseBench\-style supervisor to the first revision, renders the supervisor output as author\-facing feedback, asks GPT\-5\.5 for a second revision, and finally uses a blinded Claude Opus 4\.8 judge to compare concern\-level outcomes\. We evaluate 10 sampled papers\. The no\-supervision condition receives the raw reviews and author\-response context but no structured concern\-level supervisor output\.

Table 4:Closed\-loop revision case study\. Outcomes are concern\-level pairwise comparisons between the first and second revision\. Net is the improved minus the worse\.GLM\-5\.2 produces the largest number of useful second\-revision changes\. DeepSeek V4 Flash is weaker but still improves over the no\-supervisor condition\. Without supervisor feedback, GPT\-5\.5 still improves some concerns, but it also introduces substantially more regressions or less auditable changes\. This supports the practical claim thatAutoSupervision\-style verification is not only an offline scoring problem: better concern\-resolution supervisors can provide more useful revision guidance in a closed authoring loop\.

## 6Limitations

#### Data Diversity

The current benchmark is based on Natural Communication transparent peer\-review records, which mainly include papers that successfully completed the revision process\. This may limit the diversity of revision trajectories and overrepresent well\-addressed concerns\. Future extensions could include broader venues and review outcomes to evaluate models under more diverse scenarios\.

#### Task Boundary

AutoSupervisionfocuses on whether models can assess revision resolution from available evidence\. In practice, additional capabilities such as generating revision suggestions, prioritizing reviewer concerns, and interacting with authors or reviewers are also important directions beyond the current benchmark scope\.

## 7Conclusion

We introducedAutoSupervision, a benchmark for scientific manuscript revision verification and concern\-to\-revision grounding\.AutoSupervisionprovides structured reviewer concerns, author responses, revised manuscript evidence, resolution labels, and exact evidence block annotations\. Our evaluation shows that existing LLMs achieve strong characterization ability, but still struggle with verifying whether a revision genuinely addresses a reviewer concern and identifying the supporting manuscript evidence\.

## References

- Using the REVAS review assistant tool in the ARR\-May cycle\.Note:[https://aclrollingreview\.org/revas\-may26](https://aclrollingreview.org/revas-may26)ACL Rolling Review blog postCited by:[§2](https://arxiv.org/html/2607.27845#S2.SS0.SSS0.Px3.p1.1)\.
- Anonymous \(2022\)Transparent peer review for all\.Nature Communications13,pp\. 6173\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1038/s41467-022-33056-8)Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p6.1)\.
- Anthropic \(2026\)Introducing Claude Opus 4\.8\.Note:[https://www\.anthropic\.com/news/claude\-opus\-4\-8](https://www.anthropic.com/news/claude-opus-4-8)Accessed: 2026\-07\-27Cited by:[Table 2](https://arxiv.org/html/2607.27845#S4.T2.1.2.1.1),[§5\.2](https://arxiv.org/html/2607.27845#S5.SS2.SSS0.Px1.p1.1)\.
- L\. Cao, L\. You, and R\. Team \(2025\)CSPaper review: fast, rubric\-faithful conference feedback\.InProceedings of the 18th International Natural Language Generation Conference: System Demonstrations,L\. Flek, S\. Narayan, L\. H\. Phuong, and J\. Pei \(Eds\.\),Hanoi, Vietnam,pp\. 3–7\.External Links:[Link](https://aclanthology.org/2025.inlg-demos.2/)Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p2.1),[§2](https://arxiv.org/html/2607.27845#S2.SS0.SSS0.Px3.p1.1)\.
- L\. Cheng, L\. Bing, Q\. Yu, W\. Lu, and L\. Si \(2020\)APE: argument pair extraction from peer review and rebuttal via multi\-task learning\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),B\. Webber, T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 7000–7011\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.569/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.569)Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p4.1),[§2](https://arxiv.org/html/2607.27845#S2.SS0.SSS0.Px2.p1.1)\.
- N\. Dycke, I\. Kuznetsov, and I\. Gurevych \(2023\)NLPeer: a unified resource for the computational study of peer review\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),A\. Rogers, J\. Boyd\-Graber, and N\. Okazaki \(Eds\.\),Toronto, Canada,pp\. 5049–5073\.External Links:[Link](https://aclanthology.org/2023.acl-long.277/),[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.277)Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p4.1),[§2](https://arxiv.org/html/2607.27845#S2.SS0.SSS0.Px1.p1.1)\.
- A\. E\. Ghareeb, B\. Chang, L\. Mitchener, A\. Yiu, C\. J\. Szostkiewicz, D\. Shved, G\. J\. Gyimesi, J\. M\. Laurent, S\. M\. Wright, M\. T\. Razzak,et al\.\(2026\)A multi\-agent system for automating scientific discovery\.Nature,pp\. 1–3\.Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p2.1),[§2](https://arxiv.org/html/2607.27845#S2.SS0.SSS0.Px3.p1.1)\.
- T\. Ghosal, S\. Kumar, P\. K\. Bharti, and A\. Ekbal \(2022\)Peer review analyze: a novel benchmark resource for computational analysis of peer reviews\.Plos one17\(1\),pp\. e0259238\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1371/journal.pone.0259238)Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p4.1),[§2](https://arxiv.org/html/2607.27845#S2.SS0.SSS0.Px1.p1.1)\.
- Google DeepMind \(2026\)Gemini 3\.1 Pro model card\.Note:[https://storage\.googleapis\.com/deepmind\-media/Model\-Cards/Gemini\-3\-1\-Pro\-Model\-Card\.pdf](https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-1-Pro-Model-Card.pdf)Accessed: 2026\-07\-27Cited by:[Table 2](https://arxiv.org/html/2607.27845#S4.T2.1.5.4.1),[§5\.2](https://arxiv.org/html/2607.27845#S5.SS2.SSS0.Px1.p1.1)\.
- Y\. Guo, G\. Shang, V\. Rennard, M\. Vazirgiannis, and C\. Clavel \(2023\)Automatic analysis of substantiation in scientific peer reviews\.InFindings of the Association for Computational Linguistics: EMNLP 2023,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 10198–10216\.External Links:[Link](https://aclanthology.org/2023.findings-emnlp.684/),[Document](https://dx.doi.org/10.18653/v1/2023.findings-emnlp.684)Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p4.1)\.
- InternLM Team \(2026\)Intern\-S2\-Preview: an efficient 35b scientific multimodal foundation model\.Note:[https://huggingface\.co/internlm/Intern\-S2\-Preview](https://huggingface.co/internlm/Intern-S2-Preview)Model card; accessed: 2026\-07\-27Cited by:[Table 2](https://arxiv.org/html/2607.27845#S4.T2.1.13.12.1),[§5\.2](https://arxiv.org/html/2607.27845#S5.SS2.SSS0.Px2.p1.1)\.
- Y\. Jiang and A\. Ng \(2025\)Stanford agentic reviewer: tech overview\.Note:[https://paperreview\.ai/tech\-overview](https://paperreview.ai/tech-overview)Accessed: 2026\-07\-22Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p2.1),[§2](https://arxiv.org/html/2607.27845#S2.SS0.SSS0.Px3.p1.1)\.
- D\. Kang, W\. Ammar, B\. Dalvi, M\. van Zuylen, S\. Kohlmeier, E\. Hovy, and R\. Schwartz \(2018\)A dataset of peer reviews \(PeerRead\): collection, insights and NLP applications\.InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long Papers\),M\. Walker, H\. Ji, and A\. Stent \(Eds\.\),New Orleans, Louisiana,pp\. 1647–1661\.External Links:[Link](https://aclanthology.org/N18-1149/),[Document](https://dx.doi.org/10.18653/v1/N18-1149)Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p4.1),[§2](https://arxiv.org/html/2607.27845#S2.SS0.SSS0.Px1.p1.1)\.
- N\. N\. Kennard, T\. O’Gorman, R\. Das, A\. Sharma, C\. Bagchi, M\. Clinton, P\. K\. Yelugam, H\. Zamani, and A\. McCallum \(2022\)DISAPERE: a dataset for discourse structure in peer review discussions\.InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,M\. Carpuat, M\. de Marneffe, and I\. V\. Meza Ruiz \(Eds\.\),Seattle, United States,pp\. 1234–1249\.External Links:[Link](https://aclanthology.org/2022.naacl-main.89/),[Document](https://dx.doi.org/10.18653/v1/2022.naacl-main.89)Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p4.1),[§2](https://arxiv.org/html/2607.27845#S2.SS0.SSS0.Px2.p1.1)\.
- G\. R\. Latona, M\. H\. Ribeiro, T\. R\. Davidson, V\. Veselovsky, and R\. West \(2024\)The ai review lottery: widespread ai\-assisted peer reviews boost paper scores and acceptance rates\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.48550/arXiv.2405.02150)Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p2.1)\.
- W\. Liang, Y\. Zhang, H\. Cao, B\. Wang, D\. Y\. Ding, X\. Yang, K\. Vodrahalli, S\. He, D\. S\. Smith, Y\. Yin,et al\.\(2024\)Can large language models provide useful feedback on research papers? a large\-scale empirical analysis\.NEJM AI1\(8\),pp\. AIoa2400196\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.48550/arXiv.2310.01783)Cited by:[§2](https://arxiv.org/html/2607.27845#S2.SS0.SSS0.Px3.p1.1)\.
- J\. Lin, J\. Song, Z\. Zhou, Y\. Chen, and X\. Shi \(2023\)Moprd: a multidisciplinary open peer review dataset\.Neural Computing and Applications35\(34\),pp\. 24191–24206\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1007/s00521-023-08891-5)Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p4.1),[§2](https://arxiv.org/html/2607.27845#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Liu, A\. Mei, B\. Lin, B\. Xue, B\. Wang, B\. Xu, B\. Wu, B\. Zhang, C\. Lin, C\. Dong,et al\.\(2025\)Deepseek\-v3\. 2: pushing the frontier of open large language models\.Cited by:[Table 2](https://arxiv.org/html/2607.27845#S4.T2.1.10.9.1),[§5\.2](https://arxiv.org/html/2607.27845#S5.SS2.SSS0.Px2.p1.1)\.
- R\. Liu and N\. B\. Shah \(2023\)Reviewergpt? an exploratory study on using large language models for paper reviewing\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.48550/arXiv.2306.00622)Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p2.1),[§2](https://arxiv.org/html/2607.27845#S2.SS0.SSS0.Px3.p1.1)\.
- C\. Lu, C\. Lu, R\. T\. Lange, J\. Foerster, J\. Clune, and D\. Ha \(2024\)The ai scientist: towards fully automated open\-ended scientific discovery\.Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p2.1),[§2](https://arxiv.org/html/2607.27845#S2.SS0.SSS0.Px3.p1.1)\.
- MiniMax \(2026\)The MiniMax\-M2 series: mini activations unleashing max real\-world intelligence\.Cited by:[Table 2](https://arxiv.org/html/2607.27845#S4.T2.1.11.10.1),[§5\.2](https://arxiv.org/html/2607.27845#S5.SS2.SSS0.Px2.p1.1)\.
- Moonshot AI \(2026\)Kimi K2\.6 tech blog: advancing open\-source coding\.Note:[https://www\.kimi\.com/blog/kimi\-k2\-6](https://www.kimi.com/blog/kimi-k2-6)Accessed: 2026\-07\-27Cited by:[Table 2](https://arxiv.org/html/2607.27845#S4.T2.1.9.8.1),[§5\.2](https://arxiv.org/html/2607.27845#S5.SS2.SSS0.Px2.p1.1)\.
- A\. Mulligan, L\. Hall, and E\. Raphael \(2013\)Peer review in a changing world: an international study measuring the attitudes of researchers\.Journal of the American Society for Information Science and Technology64\(1\),pp\. 132–161\.Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p1.1)\.
- OpenAI \(2024\)GPT\-4o mini: advancing cost\-efficient intelligence\.Note:[https://openai\.com/index/gpt\-4o\-mini\-advancing\-cost\-efficient\-intelligence/](https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/)Accessed: 2026\-07\-27Cited by:[Table 2](https://arxiv.org/html/2607.27845#S4.T2.1.15.14.1),[Table 2](https://arxiv.org/html/2607.27845#S4.T2.1.16.15.1),[§5\.2](https://arxiv.org/html/2607.27845#S5.SS2.SSS0.Px1.p1.1)\.
- OpenAI \(2026\)GPT\-5\.5\.Note:[https://openai\.com/index/gpt\-5\-5\-system\-card/](https://openai.com/index/gpt-5-5-system-card/)Accessed: 2026\-07\-22Cited by:[§4\.1](https://arxiv.org/html/2607.27845#S4.SS1.SSS0.Px3.p1.1),[Table 2](https://arxiv.org/html/2607.27845#S4.T2.1.3.2.1),[§5\.2](https://arxiv.org/html/2607.27845#S5.SS2.SSS0.Px1.p1.1)\.
- PapeReview \(2026\)PapeReview: agentic ai peer\-review for research manuscripts\.Note:[https://papereview\.com/](https://papereview.com/)Accessed: 2026\-07\-22Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p2.1),[§2](https://arxiv.org/html/2607.27845#S2.SS0.SSS0.Px3.p1.1)\.
- Qwen Team \(2026a\)Qwen3\.5\-9B model card\.Note:[https://huggingface\.co/Qwen/Qwen3\.5\-9B](https://huggingface.co/Qwen/Qwen3.5-9B)Accessed: 2026\-07\-27Cited by:[Table 2](https://arxiv.org/html/2607.27845#S4.T2.1.14.13.1),[Table 2](https://arxiv.org/html/2607.27845#S4.T2.1.7.6.1),[§5\.2](https://arxiv.org/html/2607.27845#S5.SS2.SSS0.Px2.p1.1)\.
- Qwen Team \(2026b\)Qwen3\.7: the agent frontier\.Note:[https://qwen\.ai/blog?id=qwen3\.7](https://qwen.ai/blog?id=qwen3.7)Accessed: 2026\-07\-27Cited by:[Table 2](https://arxiv.org/html/2607.27845#S4.T2.1.6.5.1),[§5\.2](https://arxiv.org/html/2607.27845#S5.SS2.SSS0.Px2.p1.1)\.
- S\. Schmidgall, Y\. Su, Z\. Wang, X\. Sun, J\. Wu, X\. Yu, J\. Liu, M\. Moor, Z\. Liu, and E\. Barsoum \(2025\)Agent laboratory: using LLM agents as research assistants\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 5977–6043\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.320/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.320),ISBN 979\-8\-89176\-335\-7Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p2.1),[§2](https://arxiv.org/html/2607.27845#S2.SS0.SSS0.Px3.p1.1)\.
- C\. Shen, L\. Cheng, R\. Zhou, L\. Bing, Y\. You, and L\. Si \(2022\)MReD: a meta\-review dataset for structure\-controllable text generation\.InFindings of the Association for Computational Linguistics: ACL 2022,S\. Muresan, P\. Nakov, and A\. Villavicencio \(Eds\.\),Dublin, Ireland,pp\. 2521–2535\.External Links:[Link](https://aclanthology.org/2022.findings-acl.198/),[Document](https://dx.doi.org/10.18653/v1/2022.findings-acl.198)Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p4.1),[§2](https://arxiv.org/html/2607.27845#S2.SS0.SSS0.Px2.p1.1)\.
- B\. Wang, C\. Xu, X\. Zhao, L\. Ouyang, F\. Wu, Z\. Zhao, R\. Xu, K\. Liu, Y\. Qu, F\. Shang,et al\.\(2024\)Mineru: an open\-source solution for precise document content extraction\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.48550/arXiv.2409.18839)Cited by:[§4\.1](https://arxiv.org/html/2607.27845#S4.SS1.SSS0.Px2.p1.1)\.
- H\. Wang, T\. Fu, Y\. Du, W\. Gao, K\. Huang, Z\. Liu, P\. Chandak, S\. Liu, P\. Van Katwyk, A\. Deac,et al\.\(2023\)Scientific discovery in the age of artificial intelligence\.Nature620\(7972\),pp\. 47–60\.Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p1.1)\.
- A\. Xu, B\. Lin, B\. Xue, B\. Wang, B\. Xu, B\. Wu, B\. Zhang, C\. Lin, C\. Dong, C\. Ling,et al\.\(2026\)Deepseek\-v4: towards highly efficient million\-token context intelligence\.Cited by:[Table 2](https://arxiv.org/html/2607.27845#S4.T2.1.12.11.1),[Table 2](https://arxiv.org/html/2607.27845#S4.T2.1.8.7.1),[§5\.2](https://arxiv.org/html/2607.27845#S5.SS2.SSS0.Px2.p1.1)\.
- Z\.ai \(2026\)GLM\-5\.2: built for long\-horizon tasks\.Note:[https://z\.ai/blog/glm\-5\.2](https://z.ai/blog/glm-5.2)Accessed: 2026\-07\-27Cited by:[Table 2](https://arxiv.org/html/2607.27845#S4.T2.1.4.3.1),[§5\.2](https://arxiv.org/html/2607.27845#S5.SS2.SSS0.Px2.p1.1)\.
- R\. Zhou, L\. Chen, and K\. Yu \(2024\)Is LLM a reliable reviewer? a comprehensive evaluation of LLM on automatic paper reviewing tasks\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\),N\. Calzolari, M\. Kan, V\. Hoste, A\. Lenci, S\. Sakti, and N\. Xue \(Eds\.\),Torino, Italia,pp\. 9340–9351\.External Links:[Link](https://aclanthology.org/2024.lrec-main.816/)Cited by:[§1](https://arxiv.org/html/2607.27845#S1.p2.1),[§2](https://arxiv.org/html/2607.27845#S2.SS0.SSS0.Px3.p1.1)\.

## Appendix APrediction Schema and Evaluation Details

Given a revision episodeet=\(Mt,Rt\)e\_\{t\}=\(M\_\{t\},R\_\{t\}\), the model produces one structured prediction for each review\-response recordrt,ir\_\{t,i\}\. Each prediction consists of three groups of fields:*Characterization*,*Verification*, and*Grounding*\. Table[5](https://arxiv.org/html/2607.27845#A1.T5)summarizes all scored fields and their valid options\.

#### Characterization\.

Characterization evaluates whether a model correctly identifies the nature and scope of a reviewer comment\. It includes three aspects:comment\_kind, which describes whether a comment is a critique, question, or suggestion;point\_type, which identifies the underlying issue category; and the target scope of the comment\. We define*target scope*as the joint specification oftarget\_domain,target\_section, andtarget\_object\. Accordingly,SscopeS\_\{\\mathrm\{scope\}\}denotes the hierarchical target\-scope score computed from these three fields\.

#### Verification\.

Verification evaluates whether a model can determine whether a reviewer concern has been addressed in the revised manuscript\. It consists ofpaper\_status, which describes the observed degree of change in the revised manuscript;resolution\_label, which indicates whether the concern is resolved, partially resolved, or unresolved; andresolution\_evidence\_block\_ids, which identifies the evidence blocks supporting the resolution judgment\.

#### Grounding\.

Grounding evaluates whether a model can identify the exact revised\-manuscript evidence supporting where a reviewer concern is addressed\. It is measured usingtarget\_paper\_block\_ids, which specifies the set of revised\-manuscript evidence blocks where the concern is addressed or grounded in the paper\.

#### Component scores\.

Categorical fields are evaluated using macro\-F1 to reduce the influence of label\-frequency imbalance\. For ontology labels, macro\-F1 is computed over labels with positive gold support\. Concerns with theunverifiableresolution label are excluded according to the benchmark evaluation protocol\.

Predicted evidence block ID sets are evaluated using micro\-F1 because block identifiers are document\-specific references rather than shared categorical labels\.

The target scope is evaluated usingSscopeS\_\{\\mathrm\{scope\}\}, a hierarchical score that assigns partial credit according to the granularity of the matched target:

Sscope=1N​∑i=1N\(0\.2​Id\+0\.4​Is\+0\.4​Io\),S\_\{\\mathrm\{scope\}\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\left\(0\.2I\_\{d\}\+0\.4I\_\{s\}\+0\.4I\_\{o\}\\right\),
whereIdI\_\{d\},IsI\_\{s\}, andIoI\_\{o\}indicate exact matches of the target domain, target section, and target object, respectively\.

The three component scores are:

Cchar=13​\(Fkind\+Ftype\+Sscope\),C\_\{\\mathrm\{char\}\}=\\frac\{1\}\{3\}\\left\(F\_\{\\mathrm\{kind\}\}\+F\_\{\\mathrm\{type\}\}\+S\_\{\\mathrm\{scope\}\}\\right\),
Cver=13​\(Fres\+Fstatus\+Fevidence\),C\_\{\\mathrm\{ver\}\}=\\frac\{1\}\{3\}\\left\(F\_\{\\mathrm\{res\}\}\+F\_\{\\mathrm\{status\}\}\+F\_\{\\mathrm\{evidence\}\}\\right\),
Cground=Fground\.C\_\{\\mathrm\{ground\}\}=F\_\{\\mathrm\{ground\}\}\.
Here,FkindF\_\{\\mathrm\{kind\}\}andFtypeF\_\{\\mathrm\{type\}\}denote the macro\-F1 scores forcomment\_kindandpoint\_type, respectively\.FresF\_\{\\mathrm\{res\}\}andFstatusF\_\{\\mathrm\{status\}\}denote the macro\-F1 scores forresolution\_labelandpaper\_status\.FevidenceF\_\{\\mathrm\{evidence\}\}denotes the micro\-F1 score forresolution\_evidence\_block\_ids, measuring whether the model identifies evidence blocks that support its resolution judgment\. Finally,FgroundF\_\{\\mathrm\{ground\}\}denotes the micro\-F1 score fortarget\_paper\_block\_ids, measuring whether the model identifies the exact revised\-manuscript evidence blocks where the concern is addressed\.

The overall benchmark score is the average of the three component scores:

Coverall=13​\(Cchar\+Cver\+Cground\)\.C\_\{\\mathrm\{overall\}\}=\\frac\{1\}\{3\}\\left\(C\_\{\\mathrm\{char\}\}\+C\_\{\\mathrm\{ver\}\}\+C\_\{\\mathrm\{ground\}\}\\right\)\.
Table 5:Scored prediction fields and options\. Target scope refers jointly to target domain, target section, and target object\. Manuscript grounding refers to exact revised\-manuscript evidence blocks supporting where a concern is addressed\.

## Appendix BDetails of Dataset

### B\.1Concern Attribute Distributions

Tables[6](https://arxiv.org/html/2607.27845#A2.T6)and[7](https://arxiv.org/html/2607.27845#A2.T7)summarize the concern\-level composition of the final evaluation subset\. As expected for accepted papers, most concerns are ultimately resolved, but the dataset still retains partially resolved and unresolved cases that require models to distinguish complete revision from incomplete response\. The target\-object distribution shows that revision verification is not a prose\-only task: 1,185 concerns target figures, and additional concerns target data or analysis blocks, references, tables, and equations\. The point\-type distribution further shows that the benchmark spans a broad taxonomy of scientific revision needs, including writing and organization, claim\-evidence support, statistical reporting, mechanistic justification, reproducibility details, figure/table presentation, experimental adequacy, methodology, scope, related work, and dataset or evaluation protocol concerns\.

Table 6:Gold\-label distributions in the full benchmark\.FieldLabelCountcomment\_kindcritique4,419comment\_kindsuggestion2,843comment\_kindquestion1,528resolution\_labelresolved8,265resolution\_labelpartially resolved434resolution\_labelunresolved91target\_objecttext6,759target\_objectfigure1,231target\_objectdata365target\_objectreference276target\_objecttable78target\_objectequation81Table 7:Point\-type distribution in the full benchmark\.
### B\.2Agreement Experiment

Because the gold construction uses LLM\-assisted structured extraction, we add an agreement check on 100 randomly sampled papers from the benchmark test split\. The concern set is fixed to the GPT\-5\.5 output before running the second model\. It covers 160 revision episodes and 1,600 fixed concern instances\. Table[8](https://arxiv.org/html/2607.27845#A2.T8)summarizes agreement between GPT\-5\.5 and Claude Opus 4\.8\.

Table 8:Agreement on 100 sampled benchmark papers\. Claude Opus 4\.8 re\-annotates fixed GPT\-5\.5 concern ids using the same structured schema\.The two models agree strongly on the benchmark construction\. Target section reaches 0\.873 exact agreement and 0\.796 Cohen’sκ\\kappa, concern taxonomy reaches 0\.603 exact agreement and 0\.540κ\\kappa, and the resolution label reaches 0\.848 exact agreement\. For target manuscript evidence, the teachers achieve 0\.710 block micro\-F1 and 0\.673 mean IoU\. Resolution evidence is harder because it can include several adjacent rebuttal, follow\-up, and confirmation blocks in the review record; across all such blocks, agreement reaches 0\.441 micro\-F1 and 0\.308 mean IoU\.

## Appendix CAgentic Baseline Details

The agentic baseline uses the same base model, GPT\-4o\-mini, for all LLM calls and produces the same JSON prediction schema used by the single\-shot baselines\.

#### Concern characterization\.

The first stage reads the fixed review\-response records for an episode and predicts concern\-level characterization fields, including comment kind, point type, target domain, target section, target object, and response status\. It also generates short search phrases used by the downstream manuscript retriever\.

#### Deterministic manuscript\-block retrieval\.

For each concern, it builds a query from the reviewer concern, author response, current\-round discussion, the preliminary characterization labels, and the generated search phrases\. It then scores every revised\-manuscript block using lexical overlap, object cues such as figure/table/equation/reference mentions, and section/object consistency cues\. The top 12 ranked blocks are retained as seed candidates, and a one\-block neighbor window is added around each selected block to preserve local context\. This top\-12 setting is a recall\-oriented candidate budget: these blocks are not the final predicted evidence, but the candidate pool from which the grounding stage selects final manuscript evidence IDs\.

#### Grounding\.

The grounding stage receives each concern, its preliminary characterization, and the retrieved candidate manuscript blocks\. It selects the target manuscript block IDs, predicts paper\-side change status, and records whether manuscript grounding is available\.

#### Verification\.

The verification stage receives the concern, author response, preliminary characterization, preliminary grounding, and selected manuscript evidence\. It predicts response status, paper status, resolution label, and review\-record evidence block IDs\. If the grounding stage selects no manuscript blocks, the verification stage receives the top retrieved candidates as fallback context\.

#### Normalization and cost\.

Finally, deterministic normalization merges the characterization, grounding, and verification outputs into one prediction per concern and enforces the benchmark schema\. Grounding and verification are run in chunks of four concerns\. The agentic pipeline improves structured\-output reliability but is substantially more expensive than single\-shot GPT\-4o\-mini, using 91\.9M total tokens compared with 20\.0M tokens for the single\-shot baseline\.

## Appendix DSupervised Fine\-tuning Details

#### Data preparation\.

The supervised model is trained from Qwen3\.5\-9B using the same structured prediction format as the main benchmark\. The training data for SFT are generated from the dataset training split\. We used 809 training episodes from 500 papers with 9,211 concern instances for SFT\.

#### Training\.

Training was conducted on two NVIDIA H100 GPUs using full\-parameter fine\-tuning\. The configuration uses the qwen3\-5 model type, DeepSpeed ZeRO\-3, bfloat16 precision, a maximum sequence length of 32,768, AdamW optimization, a learning rate of \(1×10−51\\times 10^\{\-5\}\), cosine learning\-rate scheduling, a warmup ratio of 0\.03, weight decay of 0\.1, gradient checkpointing, a per\-device batch size of 1, gradient accumulation of 2, and a random seed of 42\.

#### Inference and evaluation\.

The checkpoint is served with an OpenAI\-compatible vLLM endpoint, tensor parallel size 8, and maximum model length 32,768\.

## Appendix EReproducibility Statement

In addition to the details we provide in Section[4](https://arxiv.org/html/2607.27845#S4)and Appendix[D](https://arxiv.org/html/2607.27845#A4), we also publicly release the benchmark dataset and code needed to reproduce the reported results\.

## Appendix FEthics Statement

AutoSupervisionis built from publicly available transparent peer\-review records and published articles\. These records were released by the source journal as part of its transparent review process, and our benchmark uses them only to study concern\-level revision verification and evidence grounding\.

## Appendix GLLM Usage Statement

LLMs were used to improve the grammar, wording, and readability of the manuscript\. We also accessed commercial LLMs through their APIs as part of the benchmark evaluation, with a total API cost of approximately USD 2,000\. The authors are fully responsible for the project design, experiments, analysis, citations, and final scientific claims\.

相似文章

方法是否支持声明?同行评审中的论文内验证

arXiv cs.CL

本文介绍了论文内声明验证框架,该框架利用大型语言模型评估论文中的创新声明是否得到方法学证据的支持,填补了现有自动化同行评审系统的空白。人工评估表明,该框架与人类评审员的关注点显著一致,特别是在创新相关问题上。