Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers
Summary
This paper introduces SciSlopBench, a benchmark of 390 AI-generated scientific papers detecting 'scientific slop' via structure, argument, and artifact measures (85.9% accuracy vs. 68.7% for Binoculars), and proposes SciSlopHarness, an evidentiary-grounded revision framework that reduces the AI–human gap by 63% without reward hacking.
View Cached Full Text
Cached at: 10/02/26, 09:46 AM
# Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers
Source: [https://arxiv.org/html/2610.00531](https://arxiv.org/html/2610.00531)
Yerim Oh1Young\-Jun Lee2Jaewoo Ahn1Gunhee Kim1Dongyeop Kang21Seoul National University2University of Minnesotayerim\.oh@vision\.snu\.ac\.krgunhee@snu\.ac\.krdongyeop@umn\.edu![[Uncaptioned image]](https://arxiv.org/html/2610.00531v1/figures/main_figures/emoji_globe.png)Website:[https://yerimoh\.github\.io/scientific\-slop\-demo/](https://yerimoh.github.io/scientific-slop-demo/)
###### Abstract
AI\-generated content, often called AI slop, is increasingly common everywhere, particularly in academia\. Slop in AI\-generated scientific papers, however, has more complex patterns that cannot be easily detected by existing token\-based AI detectors\. Each part of such a paper looks plausible while the scientific reasoning that connects the parts breaks down, which can mislead how readers assess the work\. We benchmark these failures asscientific slopthrough six measures across*Structure*,*Argument*, and*Artifacts*\. We constructSciSlopBenchwith 390 AI\-generated papers, mostly in computer science but spanning the life, social, and natural sciences, each paired with a human\-written paper matched by research problem and contribution type\. Our measures identify the AI paper in each pair with 85\.9% accuracy, compared with 68\.7% for Binoculars\. Higher scientific slop accompanies lower ICLR ratings and distinguishes rejected from accepted papers above chance in every year from 2017 to 2025\. Reducing these patterns, however, is not as simple as directly optimizing the measures\. We therefore proposeSciSlopHarness, a harness\-level framework that guides a fixed LLM to revise slop only where the experiment records support the change\. While standard revisions leave residual slop and direct slop\-aware prompting triggers reward hacking,SciSlopHarnessreduces the remaining AI–human gap by 63% over the strongest revision baseline without requiring human reference targets\. Overall, we demonstrate that AI\-generated scientific papers leave fundamental traces in their global reasoning, and that responsible mitigation demands strict evidentiary grounding rather than mere prose refinement\.
## 1Introduction
As LLM\-generated content fills news feeds, books, and social media, the term*AI slop*has come to name this low\-cost, high\-volume output\([Thorp, 2026](https://arxiv.org/html/2610.00531#bib.bib46)\)\. AI output that looks finished but lacks substance shifts the verification burden to the recipient, forcing them to review and revise the content, and raises concerns of misuse in journalism, education, workplaces, and academia\([Wang et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib47)\)\. AI slop is increasingly common in academia, in both papers and peer reviews\([Liang et al\., 2025](https://arxiv.org/html/2610.00531#bib.bib28);[Kobak et al\., 2025](https://arxiv.org/html/2610.00531#bib.bib18);[Liang et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib27)\)\. Paper submission platforms such as arXiv have revised their policies to restrict survey papers\([arXiv, 2025](https://arxiv.org/html/2610.00531#bib.bib3)\), and conference organizers have begun screening submissions using AI detectors\. Yet, a scientific paper is more than its*prose*\.
These concerns have made detecting AI\-generated content increasingly important\. AI agents are now used to draft and refine scientific papers, from researcher\-guided workflows\([Schmidgall et al\., 2025](https://arxiv.org/html/2610.00531#bib.bib40)\)to systems such as AI Scientist\([Lu et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib30);[Yamada et al\., 2025](https://arxiv.org/html/2610.00531#bib.bib51)\)and FARS\([Tang et al\., 2026](https://arxiv.org/html/2610.00531#bib.bib43)\)\. Such assistance can reduce writing time while also increasing the need for verification, as generated scientific text can appear credible even when the underlying data are fabricated\([Kacena et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib15);[Gao et al\., 2023](https://arxiv.org/html/2610.00531#bib.bib10)\)\. Existing detection methods address these concerns by identifying AI\-generated content through learned textual features or token probabilities\([Emi & Spero, 2024](https://arxiv.org/html/2610.00531#bib.bib9);[Mitchell et al\., 2023](https://arxiv.org/html/2610.00531#bib.bib35);[Hans et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib11);[Ma et al\., 2026](https://arxiv.org/html/2610.00531#bib.bib31)\)\. For scientific papers, however, token\-level probabilities alone are insufficient\.
Challenges\.We definescientific slopas recurring breakdowns in how scientific reasoning connects across a paper \(Figure[1](https://arxiv.org/html/2610.00531#S1.F1)\)\. Each part of a paper can look plausible in isolation, so such breakdowns are invisible to token\-level detectors and can only be identified or repaired at the level of the whole paper\. To benchmark and mitigate scientific slop, this paper addresses three challenges\.First, token\-level metrics fail to capture how scientific reasoning connects across a paper\.Sections, claims, citations, evidence, and artifacts can each appear plausible while the relationships among them break down\. We repeatedly observe such failures in end\-to\-end AI\-generated papers, and ICLR reviewers already penalize them even in human\-written submissions\.Second, detection of these patterns remains unmeasured\.Existing test sets label only the text, so the extent to which detectors, including LLMs that read the entire paper, identify these patterns has never been measured\.Third, these patterns are difficult to mitigate reliably\.Whereas token\-level signals can be removed by paraphrasing\([Krishna et al\., 2023](https://arxiv.org/html/2610.00531#bib.bib20)\), repairing these patterns requires restoring the missing relations without changing the underlying science\.
To systematically study scientific slop in AI\-generated papers, we constructSciSlopBench, the first benchmark to evaluate detection in full scientific papers rather than standalone prose\. It pairs 390 AI\-generated papers from FARS and Agents4Science 2025\([Bianchi et al\., 2026](https://arxiv.org/html/2610.00531#bib.bib5)\)with human\-written papers on the same research problem and of the same contribution type, mostly in computer science but also spanning the life, social, and natural sciences\. Every human\-written counterpart was accepted at a top\-tier venue, so a system that distinguishes the two must look beyond topic, genre, and prose quality\. To mitigate the identified scientific slop, we introduceSciSlopHarness, which translates slop diagnoses into iterative revision guidance for an LLM\. In each round, it checks edits against the original manuscript and experiment records, then uses review feedback to repair remaining gaps in scientific reasoning\.
Figure 1:Scientific slop captures failures beyond token\-level AI detection\. While traditional detectors assess local text, scientific slop analysis examines how scientific reasoning connects, revealing paper\-level failures that plausible prose can conceal\.Contributions\.Our results formalize scientific slop as paper\-level breakdowns in scientific reasoning and develop measures to quantify it\. These measures reduce detection error by 55% compared with Binoculars\([Hans et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib11)\), the strongest token\-level baseline\. Moreover, measures of scientific slop align with human evaluations: AI\-likeness drops from 26% to 12% as review scores rise from 2 to 9\. These measures also distinguish rejected from accepted papers in every ICLR year from 2017 to 2025\. For mitigation, general revision leaves slop unresolved, while direct slop\-aware revision lowers the scores without fixing the reasoning\.SciSlopHarness, which repairs scientific reasoning against grounding evidence, instead moves all six measures toward human levels without using human targets\. These results show that AI\-generated papers leave measurable traces in how scientific reasoning connects across a paper, and that grounding revision in scientific records enables these traces to guide scientific repair\. Our contributions are:
- •We define and measure scientific slop through six types of breakdown in scientific reasoning acrossSTRUCTURE,ARGUMENT, andARTIFACTS\. Individual measures achieve a pairwise accuracy of up to 0\.905 in distinguishing AI\-generated from human\-written papers\.
- •We constructSciSlopBenchto evaluate AI detection based on scientific reasoning across a whole paper rather than standalone prose\. The strongest token\-level detector, Binoculars, achieves a pairwise accuracy of only 0\.687, while full\-text AI reviewers achieve at most 0\.685\.
- •We introduceSciSlopHarness, which verifies repairs against scientific records rather than slop scores\. It reduces the remaining AI–human gap by 63% over the strongest revision baseline, while general revision leaves scientific slop and direct slop\-aware revision overcorrects\.
## 2Related Work
Detecting machine generated text\.Detection methods read local text spans\. DetectGPT measures probability curvature\([Mitchell et al\., 2023](https://arxiv.org/html/2610.00531#bib.bib35)\), Binoculars compares model perplexity\([Hans et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib11)\), NTS measures sensitivity to sampling temperature\([Ma et al\., 2026](https://arxiv.org/html/2610.00531#bib.bib31)\), and Pangram trains classifiers on generated text\([Emi & Spero, 2024](https://arxiv.org/html/2610.00531#bib.bib9)\)\. Watermarking and paraphrasing studies also work on surface signals\([Kirchenbauer et al\., 2023](https://arxiv.org/html/2610.00531#bib.bib17);[Krishna et al\., 2023](https://arxiv.org/html/2610.00531#bib.bib20)\)\. Unlike these tools, our measures read relations across a whole paper rather than local text spans\.
Automated review and revision\.Recent systems generate critiques or revise drafts directly\. The AI Scientist evaluates its own papers with a simulated review process\([Lu et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib30)\), CycleReviewer learns from real peer reviews\([Weng et al\., 2025](https://arxiv.org/html/2610.00531#bib.bib49)\), and OpenReviewer and DeepReview treat review generation as a primary task\([Idahl & Ahmadi, 2025](https://arxiv.org/html/2610.00531#bib.bib13);[Zhu et al\., 2025](https://arxiv.org/html/2610.00531#bib.bib54)\)\. Self\-Refine closes the loop by having one model critique and revise its own output\([Madaan et al\., 2023](https://arxiv.org/html/2610.00531#bib.bib32)\)\. RefineBench scores refinement against checklists and finds that models improve little without feedback\([Lee et al\., 2026c](https://arxiv.org/html/2610.00531#bib.bib25)\)\. REPAIR grounds each refinement round in documents verified through scientific APIs\([Oh & Kim, 2026](https://arxiv.org/html/2610.00531#bib.bib37)\)\. However, these systems offer general feedback without isolating specific recurring patterns\. We use the AI Scientist reviewer and CycleReviewer as detection baselines and a reviewer\-based refinement loop as a revision baseline, and ask whether general review and revision removes the flaws we define\.
Measurement under active removal\.Recent work improves an LLM system by changing the information, checks, and tools around a fixed model rather than the model itself\([Lee et al\., 2026a](https://arxiv.org/html/2610.00531#bib.bib23)\)\. Metanapplies one fixed improvement operation recursively to its own outputs\([Kim et al\., 2026](https://arxiv.org/html/2610.00531#bib.bib16)\), and Evolution Fine\-Tuning trains the model on evolutionary search trajectories\([Lee et al\., 2026b](https://arxiv.org/html/2610.00531#bib.bib24)\)\.SciSlopHarnessfollows this design, keeping the editor fixed and adding located findings before the edit and a record\-grounded review after it\. Mutation testing\([DeMillo et al\., 1978](https://arxiv.org/html/2610.00531#bib.bib7);[Jia & Harman, 2010](https://arxiv.org/html/2610.00531#bib.bib14)\)and adversarial stylometry\([Brennan et al\., 2012](https://arxiv.org/html/2610.00531#bib.bib6);[Potthast et al\., 2016](https://arxiv.org/html/2610.00531#bib.bib38)\)judge a property by trying to remove it, and our slop\-aware baseline plays that role\.
## 3Defining and Measuring scientific slop
Table 1:Measurement specification of SciSlop patterns\. Each score is the share of units that show the pattern, written as numerator / denominator \(Eq\.[1](https://arxiv.org/html/2610.00531#S3.E1)\)\. The last column states what a score of 1 means\. Full rules are in Appendix[A](https://arxiv.org/html/2610.00531#A1)\.### 3\.1Scientific Slop: Definition and Measures
[Kommers et al\. \(2026\)](https://arxiv.org/html/2610.00531#bib.bib19)identify superficial competence as a characteristic of AI slop, whose apparent quality conceals a lack of substance\. In scientific papers, we examine this gap in the connections needed to organize the study, substantiate its claims, and make its methods and evidence inspectable\. We definescientific slopas recurring failures in these connections that hinder readers from following the paper’s scientific reasoning\. We group these failures intoSTRUCTURE,ARGUMENT, andARTIFACTS\. Table[1](https://arxiv.org/html/2610.00531#S3.T1)specifies six observable patterns and their counting rules\.
STRUCTURE
concerns how sections contribute to the paper’s scientific argument\.*Cross\-section references*targets missing connections between them\. Rhetorical structure theory explains how parts of a text are connected, while endophoric references make these connections visible to readers\([Mann & Thompson, 1988](https://arxiv.org/html/2610.00531#bib.bib33);[Hyland, 2023](https://arxiv.org/html/2610.00531#bib.bib12)\)\. We therefore measure sections and labeled objects that no other section refers to\.*Macro redundancy*captures when later sections repeat earlier material instead of developing the scientific argument\. Methods, results, and conclusions serve different argumentative roles\([Teufel et al\., 2009](https://arxiv.org/html/2610.00531#bib.bib45)\), so each section should add to rather than restate the argument\. We measure such reuse with cross\-sectionnn\-gram overlap\([Welleck et al\., 2020](https://arxiv.org/html/2610.00531#bib.bib48)\)\.
ARGUMENT
evaluates how a paper justifies its claims and establishes its position among prior studies\.*Argument graph*targets breaks in the reasoning that connects a paper’s claims to their support\. Argument mining represents this reasoning through relations between claims and premises\([Stab & Gurevych, 2017](https://arxiv.org/html/2610.00531#bib.bib42)\)\. We measure how often a claim appears before its supporting context\.*Citation isolation*captures prior work that is cited without explaining its role in the paper’s argument\. Citations can show what a paper builds on, differs from, or compares against\([Teufel et al\., 2006](https://arxiv.org/html/2610.00531#bib.bib44)\)\. We measure citations that provide none of these relations\.
Figure 2:SciSlopBenchpipeline\.We pair each AI\-generated paper with a human\-written paper matched by research problem and contribution type\. We measure six scientific slop patterns in both papers, then compare the mean slop scores of each matched pair\.ARTIFACTS
examines whether readers can inspect how a method works and what its reported results represent\.*Figure exposition*captures method diagrams that obscure how the method works\. Diagrams support inference by showing related information together\([Larkin & Simon, 1987](https://arxiv.org/html/2610.00531#bib.bib22)\), but unnecessary explanations can interfere with this role\([Mayer & Moreno, 2003](https://arxiv.org/html/2610.00531#bib.bib34)\)\. We flag experimental details or claims about results that make the method’s steps and connections harder to follow\.*Evidence gap*occurs when a paper reports aggregate results but provides no concrete examples\. Without these examples, readers cannot inspect individual successes and failures to better understand model behavior\([Lipton & Steinhardt, 2019](https://arxiv.org/html/2610.00531#bib.bib29)\)\. We check for this gap in papers with result tables in the main text, looking for inputs, outputs, or cases throughout the paper and appendix\.
For paperppand itemjj, letUj\(p\)U\_\{j\}\(p\)be the units of analysis andhj\(u\)h\_\{j\}\(u\)indicate whether unituumeets the counting criterion in Table[1](https://arxiv.org/html/2610.00531#S3.T1)\. The item score is
sj\(p\)=1\|Uj\(p\)\|∑u∈Uj\(p\)hj\(u\)\.s\_\{j\}\(p\)=\\frac\{1\}\{\|U\_\{j\}\(p\)\|\}\\sum\_\{u\\in U\_\{j\}\(p\)\}h\_\{j\}\(u\)\.\(1\)
### 3\.2SciSlopBench
Existing AI text detection benchmarks\([Wang et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib47);[Li et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib26);[Dugan et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib8);[Wu et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib50)\)typically evaluate standalone prose, leaving relations among a paper’s sections, tables, and code untested\. To study scientific slop systematically, we introduceSciSlopBench, a benchmark for evaluating how well systems distinguish AI\-generated papers from matched human\-authored papers using whole\-paper evidence\. As shown in Fig\.[2](https://arxiv.org/html/2610.00531#S3.F2),SciSlopBenchpairs AI\-generated papers with human\-written papers matched by research problem and contribution type, gives each paper one slop score from the six patterns, and checks whether the human\-written paper scores lower than its AI\-generated counterpart\.
#### 3\.2\.1Benchmark Construction
Each instance inSciSlopBenchpairs an AI\-generated paper with a human\-authored paper matched by paper type and topic\. We evaluate whether these scores rank the human\-authored paper above its AI\-generated counterpart\. The reference labels reflect how the papers were produced\.
##### Dataset construction\.
We draw AI\-generated papers from FARS and Agents4Science 2025 submissions\. FARS takes a research idea and produces a finished paper, running the experiments and writing the manuscript itself\. For each FARS paper, we identify cited papers of the same paper type that meet our human\-authorship screening criteria and have main\-track acceptance at a top\-tier venue\. We then select the closest topical match, yielding 143 pairs matched on both the kind of contribution and the research topic\. For Agents4Science, we also include papers available as PDFs, which we convert into a form suitable for slop measurement\. Data collection, matching, and document\-processing details are provided in Appendix[B](https://arxiv.org/html/2610.00531#A2)\.
##### Benchmark statistics\.
SciSlopBenchcombines 143 matched pairs from FARS with 247 pairs from Agents4Science 2025 submissions, spanning computer and data sciences, life and health sciences, social sciences, natural sciences, and other fields\. We report slop measures as normalized rates to account for differences in paper length and the number of measured units\. Source\-specific composition, acceptance counts, and detailed topic distributions are provided in Appendix[B](https://arxiv.org/html/2610.00531#A2)\.
##### Evaluation metrics\.
We use PairAcc and AUROC as our main evaluation metrics\. PairAcc measures how often a system ranks the human\-authored paper above its AI\-generated counterpart, with score directions fixed in advance and ties receiving half credit\. AUROC measures how well the scores distinguish the two classes across classification thresholds\. We also report the true\-positive rate at a 5% false\-positive rate \(TPR\), the share of AI papers flagged when the threshold lets through 5% of human papers\. For the ICLR analysis in Section[5\.2](https://arxiv.org/html/2610.00531#S5.SS2), each system’s score is turned into an AI probability by a logistic fit on the FARS pairs \(Appendix[E](https://arxiv.org/html/2610.00531#A5)\)\.
Figure 3:SciSlopHarnesspipeline\.We locate scientific slop instances in the manuscript and use them to guide localized revision\. We then review each proposed change against the original manuscript, experiment outputs, and code, retain only supported changes, and repeat the process on the remaining slop\.
## 4SciSlopHarness: Mitigating Scientific Slop
This section describesSciSlopHarness, our method for reducing scientific slop by repairing the connections in scientific reasoning\. As shown in Fig\.[3](https://arxiv.org/html/2610.00531#S3.F3),SciSlopHarnesslocates scientific slop, uses the detected instances to guide manuscript revision, and reviews each proposed change against the original manuscript and the paper’s experiment records before retaining it\. The remaining slop and review feedback then guide the next revision round\.
##### Task\.
Given an original manuscriptM0M\_\{0\}and its experiment records, the result files its code produced,SciSlopHarnessrevises the manuscript for up to three rounds\. Following harness\-level optimization\([Lee et al\., 2026a](https://arxiv.org/html/2610.00531#bib.bib23)\), we keep the revision model fixed and change only the information and checks around it\. At roundrr, the revision model receivesMr−1M\_\{r\-1\}withSciSlop\.md, which describes how each type of scientific slop should be revised\. It also receivesSLOP\_FINDINGS\.md, which lists the detected problems and their locations\. The revision model edits the manuscript to produce a draftM~r\\widetilde\{M\}\_\{r\}\. A review model then compares each change withMr−1M\_\{r\-1\},M0M\_\{0\}, experiment outputs, and the code\. Changes consistent with the reported results, existing evidence, and cited work are kept inMrM\_\{r\}\. The remaining changes are reverted to their version inMr−1M\_\{r\-1\}, so each round proceeds asMr−1⟶revisionM~r⟶reviewMrM\_\{r\-1\}\\overset\{\\mathrm\{revision\}\}\{\\longrightarrow\}\\widetilde\{M\}\_\{r\}\\overset\{\\mathrm\{review\}\}\{\\longrightarrow\}M\_\{r\}\. Appendix[C](https://arxiv.org/html/2610.00531#A3)provides implementation details\.
### 4\.1Localized Revision
We first locate where a connection in scientific reasoning breaks down so that the revision model knows what part of the manuscript to revise\. For each detected problem, the model receives its location and guidance on how to revise it\. The locations come from measurements of the manuscriptMr−1M\_\{r\-1\}and are stored inSLOP\_FINDINGS\.md\. The revision guidance is fixed across manuscripts and stored inSciSlop\.md\. The model uses both to edit theLaTeXsource and produceM~r\\widetilde\{M\}\_\{r\}\.
We use these findings to guide revision, not to require every scientific slop measure to decrease\. The model can leave a finding unchanged when the suggested connection is not needed in the paper\. For example, it adds a cross\-section reference only when another section actually uses the referenced object\. We therefore provide the findings and revision guidance without a target score\.
### 4\.2Reviewing Proposed Revisions
Localized guidance tells the revision model what to change, but it cannot verify whether the new content is supported by the research\. We therefore review each proposed change before keeping it in the manuscript\. For each changed paragraph inM~r\\widetilde\{M\}\_\{r\}, the review model checks the original manuscript, experiment outputs, and code for the claims, numbers, citations, and examples affected by the revision\. It keeps a change only when these sources support the revised content\. Otherwise, the paragraph is restored to its version inMr−1M\_\{r\-1\}\. The remaining changes formMrM\_\{r\}\.
The review checks that revisions preserve the paper’s reported results, citations, and available evidence\. A changed result is reverted if it conflicts with the experiment outputs\. A new example is kept only when the same example exists in the available records\. Changes that remove citations or add unnecessary references are also reverted\.
We repeat revision and review for up to three rounds\. We stop when no problems remain for revision or when review produces no further changes\. Before returning the final manuscript, we check that reported numbers and cited works are preserved and that citations and references remain valid\.
Table 2:Paper discrimination onSciSlopBench\. TPR is at 5% FPR, and bold and underline mark the best two\.
Table 3:Scientific slop measures by plane\. TPR is at 5% FPR, a dagger marks a measure that needs a model, and the last row is the score of Table[3](https://arxiv.org/html/2610.00531#S4.T3)\.
## 5Experiments
### 5\.1Experimental Setup
##### Baselines\.
All existing methods assign a single score to each paper\. We evaluate Binoculars\([Hans et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib11)\), DetectGPT\([Mitchell et al\., 2023](https://arxiv.org/html/2610.00531#bib.bib35)\)and NTS\([Ma et al\., 2026](https://arxiv.org/html/2610.00531#bib.bib31)\)as AI text detectors, and use the overall ratings from the AI Scientist reviewer\([Lu et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib30)\)and CycleReviewer\([Weng et al\., 2025](https://arxiv.org/html/2610.00531#bib.bib49)\)as automated review baselines\. For revision, we compare four baselines that share the editor model and round protocol\. Base prompting, Claude Code, and reviewer\-based refinement receive only a generic improvement instruction, and we refer to these three as general revision\. Slop\-aware revision instead receives explicit slop definitions and detected locations, using the same editor and tool configuration as reviewer\-based refinement\. We compare these baselines with ourSciSlopHarness, using the negative of the aggregate item score so that higher values favor human authorship, as for the other methods\. Details of the revision baselines are provided in Appendix[D](https://arxiv.org/html/2610.00531#A4), and the distances behind the reported gap reduction in Appendix[C\.5](https://arxiv.org/html/2610.00531#A3.SS5)\.
##### Implementation details\.
Text detectors read the prose view\. Other methods read the paper source, with code access provided where needed\. Binoculars uses the Falcon\-7B\([Almazrouei et al\., 2023](https://arxiv.org/html/2610.00531#bib.bib1)\)base and instruct models, while DetectGPT uses t5\-3b\([Raffel et al\., 2020](https://arxiv.org/html/2610.00531#bib.bib39)\)to generate 100 perturbations\. For inputs exceeding Falcon’s 2048\-token context limit, we score windows without overlap and average their scores\. The AI Scientist reviewer uses Qwen2\.5\-32B\-Instruct\([Yang et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib52)\), and CycleReviewer uses its released 8B checkpoint\.
### 5\.2Does Scientific Slop Mark Provenance, Quality, or Both?
We test whether scientific slop captures overlooked weaknesses already reflected in lower human review scores\. We first assess how well slop measures identify AI\-generated papers, then examine whether the same patterns matter to human peer review\.
##### Scientific slop identifies AI\-generated papers by what detectors and reviewers never read\.
Tables[3](https://arxiv.org/html/2610.00531#S4.T3)and[3](https://arxiv.org/html/2610.00531#S4.T3)show that scientific slop identifies AI\-generated papers more accurately than the listed detectors and automated reviewers that read the full paper\. The best of each reaches 0\.687 and 0\.685, while scientific slop reaches 0\.859, cutting the error by more than half\.*Cross\-section references*alone reaches 0\.905 without a model and, at 5% FPR, flags 65% of AI\-generated papers against 24% for Binoculars\. These results show that measuring patterns across a paper provides an effective signal for AI detection on this corpus\. By defining these patterns as specific defects, scientific slop further connects detection to concrete targets for revision\.
Figure 4:AI likeness against ICLR review scores \(a\), assessment dimensions \(b\), and annual acceptance decisions \(c\)\. Panel \(a\) adjusts probabilities for differences between years, panel \(b\) shows\|ρ\|\|\\rho\|withblueandredfor significant negative and positive correlations \(p<0\.05p<0\.05\) and gray for estimates that are not significant, and panel \(c\) uses rejection AUROC with0\.50\.5indicating chance\.\(a\)Review scores\(b\)Review dimensions\(c\)Acceptance decisions
##### Scientific slop reads human\-written papers the way their reviewers do\.
Figure[4](https://arxiv.org/html/2610.00531#S5.F4)\(a\) compares each system’s AI probability with human review ratings\. Scientific slop shows a clearerrelationship with review scoresthan the evaluated detectors and automated reviewers\. Its AI probability declines more sharply as ratings rise, indicating that the score also captures differences in reviewer\-assessed quality\. Figure[4](https://arxiv.org/html/2610.00531#S5.F4)\(b\) examines individual review dimensions\. Only scientific slop consistently assigns higher scores to papers rated lower across overall rating, soundness, presentation, and contribution, with statistical support in all four dimensions\. This links scientific slop tosoundness and contributionas well as presentation\. Figure[4](https://arxiv.org/html/2610.00531#S5.F4)\(c\) compares the separation of rejected and accepted papers across years\. Scientific slop exceeds chance in all nine years, whereas detectors fall below chance in most comparisons\. The relationship thus extends toacceptance decisionsacross changing review scales\. Together, these findings suggest that the weaknesses we define as scientific slop are already reflected in lower human evaluations\. Scientific slop makes these recurring weaknesses explicit and measurable\. Measurement details are provided in Appendix[E](https://arxiv.org/html/2610.00531#A5)\.
### 5\.3DoesSciSlopHarnessRemove the Slop That General Revision Leaves?
We evaluate whetherSciSlopHarnessremoves scientific slop that survives general revision without receiving human anchors as revision targets\. We runSciSlopHarnessand the four revision baselines on every AI\-generated paper inSciSlopBench, revising each paper for three rounds under the shared protocol of Appendix[D\.1](https://arxiv.org/html/2610.00531#A4.SS1)\. We first compare it with general revision, then examine whether repairs guided by connections in scientific reasoning approach the reference levels observed in human papers\.
Figure 5:Scientific slop across revision rounds forSciSlopHarnessand baselines\. Values closer to zero \(green\) are closer to matched human means, with round 0 denoting original manuscripts and the scale below zero compressed\.##### SciSlopHarnessreduces the AI–human gap that general revision leaves and slop\-aware revision overcorrects\.
Figure[5](https://arxiv.org/html/2610.00531#S5.F5)shows thatSciSlopHarnessends closest to the human mean on all six measures, despite never receiving these means as revision targets\. For*Macro redundancy*,*Cross\-section references*, and*Citation isolation*, general revision leaves substantial slop, whereas slop\-aware revision pushes scores past the human means and increasingly away from them, suggesting overcorrection\.SciSlopHarnessreduces the defects left by general revision while remaining closer to human levels\. Its gains also extend to*Evidence gap*,*Argument graph*, and*Figure exposition*, where the baselines make little or no progress\. The harness evaluates each defect in the context of the paper’s scientific reasoning\. The skill guides edits that connect claims with supporting evidence where needed, while review checks whether each edit contributes to the argument and preserves existing information\. Together, they address defects left by feedback alone while limiting overcorrection\. Summed over the six measures, the remaining gap to the human means is 63% smaller than for Claude Code, the closest baseline \(Appendix[C\.5](https://arxiv.org/html/2610.00531#A3.SS5)\)\. The approach toward human levels thus supports scientific reasoning as a criterion for deciding what to repair and what to preserve, without specifying a target score\.
### 5\.4Is the Slop Removed, or Only the Score?
We test whether reductions in scientific slop reflect valid repairs rather than lower scores alone\. We compare the revisions ofSciSlopHarnesswith slop\-aware revision to see what produces these reductions, and ablate its components\.
Table 4:The same passage before revision, after slop\-aware revision, and afterSciSlopHarness, markingthe located slop,what a revision lost or left, andwhat it kept or repaired\. Brackets replace citation and reference commands\.##### SciSlopHarnessremoves the slop where slop\-aware revision removes only the score\.
Table[4](https://arxiv.org/html/2610.00531#S5.T4)shows how slop\-aware revision andSciSlopHarnessrepair the same detected problems differently\. For Macro redundancy, slop\-aware revision reaches a zero score by removing the repeated sentence together with four citations, losing the attribution to prior work\. In contrast,SciSlopHarnessremoves only the repetition and preserves all four citations, reducing redundancy without losing the information needed for the argument\. For Cross\-section references, slop\-aware revision approaches zero by adding references to three sections, three tables, and two figures throughout the conclusion\. These references improve the score, but most only tell the reader where to look without connecting the evidence to a claim\. In contrast,SciSlopHarnessadds one reference where the threshold table supports the robustness claim, and states what the table shows, that everyτ≥32\\tau\\geq 32beats the greedy baseline\.SciSlopHarnessreduces the same slop by preserving necessary information and repairing the missing connections between claims and evidence\. We observe the same distinction for the remaining measures in Appendix[F](https://arxiv.org/html/2610.00531#A6)\.
##### The fullSciSlopHarnessis needed for reliable slop reduction\.
Figure[6](https://arxiv.org/html/2610.00531#S5.F6)shows that removing any part of the harness introduces a different failure\. Definitions and locations reduce the distance to 0\.11, but without review this comes with guard breaks in 23% of rounds\.
Figure 6:Ablation ofSciSlopHarness\.Review alone leaves most slop unrepaired at 0\.30 and still breaks a guard in 17% of rounds\. Definitions with review are safer, but without located instances the distance remains 0\.25\. In contrast, the fullSciSlopHarnessreduces the distance to 0\.17 with no guard breaks\. The ablations show why the components must work together: the harness must identify where to repair, guide the repair, and verify that the resulting edit preserves the scientific record\.
## 6Discussion
##### Conclusion\.
We introducescientific slopas measurable breakdowns in the reasoning that connects a paper’s parts\. Its measures distinguish AI\-generated from human papers and are associated with lower scores across all four ICLR review dimensions\.SciSlopHarnessuses the identified defects to guide revisions checked against the paper’s own experiment outputs and code\. It brings all six measures closest to human levels among the evaluated revision methods, without receiving human reference targets\. Scientific slop connects detection to weaknesses reflected in human reviews and to repairs supported by scientific records\.
##### Limitations and future directions\.
The six scientific slops offer an initial view of how connections within a paper reveal defects and guide repair\. Future work can extend this view to how a study’s question, methods, and evidence support its conclusions as a whole\. Our evaluation focuses on FARS and Agents4Science 2025, with most papers drawn from computer science\. Extending the paired evaluation to new generators and disciplines can test how broadly this view applies\.
## AI Use Statement
LLMs are used in this work both as the subject of study and as components of our experiments\. SciSlopBench includes third\-party AI\-generated papers, our measures use Qwen2\.5 models, and SciSlopHarness and its baselines use Claude models as editors and reviewers\. We also used Claude \(including Claude Code\) and ChatGPT during implementation and manuscript preparation\. They assisted with writing and debugging code for the measurement pipeline, harness, baselines, and analysis scripts, as well as editing the manuscript for grammar, clarity, and concision\. In some cases, they were used to produce initial text that was subsequently edited by the authors\. The authors reviewed and tested all assisted code and reviewed the final manuscript\. All reported results were obtained by running the released code\.
We take responsibility for the final content of this work, including text, claims or artifacts produced with the aid of generative AI\.
## Ethics statement
This work aims to support scientific scrutiny by making failures in reasoning identifiable and open to inspection\. Scientific slop measures give authors and reviewers concrete points to examine in a paper, including weaknesses that fluent prose can obscure\. These findings concern the paper’s reasoning and do not establish an author’s intent or misconduct\.
Correcting a reasoning failure remains valuable even when it weakens a signal of AI authorship\. The criterion for responsible repair is whether the study’s records support the change\.SciSlopHarnessapplies this criterion by checking revisions against the manuscript and experiment records and recording preservation checks for citations and reported numbers\. When no concrete example exists in the records, the revision acknowledges that absence\. This makes the basis for a revision available for inspection alongside its effect on the score\.
Our ICLR analysis uses public manuscripts and review information and reports aggregate relationships between measured slop and review outcomes\. The benchmark’s data sources and processing procedures are documented to support scrutiny of its construction\.
#### Acknowledgments
This work was supported by Institute of Information & communications Technology Planning & Evaluation \(IITP\) grant funded by the Korea government\(MSIT\) \(No\.RS\-2026\-25524173, Ultro\-Long\-Term Hierarchical Memory and Reasoning Architecture for Next\-Generation Omnimodal Agents, the Institute of Information & Communications Technology Planning & Evaluation\(IITP\) grant funded by the Korea government\(MSIT\) \(RS\-2025\-25442338, AI star Fellowship Support Program\(Seoul National Univ\.\)\), the IITP\(Institute of Information & Communications Technology Planning & Evaluation\)\-ITRC\(Information Technology Research Center\) grant funded by the Korea government\(Ministry of Science and ICT\)\(IITP\-2025\-RS\-2024\-00437633\), Basic Science Research Program through the National Research Foundation of Korea\(NRF\) funded by the Ministry of Education\(RS\-2023\-00274280\), National IT Industry Promotion Agency\(NIPA\) grant funded by the Korea government\(MSIT\), and the AI Seoul Tech Research Support Program of the Seoul Future Foundation\. Gunhee Kim and Dongyeop Kang are the corresponding author\.
## References
- Almazrouei et al\. \(2023\)Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo\.The Falcon series of open language models, 2023\.URL[https://arxiv\.org/abs/2311\.16867](https://arxiv.org/abs/2311.16867)\.
- Anthropic \(2025\)Anthropic\.Claude code\.[https://claude\.com/claude\-code](https://claude.com/claude-code), 2025\.
- arXiv \(2025\)arXiv\.Attention authors: Updated practice for review articles and position papers in arXiv CS category\.[https://blog\.arxiv\.org/2025/10/31/attention\-authors\-updated\-practice\-for\-review\-articles\-and\-position\-papers\-in\-arxiv\-cs\-category/](https://blog.arxiv.org/2025/10/31/attention-authors-updated-practice-for-review-articles-and-position-papers-in-arxiv-cs-category/), October 2025\.Accessed 2026\-09\-21\.
- Bai et al\. \(2025\)Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin\.Qwen2\.5\-VL Technical Report, 2025\.URL[https://arxiv\.org/abs/2502\.13923](https://arxiv.org/abs/2502.13923)\.
- Bianchi et al\. \(2026\)Federico Bianchi, Owen Queen, Nitya Thakkar, Eric Sun, and James Zou\.Exploring the use of AI authors and reviewers at Agents4Science\.*Nature Biotechnology*, 44\(1\):11–14, 2026\.
- Brennan et al\. \(2012\)Michael Brennan, Sadia Afroz, and Rachel Greenstadt\.Adversarial stylometry: Circumventing authorship recognition to preserve privacy and anonymity\.*ACM Transactions on Information and System Security \(TISSEC\)*, 15\(3\):1–22, 2012\.
- DeMillo et al\. \(1978\)Richard A DeMillo, Richard J Lipton, and Frederick G Sayward\.Hints on test data selection: Help for the practicing programmer\.*Computer*, 11\(4\):34–41, 1978\.
- Dugan et al\. \(2024\)Liam Dugan, Alyssa Hwang, Filip Trhlík, Andrew Zhu, Josh Magnus Ludan, Hainiu Xu, Daphne Ippolito, and Chris Callison\-Burch\.Raid: A shared benchmark for robust evaluation of machine\-generated text detectors\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 12463–12492, 2024\.
- Emi & Spero \(2024\)Bradley Emi and Max Spero\.Technical report on the pangram ai\-generated text classifier, 2024\.
- Gao et al\. \(2023\)Catherine A Gao, Frederick M Howard, Nikolay S Markov, Emma C Dyer, Siddhi Ramesh, Yuan Luo, and Alexander T Pearson\.Comparing scientific abstracts generated by chatgpt to real abstracts with detectors and blinded human reviewers\.*NPJ digital medicine*, 6\(1\):75, 2023\.
- Hans et al\. \(2024\)Abhimanyu Hans, Avi Schwarzschild, Valeriia Cherepanova, Hamid Kazemi, Aniruddha Saha, Micah Goldblum, Jonas Geiping, and Tom Goldstein\.Spotting llms with binoculars: zero\-shot detection of machine\-generated text\.In*Proceedings of the 41st International Conference on Machine Learning*, ICML’24\. JMLR\.org, 2024\.
- Hyland \(2023\)Ken Hyland\.*Metadiscourse*\.Routledge, 2023\.
- Idahl & Ahmadi \(2025\)Maximilian Idahl and Zahra Ahmadi\.Openreviewer: A specialized large language model for generating critical scientific paper reviews\.In*Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(System Demonstrations\)*, pp\. 550–562, 2025\.
- Jia & Harman \(2010\)Yue Jia and Mark Harman\.An analysis and survey of the development of mutation testing\.*IEEE transactions on software engineering*, 37\(5\):649–678, 2010\.
- Kacena et al\. \(2024\)Melissa A Kacena, Lilian I Plotkin, and Jill C Fehrenbacher\.The use of artificial intelligence in writing scientific review articles\.*Current Osteoporosis Reports*, 22\(1\):115–121, 2024\.
- Kim et al\. \(2026\)Zae Myung Kim, Young\-Jun Lee, Seungyeon Jwa, and Dongyeop Kang\.Metan: Recursive self\-improvement through emergent depth\.In*Advances in Neural Information Processing Systems*, 2026\.
- Kirchenbauer et al\. \(2023\)John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein\.A watermark for large language models\.In*International conference on machine learning*, pp\. 17061–17084\. PMLR, 2023\.
- Kobak et al\. \(2025\)Dmitry Kobak, Rita González\-Márquez, Emőke\-Ágnes Horvát, and Jan Lause\.Delving into llm\-assisted writing in biomedical publications through excess vocabulary\.*Science Advances*, 11\(27\):eadt3813, 2025\.
- Kommers et al\. \(2026\)Cody Kommers, Eamon Duede, Julia Gordon, Ari Holtzman, Tess McNulty, Spencer Stewart, Lindsay Thomas, Richard Jean So, and Hoyt Long\.Why slop matters\.*ACM AI Lett\.*, 1\(1\), March 2026\.doi:10\.1145/3786777\.URL[https://doi\.org/10\.1145/3786777](https://doi.org/10.1145/3786777)\.
- Krishna et al\. \(2023\)Kalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting, and Mohit Iyyer\.Paraphrasing evades detectors of AI\-generated text, but retrieval is an effective defense\.In*Advances in Neural Information Processing Systems*, volume 36, pp\. 27469–27500, 2023\.
- Kwon et al\. \(2023\)Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E\. Gonzalez, Hao Zhang, and Ion Stoica\.Efficient memory management for large language model serving with PagedAttention\.In*Proceedings of the 29th Symposium on Operating Systems Principles*, 2023\.
- Larkin & Simon \(1987\)Jill H\. Larkin and Herbert A\. Simon\.Why a diagram is \(sometimes\) worth ten thousand words\.*Cogn\. Sci\.*, 11\(1\):65–100, 1987\.doi:10\.1111/J\.1551\-6708\.1987\.TB00863\.X\.URL[https://doi\.org/10\.1111/j\.1551\-6708\.1987\.tb00863\.x](https://doi.org/10.1111/j.1551-6708.1987.tb00863.x)\.
- Lee et al\. \(2026a\)Yoonho Lee, Roshen Nair, Qizheng Zhang, Kangwook Lee, Omar Khattab, and Chelsea Finn\.Meta\-harness: End\-to\-end optimization of model harnesses\.*arXiv preprint arXiv:2603\.28052*, 2026a\.
- Lee et al\. \(2026b\)Young\-Jun Lee, Seungone Kim, Minki Kang, Alistair Cheong Liang Chuen, Zerui Chen, Seungho Han, Taehee Jung, and Dongyeop Kang\.Evolution fine\-tuning: Learning to discover across 371 optimization tasks\.In*Advances in Neural Information Processing Systems*, 2026b\.
- Lee et al\. \(2026c\)Young\-Jun Lee, Seungone Kim, Byung\-Kwan Lee, Minkyeong Moon, Yechan Hwang, Jong Myoung Kim, Graham Neubig, Sean Welleck, and Ho\-Jin Choi\.RefineBench: Evaluating refinement capability of language models via checklists\.In*International Conference on Learning Representations*, 2026c\.
- Li et al\. \(2024\)Yafu Li, Qintong Li, Leyang Cui, Zhilin Wang, Longyue Wang, Linyi Yang, Shuming Shi, Yue Zhang, et al\.Mage: Machine\-generated text detection in the wild\.In*Proceedings of the 62nd annual meeting of the association for computational linguistics \(Volume 1: Long Papers\)*, pp\. 36–53, 2024\.
- Liang et al\. \(2024\)Weixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp, Hancheng Cao, Xuandong Zhao, Lingjiao Chen, Haotian Ye, Sheng Liu, Zhi Huang, et al\.Monitoring ai\-modified content at scale: A case study on the impact of chatgpt on ai conference peer reviews\.*arXiv preprint arXiv:2403\.07183*, 2024\.
- Liang et al\. \(2025\)Weixin Liang, Yaohui Zhang, Zhengxuan Wu, Haley Lepp, Wenlong Ji, Xuandong Zhao, Hancheng Cao, Sheng Liu, Siyu He, Zhi Huang, et al\.Quantifying large language model usage in scientific papers\.*Nature Human Behaviour*, 9\(12\):2599–2609, 2025\.
- Lipton & Steinhardt \(2019\)Zachary C\. Lipton and Jacob Steinhardt\.Troubling trends in machine learning scholarship, 2019\.URL[https://doi\.org/10\.1145/3317287\.3328534](https://doi.org/10.1145/3317287.3328534)\.
- Lu et al\. \(2024\)Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha\.The ai scientist: Towards fully automated open\-ended scientific discovery, 2024\.
- Ma et al\. \(2026\)Shixuan Ma, Jiahao Li, Zhendong Mao, and Quan Wang\.Zero\-shot detection of LLM\-generated text using temperature sensitivity\.In*Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 37664–37679, 2026\.
- Madaan et al\. \(2023\)Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al\.Self\-refine: Iterative refinement with self\-feedback\.In*Advances in Neural Information Processing Systems*, volume 36, pp\. 46534–46594, 2023\.
- Mann & Thompson \(1988\)William Mann and Sandra Thompson\.Rhetorical structure theory: Toward a functional theory of text organization\.*Text \- Interdisciplinary Journal for the Study of Discourse*, 8:243–281, 01 1988\.doi:10\.1515/text\.1\.1988\.8\.3\.243\.
- Mayer & Moreno \(2003\)Richard E\. Mayer and Roxana Moreno\.Nine ways to reduce cognitive load in multimedia learning\.*Educational Psychologist*, 38\(1\):43–52, 2003\.doi:10\.1207/S15326985EP3801\\\_6\.
- Mitchell et al\. \(2023\)Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D Manning, and Chelsea Finn\.Detectgpt: Zero\-shot machine\-generated text detection using probability curvature\.In*International conference on machine learning*, pp\. 24950–24962\. PMLR, 2023\.
- Mosca et al\. \(2023\)Edoardo Mosca, Mohamed Hesham Ibrahim Abdalla, Paolo Basso, Margherita Musumeci, and Georg Groh\.Distinguishing fact from fiction: A benchmark dataset for identifying machine\-generated scientific papers in the LLM era\.In*Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing*, 2023\.
- Oh & Kim \(2026\)Yerim Oh and Gunhee Kim\.REPAIR: Resolving long\-tail confusion in scientific retrievers via fact\-verified iterative refinement\.In*Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing*, 2026\.
- Potthast et al\. \(2016\)Martin Potthast, Matthias Hagen, and Benno Stein\.Author obfuscation: Attacking the state of the art in authorship verification\.*CLEF \(Working Notes\)*, pp\. 716–749, 2016\.
- Raffel et al\. \(2020\)Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J\. Liu\.Exploring the limits of transfer learning with a unified text\-to\-text transformer\.*Journal of Machine Learning Research*, 21\(140\):1–67, 2020\.URL[https://jmlr\.org/papers/v21/20\-074\.html](https://jmlr.org/papers/v21/20-074.html)\.
- Schmidgall et al\. \(2025\)Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Michael Moor, Zicheng Liu, and Emad Barsoum\.Agent laboratory: Using llm agents as research assistants\.*Findings of the Association for Computational Linguistics: EMNLP 2025*, pp\. 5977–6043, 2025\.
- Singh et al\. \(2023\)Amanpreet Singh, Mike D’Arcy, Arman Cohan, Doug Downey, and Sergey Feldman\.SciRepEval: A multi\-format benchmark for scientific document representations\.In*Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing*, 2023\.
- Stab & Gurevych \(2017\)Christian Stab and Iryna Gurevych\.Parsing argumentation structures in persuasive essays\.*Comput\. Linguistics*, 43\(3\):619–659, 2017\.doi:10\.1162/COLI\\\_A\\\_00295\.URL[https://doi\.org/10\.1162/COLI\_a\_00295](https://doi.org/10.1162/COLI_a_00295)\.
- Tang et al\. \(2026\)Qiong Tang, Tianxiang Sun, Xiangkun Hu, Xiangyang Liu, Yiran Chen, Yunfan Shao, Bobo Li, Changze Lv, Cheng Xu, Chengsong Huang, et al\.Fars: A fully automated research system deployed at scale\.*arXiv preprint arXiv:2606\.31651*, 2026\.
- Teufel et al\. \(2006\)Simone Teufel, Advaith Siddharthan, and Dan Tidhar\.Automatic classification of citation function\.In Dan Jurafsky and Éric Gaussier \(eds\.\),*EMNLP 2006, Proceedings of the 2006 Conference on Empirical Methods in Natural Language Processing, 22\-23 July 2006, Sydney, Australia*, pp\. 103–110\. ACL, 2006\.URL[https://aclanthology\.org/W06\-1613/](https://aclanthology.org/W06-1613/)\.
- Teufel et al\. \(2009\)Simone Teufel, Advaith Siddharthan, and Colin Batchelor\.Towards domain\-independent argumentative zoning: Evidence from chemistry and computational linguistics\.In*Proceedings of the 2009 conference on empirical methods in natural language processing*, pp\. 1493–1502, 2009\.
- Thorp \(2026\)H Holden Thorp\.Resisting AI slop\.*Science*, 391\(6780\):5–5, 2026\.
- Wang et al\. \(2024\)Yuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su, Artem Shelmanov, Akim Tsvigun, Chenxi Whitehouse, Osama Mohammed Afzal, Tarek Mahmoud, Toru Sasaki, et al\.M4: Multi\-generator, multi\-domain, and multi\-lingual black\-box machine\-generated text detection\.In*Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 1369–1407, 2024\.
- Welleck et al\. \(2020\)Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston\.Neural text generation with unlikelihood training\.In*8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26\-30, 2020*\. OpenReview\.net, 2020\.URL[https://openreview\.net/forum?id=SJeYe0NtvH](https://openreview.net/forum?id=SJeYe0NtvH)\.
- Weng et al\. \(2025\)Yixuan Weng, Minjun Zhu, Guangsheng Bao, Hongbo Zhang, Jindong Wang, Yue Zhang, and Linyi Yang\.Cycleresearcher: Improving automated research via automated review\.In*International Conference on Learning Representations*, volume 2025, pp\. 3669–3709, 2025\.
- Wu et al\. \(2024\)Junchao Wu, Runzhe Zhan, Derek F Wong, Shu Yang, Xinyi Yang, Yulin Yuan, and Lidia S Chao\.Detectrl: Benchmarking llm\-generated text detection in real\-world scenarios\.*Advances in Neural Information Processing Systems*, 37:100369–100401, 2024\.
- Yamada et al\. \(2025\)Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, and David Ha\.The ai scientist\-v2: Workshop\-level automated scientific discovery via agentic tree search\.*arXiv preprint arXiv:2504\.08066*, 2025\.
- Yang et al\. \(2024\)An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tianyi Tang, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu\.Qwen2\.5 Technical Report, 2024\.URL[https://arxiv\.org/abs/2412\.15115](https://arxiv.org/abs/2412.15115)\.
- Yu et al\. \(2023\)Peipeng Yu, Jiahan Chen, Xuan Feng, and Zhihua Xia\.CHEAT: A large\-scale dataset for detecting ChatGPT\-written abstracts\.*arXiv preprint arXiv:2304\.12008*, 2023\.
- Zhu et al\. \(2025\)Minjun Zhu, Yixuan Weng, Linyi Yang, and Yue Zhang\.Deepreview: Improving llm\-based paper review with human\-like deep thinking process\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 29330–29355, 2025\.
## Appendix AMeasurement Specification of the Slop Measures
This appendix provides the detailed measurement procedure for each measure in Table[1](https://arxiv.org/html/2610.00531#S3.T1)\.
##### Document views\.
We expand all`\\input`files and read each paper as a singleLaTeXdocument\. We define the paper body as the text before the appendix, acknowledgements, or bibliography, whichever comes first, and apply the same rule to both papers in each pair\. We construct two views of this body\. The*structure view*, used by the six scientific slop measures, preserves sections, labels, references, citations, and floats\. The*prose view*, used by the text detectors, removes floats and captions and replaces mathematics, citations, and references with sentinel tokens\. For human papers, we construct these views directly from theLaTeXe\-print\. Among the 247 Agents4Science submissions, 18 provide manuscript source\. For the remaining 229, we reconstruct the sameLaTeXstructure from the PDF, recovering headings, float labels, citations, and references\.
##### Macro redundancy\.
A sentence is recycled when more than half of its tokens sit inside 8\-grams that an earlier section already used\. Over the prose view of the abstract and the body sections, for a sentenceuuof at least eight tokens, withC\(u\)C\(u\)the tokens ofuuthat lie in an 8\-gram first seen in an earlier sentence of a different section,
h\(u\)=\[\|C\(u\)\|/\|u\|≥0\.5\]\.h\(u\)=\\mathbf\{1\}\\\!\\left\[\\,\|C\(u\)\|/\|u\|\\geq 0\.5\\,\\right\]\.\(2\)
##### Cross\-section references\.
Every body section, and every figure, table, equation, algorithm, or theorem labeled inside one, is an object the paper declares\.R\(o\)R\(o\)collects the sections other than its home that point tooothrough a reference command, a printed section number, or a named reference, and
h\(o\)=\[R\(o\)=∅\]\.h\(o\)=\\mathbf\{1\}\\\!\\left\[\\,R\(o\)=\\emptyset\\,\\right\]\.\(3\)Pointing at a table from the section that holds it is neither rewarded nor punished, since only other sections count\. Pointers into the appendix, and a roadmap in the first section, the longest run of forward pointers to later sections, are recorded apart because neither is a trace of one section having read another\.
##### Argument graph\.
Does the context that supports a claim come before the claim? A language model labels each sentence of the introduction, by a forced choice over the numbered introduction with the majority of three greedy runs, and the key claims are the sentences labeled superiority, prior limitation, or design choice\. The supporting context ofsis\_\{i\}is the sentence that raises its likelihood most under a second model that is never prompted,
pmi\(j→i\)=1\|si\|−1∑t=2\|si\|\[logP\(sit∣sj,si<t\)−logP\(sit∣si<t\)\],π\(i\)=argmaxj≠ipmi\(j→i\),\\begin\{gathered\}\\mathrm\{pmi\}\(j\\\!\\to\\\!i\)=\\frac\{1\}\{\|s\_\{i\}\|\-1\}\\sum\_\{t=2\}^\{\|s\_\{i\}\|\}\\Big\[\\log P\\\!\\left\(s\_\{i\}^\{t\}\\mid s\_\{j\},\\,s\_\{i\}^\{<t\}\\right\)\-\\log P\\\!\\left\(s\_\{i\}^\{t\}\\mid s\_\{i\}^\{<t\}\\right\)\\Big\],\\\\ \\pi\(i\)=\\arg\\max\_\{j\\neq i\}\\,\\mathrm\{pmi\}\(j\\\!\\to\\\!i\),\\end\{gathered\}\(4\)and the claim is marked when that context comes after it,
h\(i\)=\[π\(i\)\>i\]\.h\(i\)=\\mathbf\{1\}\\\!\\left\[\\,\\pi\(i\)\>i\\,\\right\]\.\(5\)
##### Citation isolation\.
A citing sentence weaves when it cites two works, names a second work outside a citation command, or links two works by a cue from a frozen list, and it is isolated otherwise,
h\(u\)=\[udoes not weave\]\.h\(u\)=\\mathbf\{1\}\\\!\\left\[\\,u\\ \\text\{does not weave\}\\,\\right\]\.\(6\)The score runs over the citing sentences of the introduction and the related work\. Relating a work only to the present paper is not weaving, and fewer than eight citing sentences flags the score weak\.
##### Figure exposition\.
The method figure is read as text\. One vision model transcribes it, with the same prompt on both sides, and six kinds of exposition are checked on the transcript, a notation key, a comparison arm, experimental content, an evaluative mark, a thesis box, and enumerated stages, withh\(k\)=1h\(k\)=1when the figure carries kindkk\. Five kinds are frozen patterns\. Experimental content, whether a number is a setting of the run or a constant of the method, is decided by one reader, on 33 clear and 25 borderline figures of the FARS half and 13 and 9 of the Agents4Science half\. On the FARS half the method figure is the pipeline’s overview image, and on the Agents4Science half both sides are located by the same caption gate\.
##### Evidence gap\.
Here the unit is the paper\. It is applicable when its body holds a result table with at least two data rows and four numeric cells, andh\(p\)=1h\(p\)=1when no exhibit appears anywhere in it, appendix included\. An exhibit is a display environment carrying consumed or produced material, a caption announcing an example, a case, or a failure, or a long quotation outside the related work, and one exhibit anywhere closes the gap\.
##### Aggregate score\.
The SciSlop score of a paper averages the item scores of Eq\.[1](https://arxiv.org/html/2610.00531#S3.E1)within each plane and then averages the planes\. With𝒥k\(p\)\\mathcal\{J\}\_\{k\}\(p\)the items of planekkthat apply to paperppand𝒦\(p\)\\mathcal\{K\}\(p\)the planes with at least one such item,
S\(p\)=1\|𝒦\(p\)\|∑k∈𝒦\(p\)1\|𝒥k\(p\)\|∑j∈𝒥k\(p\)sj\(p\)\.S\(p\)=\\frac\{1\}\{\|\\mathcal\{K\}\(p\)\|\}\\sum\_\{k\\in\\mathcal\{K\}\(p\)\}\\frac\{1\}\{\|\\mathcal\{J\}\_\{k\}\(p\)\|\}\\sum\_\{j\\in\\mathcal\{J\}\_\{k\}\(p\)\}s\_\{j\}\(p\)\.\(7\)This is the SciSlop row of Tables[3](https://arxiv.org/html/2610.00531#S4.T3)and[3](https://arxiv.org/html/2610.00531#S4.T3); the negative ofS\(p\)S\(p\)is used wherever a higher score should favor human authorship\.
##### Models and serving\.
*Argument graph*labels sentences with Qwen2\.5\-32B\-Instruct\([Yang et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib52)\)and reads probabilities from Qwen2\.5\-7B\-Instruct, and*Figure exposition*transcribes with Qwen2\.5\-VL\-32B\-Instruct\([Bai et al\., 2025](https://arxiv.org/html/2610.00531#bib.bib4)\)\. Models run on vLLM\([Kwon et al\., 2023](https://arxiv.org/html/2610.00531#bib.bib21)\)in bfloat16 with greedy decoding, a label or scalar is the majority or median of three runs, every call is cached under a hash of prompt and model, and no measurement calls a paid or closed model\.
## Appendix BSciSlopBenchConstruction Details
SciSlopBenchholds 390 pairs\. Each pair is one AI\-generated paper and one human\-written paper, its anchor, of the same kind of contribution on the same or a neighboring problem, and every anchor is a paper whose human authorship and peer acceptance can be shown\. The AI papers come from two sources, 143 written end to end by FARS\([Tang et al\., 2026](https://arxiv.org/html/2610.00531#bib.bib43)\)and the 247 submissions to Agents4Science 2025, and the same pairing rule is applied to both\. Tables[3](https://arxiv.org/html/2610.00531#S4.T3)and[3](https://arxiv.org/html/2610.00531#S4.T3)pool the 390 pairs into one evaluation\. Appendix[B\.1](https://arxiv.org/html/2610.00531#A2.SS1)gives the pairing rule, Appendix[B\.2](https://arxiv.org/html/2610.00531#A2.SS2)the FARS half, Appendix[B\.3](https://arxiv.org/html/2610.00531#A2.SS3)the Agents4Science half and how its PDFs are read, and Appendix[B\.4](https://arxiv.org/html/2610.00531#A2.SS4)the relation to existing detection benchmarks\.
### B\.1Pairing rule
A human paper can anchor an AI paper only if it passes four conditions\.
- H1Human authorship\.Its first arXiv version is dated 2024 or earlier, or it is an ICLR 2026 submission whose Pangram AI share is at most 0\.05\. The date stands in for a check we cannot run on older papers, and Pangram is the only check available for papers written after 2024\.
- H2Acceptance\.It was accepted to the main track of a top\-tier venue, shown by the venue record, the arXiv comment or journal reference, or OpenReview\. Rejected submissions, workshop papers, and preprints do not pass\.
- H3Contribution type\.It is the same kind of paper as the AI paper, one of method, framework, benchmark, dataset, analysis, or survey\.
- H4Usable source\.Its arXiv e\-print expands in document order to a body of at least 2,500 words in at least four sections, with the appendix separable from the body\.
Within each contribution type, we select the human paper closest to the AI paper in research problem\. A language model assigns topic similarity from the title and abstract on four tiers: 3 for the same problem and contribution type, 2 for the same problem, 1 for the same area, and 0 for papers cited only as tools, models, or datasets\. We select the highest tier and break ties by experiment\-section citation, availability of public review scores, a human\-to\-AI body\-length ratio within 2\.5, and finally SPECTER2 cosine similarity\([Singh et al\., 2023](https://arxiv.org/html/2610.00531#bib.bib41)\)\. We use cosine only as the final tie\-break because candidate similarities are tightly clustered between 0\.90 and 0\.97\.
We search two candidate pools in order\. We first consider papers cited by the AI paper, since these are the works against which the paper positions its contribution\. If no cited paper satisfies the pairing conditions, we search 2,690 accepted ICLR 2026 papers with an arXiv identifier and Pangram share at most 0\.05, retrieve the three nearest by TF\-IDF over title and abstract, and apply the same criteria\. Assignment is one\-to\-one within each source, with contested anchors assigned to the closer match\. We retain the candidate pool, similarity tier, condition outcomes, and alternative candidates for each pair so that every match can be independently checked\.
### B\.2FARS: collection and matching
##### Source\.
FARS contains 166 papers generated end to end from a research idea, including experiments and manuscript writing\. Of these, 165 provideLaTeXsource and 160 provide code\. Most are method papers \(123\), followed by analysis \(38\), framework \(4\), and benchmark \(1\)\.
##### Matching\.
Of the 166 FARS papers, 161 cite at least one arXiv paper, yielding 1,946 candidates from 1,197 unique papers\. Applying H1–H4 leaves candidates for 124 papers\. H3 removes the most because many citations are benchmarks or base models rather than papers of the same contribution type\. One\-to\-one assignment then yields 121 pairs from the cited pool, resolving 16 shared\-anchor conflicts in favor of the closer match\. For the remaining 45 papers we search the widened ICLR 2026 pool, which pairs 23 more; the other 22 \(13 analysis, 7 method, 2 framework\) have no same\-type anchor in either pool and are not paired\. One of the 144 pairs is dropped because its FARS paper ships noLaTeXsource, leaving the 143 pairs of the benchmark, 120 from the cited pool and 23 from the widened pool \(Table[5](https://arxiv.org/html/2610.00531#A2.T5)\)\.
Table 5:Retrieval audit of the assignment with SPECTER2 similarity between the AI paper and the human anchor\. Rank is the position of the assigned anchor among all 143 anchors for that AI paper, and the reference row gives the mean cosine over every non\-assigned pair\.
### B\.3Agents4Science: collection, matching, and document processing
##### Collection\.
Agents4Science 2025 provides 247 primarily AI\-authored submissions, and we include all of them\. Most are in computer and data sciences \(178\), particularly AI and machine learning \(154\)\. By contribution type, 171 are method papers, 46 analysis, 19 framework, 6 survey, and 5 benchmark\. Acceptance decisions, reviewer scores, and reported AI involvement are retained as metadata but are not used for pairing\. Tables[7](https://arxiv.org/html/2610.00531#A2.T7)and[7](https://arxiv.org/html/2610.00531#A2.T7)report the full topic distribution\.
Table 6:Primary topics and acceptance status of the 247 Agents4Science 2025 papers listed in the public submission table \(Acc\./Rej\.: accepted/rejected\)\. One rejected submission has no domain annotation\.
Table 7:Secondary topics within each primary category for the 247 Agents4Science 2025 papers listed in the public submission table\. Rejected counts are the difference between total and accepted counts\. Human\-Computer Interaction combines two punctuation variants\. Biomedical Engineering and Operations Research each appear under two primary categories\.
##### Matching\.
Every one of the 247 submissions is paired with a human anchor that passes H1, H2, and H4\. The full pairing rule, which also requires the same contribution type, pairs 185 of them: 66 anchors cited by the submission and 119 from the widened ICLR 2026 pool at tier 2 or higher\. The remaining 62 submissions cover topics poorly represented in ICLR, including physics, astronomy, biology, medicine, finance, and human\-computer interaction, so for these we keep H1, H2, and H4 fixed and take the closest admissible anchor in a recorded order\. Of the 62, 19 share the research problem, 21 share the contribution type and the area, 16 share the area, and 6 are matched to the nearest TF\-IDF candidate\. The last 22 are thus still paired with an accepted, human\-written paper, at the area level or by nearest neighbor, rather than left without a counterpart\. Overall, 206 of the 247 anchors match the contribution type and 196 match the research problem\. We record the match level for every pair, allowing evaluation on the strict 185 pairs, the 196 problem\-matched pairs, or all 247\.
##### Document processing\.
We read human anchors from their arXivLaTeXsource\. Among the AI submissions, 18 haveLaTeXsource consistent with the submitted PDF and are read directly; the remaining 229 are reconstructed from PDF into theLaTeXelements required by our measures\. The reconstruction extracts section headings, figure and table captions, citations, cross\-references, and result\-table cells using PyMuPDF\. It does not infer content absent from the PDF\. On the 18 submissions available in both formats, reconstruction recovers a median of 20 body headings compared with 21 in the source\. We record the processing route for every paper\.
### B\.4Comparison with existing benchmarks
Benchmark↓\\downarrowProperty→\\rightarrowScientifictextWholepaperTables andfiguresReleasedcodeExecutedresearchMatchedhuman pairMultiplegenresAdversarialvariantsRAID\([Dugan et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib8)\)✓\\checkmark––––✓\\checkmark✓\\checkmark✓\\checkmarkDetectRL\([Wu et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib50)\)✓\\checkmark––––○\\bigcirc✓\\checkmark✓\\checkmarkM4\([Wang et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib47)\)✓\\checkmark––––✓\\checkmark✓\\checkmark–MAGE\([Li et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib26)\)✓\\checkmark––––○\\bigcirc✓\\checkmark–CHEAT\([Yu et al\., 2023](https://arxiv.org/html/2610.00531#bib.bib53)\)✓\\checkmark––––✓\\checkmark–○\\bigcircIDMGSP\([Mosca et al\., 2023](https://arxiv.org/html/2610.00531#bib.bib36)\)✓\\checkmark○\\bigcirc–––✓\\checkmark–○\\bigcircSciSlopBench\(ours\)✓\\checkmark✓\\checkmark✓\\checkmark✓\\checkmark✓\\checkmark✓\\checkmark–△\\triangleTable 8:Comparison with existing machine\-generated text benchmarks\. \(✓\\checkmark\) the benchmark instances carry the property, \(○\\bigcirc\) partly, such as a human text from the same source rather than a matched counterpart, \(△\\triangle\) studied in the paper but not in the benchmark instances, \(–\) absent\.SciSlopBenchis designed to evaluate scientific failures that passage\-level detection benchmarks cannot represent \(Table[8](https://arxiv.org/html/2610.00531#A2.T8)\)\. Existing benchmarks primarily treat passages or abstracts as independent text instances\. In contrast,SciSlopBenchpreserves complete papers and pairs each AI paper with a human paper addressing the same research problem\. This makes relations across sections, claims, citations, figures, tables, and evidence directly measurable rather than reducing detection to local writing patterns\. The FARS half further includes the code that produced the reported results, enabling checks between manuscript claims and experimental records\. While existing benchmarks provide broader coverage of genres and adversarial rewrites,SciSlopBenchprovides the paper\-level structure and matched scientific context required to evaluate scientific slop\.
## Appendix CSciSlopHarnessImplementation Details
This appendix gives the implementation ofSciSlopHarnessused for Figure[5](https://arxiv.org/html/2610.00531#S5.F5)\. Table[9](https://arxiv.org/html/2610.00531#A3.T9)lists the configuration, and the following subsections describe each stage in the form of the harness summaries of[Lee et al\. \(2026a\)](https://arxiv.org/html/2610.00531#bib.bib23)\. Figure[5](https://arxiv.org/html/2610.00531#S5.F5)reports every AI\-generated paper inSciSlopBench, each revised for three rounds, and the baselines of Appendix[D](https://arxiv.org/html/2610.00531#A4)are run on the same papers\.
### C\.1Configuration
The editor, the six definitions, and their measurers stay fixed across all conditions, and so does the reviewer\. A round consists of one editor pass, one review pass, and one measurement\.
Table 9:Configuration of the default revision procedure\.
### C\.2Localized revision
##### Overview\.
The editor receives two files\.SciSlop\.mdsays how each pattern is repaired, andSLOP\_FINDINGS\.mdsays where the remaining instances are\. It revises one source file per call\. Table[10](https://arxiv.org/html/2610.00531#A3.T10)lists the repair direction of each pattern\.
- •Skill file\.SciSlop\.mdholds the six definitions, each with a repair direction and the format of its located instances\. Its global rules are short\. Do not fabricate experiments, numbers, citations, or examples\. Do not add or remove cited works and do not change reported numbers\. Keep the method, the experiments, and the results as they are\. Make the smallest edit that removes the pattern\. Treat listed locations as candidates and leave one alone when the repair would add text the argument does not need\. The file has no score formula and no target\.
- •Located feedback\.SLOP\_FINDINGS\.mdlists the remaining instances by pattern, up to 8 per pattern with the count of the rest\. From round 2 onward, the previous round’s review notes are appended: the reverted changes with the reviewer’s reasons, and the retired instances\.
- •Editor call\.For each ofmain\.texandsections/\*\.tex, one call sendsSciSlop\.md,SLOP\_FINDINGS\.md, the other files as read\-only context, and the file to edit\. The editor edits only its own file\. It receives no human paper, no score, and no formula\.
- •Edit blocks\.The call returns edit blocks, each quoting an original passage and giving its replacement, or the wordNOCHANGE\. A block is applied only when the quoted passage occurs exactly once in the file\. An edited file is accepted only if it keeps the file’s structural commands and its length stays within 0\.5 to 2\.5 times the original\.
- •Figure editing\.The method figure is a raster image, so the editor lists the expository phrases to erase and a deterministic step erases them\. The controller keeps only phrases the measurer had classified as expository\. A vision\-language model returns a bounding box for each phrase and an inpainting model fills the box from its surroundings\. An edit report records the erased phrases, the boxes, and the pixels changed outside them\.
- •Records\.Before each round, the controller copies the project’s own experiment records fromexp/EXPERIMENT\_RESULTSintomaterials/\. Once per paper, a specimen search checks whether these records hold a displayable instance and must quote it verbatim from a named file\.
Table 10:Repair directions and location formats ofSciSlop\_v0\.5\.md\. All six entries supply feedback and enter the stopping rule\.
### C\.3Review against scientific records
##### Overview\.
The reviewer reads every change of the round as a before\-and\-after pair and keeps or reverts each one\. It checks the change against the original manuscript and the project records\.
- •Changes\.The controller alignsMr−1M\_\{r\-1\}andM~r\\widetilde\{M\}\_\{r\}paragraph by paragraph\. Each differing paragraph becomes a numbered change with its before\-and\-after text\. A figure edit becomes one further change described by its edit report\.
- •Review call\.Changes are reviewed three at a time, in parallel\. Each call receives the rules below, its three changes, the full revised and original text of the files they touch, and the lines of the records that contain any number used in the revised text\. It returns keep or revert, a category, and a reason quoting the decisive words for each change\.
- •Revert rules\.The revised text says something that neither the manuscript nor the records support\. A reported number is changed, rounded, dropped, or introduced without an exact match\. A reference, section pointer, or citation is added to a sentence that does not use it\. Two cited works are joined by a relation the manuscript never stated\. A figure, table, label, citation command, or result is removed\. A figure edit erased a phrase the check did not list or a component of the mechanism\. A new exhibit is not an exact copy of a record with its source named\. Citing sentences are merged without a supported relation\.
- •Keep rules\.The change repairs a listed pattern with material already in the manuscript or records, in a sentence that uses what it refers to, and leaves every number and cited work as it was\. A sentence stating that no individual case was inspected and that results rest on aggregate scores is kept when the records hold no specimen\.
- •Application\.A change mixing supported and unsupported edits is reverted in full, with the reason naming the unsupported part\. Only an explicit keep retains a change\. Everything else is restored fromMr−1M\_\{r\-1\}\.
- •Retirement\.A separate call proposes instances to leave alone for the rest of the procedure\. It must look for a home for the object in every other section and name the sections it checked\. The controller applies retirement only to section instances and to the evidence\-gap observation\. Retired instances leave the feedback but stay in the measurements\.
### C\.4Measurement, preservation, and stopping
- •Measurement after review\.The retained manuscript is measured after review, so reverted edits cannot contribute to the reported reduction\. The four deterministic measurers run every round\. The argument graph is remeasured whenever the introduction changed, with the same labeling server and estimator as the benchmark\. The figure exposition is rescored from the transcript of the current image\. A manuscript whose measured evidence gap is 1\.0 and whose text states that no individual case was retained or inspected and that the claims rest on aggregate scores is recorded as an acknowledged gap of 0\.5\. The matched human values are the unmodified measurer values\.
- •Preservation checks\.Table[11](https://arxiv.org/html/2610.00531#A3.T11)lists the checks computed againstM0M\_\{0\}after every round\. A hard violation disqualifies the round as successful mitigation\. Audit flags mark changes that may be legitimate but require inspection, such as a removed recycled sentence that also removed a repeated number\. Numbers are compared as tokens after removing comments and the arguments of citation, reference, label, input, and graphics commands\.
- •Stopping\.The procedure stops when no instance remains after retirement, when the editor changes nothing, when review restores the whole round twice in a row, or after three rounds\. The last retained manuscript is returned with its per\-round trees, measurements, verdicts, retirements, edit reports, and preservation records\.
Table 11:Preservation checks recorded after every round\.Hardconstraints define admissible mitigation\.Auditflags require examination of the underlying change\.
### C\.5Distance to human means
Table 12:Distance from the human mean after the third revision round \(revised mean−\-human mean\) on the all papers of Figure[5](https://arxiv.org/html/2610.00531#S5.F5)\. Bold marks the smallest distance per row\.Table[12](https://arxiv.org/html/2610.00531#A3.T12)gives the values behind the 63% reduction reported in Section[5\.3](https://arxiv.org/html/2610.00531#S5.SS3)\. For each measure, the distance is the mean score of the revised papers after the third round minus the matched human mean, the quantity plotted in Figure[5](https://arxiv.org/html/2610.00531#S5.F5)\. The gap of a method is the sum of its absolute distances over the six measures\. Claude Code has the smallest gap among the four baselines, so the reduction is computed against it\.SciSlopHarnesscloses the gap from 1\.395 to 0\.519, a reduction of 63%, and ends closest to the human mean on every measure\. Slop\-aware revision passes below the human mean on Macro redundancy, Cross\-section references, and Citation isolation, so its gap grows over rounds although its scores fall\.
## Appendix DRevision Baseline Details
This appendix specifies the four revision baselines compared withSciSlopHarnessin Section[5](https://arxiv.org/html/2610.00531#S5): base prompting, Claude Code, reviewer\-based refinement, and slop\-aware revision\. The baselines differ only in the interface through which the editor acts and in the instruction it receives\. The editor model, the manuscript state, the round protocol, and the measurement are shared\. Appendix[C](https://arxiv.org/html/2610.00531#A3)specifiesSciSlopHarnessitself\.
### D\.1Shared protocol
##### Editor and rounds\.
Every baseline usesclaude\-haiku\-4\-5\-20251001as the editor, the same model as the defaultSciSlopHarnesseditor in Table[9](https://arxiv.org/html/2610.00531#A3.T9)\. Each paper is revised for three rounds\. Roundrrstarts from a copy of the manuscript retained after roundr−1r\-1, and round 1 starts from the original AI\-generated manuscript\. The copied tree contains the root\.tex,\.bib,\.bst, and\.styfiles andsections/\*\.tex\. Figures and compiled PDFs are not copied, because the editor cannot read them and the measurers do not use them\. Each round issues one editor call, which for agent conditions is one agent session that may contain several tool interactions\.
##### Information given to the editor\.
No baseline receives human papers, matched human reference values, slop scores, score formulas, or thresholds\. All four baselines are instructed not to fabricate experiments, numbers, or citations\. The boxes below show each baseline’s system and user messages, including its exact constraints\. Slop\-aware revision additionally receives the pattern definitions and located instances described in Appendix[D\.5](https://arxiv.org/html/2610.00531#A4.SS5)\.
##### Execution record\.
If the editor makes no changes, we retry once\. If the retry also makes no changes, we exclude that round\. We save each prompt and revision and evaluate revisions with the same metrics\. We also check for content loss and citations missing from the bibliography\.
### D\.2Base prompting
Base prompting is a plain text condition without tools\. The editor receives the completeLaTeXsource of the paper inside the prompt and returns the complete revised source of every file inside file markers\. The user instruction is the generic improvement instruction followed by the format specification\.
Base promptingSystem\.You are an expert editor of machine\-learning research papers written in LaTeX\. You return revised LaTeX source only\.User\.Revise this paper so that it is a better research paper: improve its clarity, organization, argumentation, and presentation, and make sure every part of the paper does its job for the reader\. Keep the method, the experiments, and the reported results as they are\. Do not fabricate experiments, numbers, or citations; if evidence is missing, narrow the claim instead\. Do not add citations that are not already in the paper, and do not change reported numbers\.This is revision round⟨r⟩\\langle r\\rangle\. The complete LaTeX source of the paper follows, file by file\. Return the revised source of every file, in the same order, using exactly this format and nothing else:<<<FILE \{path\}\>\>\>\(the complete revised content of that file\)<<<END \{path\}\>\>\>Return all⟨n⟩\\langle n\\ranglefiles \(⟨\\langlepaths⟩\\rangle\)\. Keep every\\input,\\label,\\cite,\\ref, environment and preamble line that a file already has, unless the revision moves or removes prose; never remove\\begin\{document\},\\end\{document\},\\documentclass, or\\inputlines\. Do not abbreviate or elide any part of a file\.⟨\\langlesource of each file, preceded by===== FILE \{path\} =====⟩\\rangle
A returned file replaces the original only if it is present in the output, its length lies between 0\.5 and 2\.5 times the original, it preserves every\\input,\\documentclass,\\begin\{document\}, and\\end\{document\}token of the original, and it is not identical to the original\. A file that fails any check keeps its previous content, and the reason is logged\. These checks protect the compilability of the tree and do not inspect the prose\.
### D\.3Claude Code
Claude Code\([Anthropic, 2025](https://arxiv.org/html/2610.00531#bib.bib2)\)runs with its stock system prompt in the paper directory, with the same editor model as the other conditions\. It runs in restricted mode, which removes shell and code execution and ignores user configuration, so the available tools are file reading, searching, and editing\. The instruction is the generic improvement instruction of base prompting with one added sentence\.
Claude CodeSystem\.The stock Claude Code system prompt is used without a custom replacement\.User\.Revise this paper so that it is a better research paper: improve its clarity, organization, argumentation, and presentation, and make sure every part of the paper does its job for the reader\. Keep the method, the experiments, and the reported results as they are\. Do not fabricate experiments, numbers, or citations; if evidence is missing, narrow the claim instead\. Do not add citations that are not already in the paper, and do not change reported numbers\.Make the changes directly by editing the \.tex files in the current directory; do not just give advice\.
### D\.4Reviewer\-based refinement
Reviewer\-based refinement alternates an automated review with an editing session\. At each round, Qwen2\.5\-32B\-Instruct\([Yang et al\., 2024](https://arxiv.org/html/2610.00531#bib.bib52)\)reads the current manuscript source and returns a structured review\. The review is regenerated from the current manuscript each round, so later rounds respond to the revised paper rather than to the original\. The source is passed as text truncated to 50,000 characters\. If the model returns no valid JSON, the call is repeated at 24,000 and then 14,000 characters\.
Reviewer\-based refinement / Review generationSystem\.No separate system message is supplied\.User\.You are a knowledgeable, critical, but fair reviewer for a top machine\-learning conference\. Read the paper draft below and write a review of its overall quality: the significance of the problem, the soundness of the approach, the strength of the evidence, and the clarity of the writing\.Respond with ONLY compact JSON in exactly this format:\{"summary": "<2\-4 sentence summary of the paper and its contributions\>","strengths": \["<strength 1\>", "<strength 2\>", "<strength 3\>"\],"weaknesses": \["<weakness 1\>", "<weakness 2\>", "<weakness 3\>"\],"questions": \["<question 1\>", "<question 2\>", "<question 3\>"\],"rating": <integer 1\-10, where 1 = trivial or wrong and 10 = seminal\>\}PAPER \(LaTeX source, possibly truncated\):⟨\\langlemanuscript source⟩\\rangle
The editor is an agent with a neutral system prompt and with read, edit, and write tools\. Web search is denied and shell commands are not permitted\. The review prompt and the editing instruction are general\. Neither mentions SciSlop, the six patterns, or any measured quantity\.
Reviewer\-based refinement / Manuscript editingSystem\.You are revising a LaTeX research paper\. Source files \(main\.tex, sections/\*\.tex, math\_commands\.tex, a \.bib file\) are in your working directory\. Use Read/Edit/Write to modify the \.tex files in place\. Do not create files outside the paper directory\. When done, finish your turn; do not run latex and do not summarize\.User\.A reviewer left the following review of this paper\. Revise the paper to address the review\.⟨\\langlereview JSON⟩\\rangleDo not fabricate experiments, numbers, or citations; if evidence is missing, narrow the claim instead\.
### D\.5Slop\-aware revision
Slop\-aware revision gives the editor a list of detected slop problems and instructions for fixing them in place of a general review\. The editor interface, system prompt, tools, and budget are the same as in reviewer\-based refinement\. At the start of each round, we check the current manuscript and prepare one prompt containing the detected problems\. For each problem, the prompt gives its name, definition, repair instructions, and locations in the manuscript\. The definitions and instructions follow the measurer specifications\. We check the first four patterns of Table[13](https://arxiv.org/html/2610.00531#A4.T13), adding*Argument graph*and*Figure exposition*for the comparison withSciSlopHarness\. The prompt lists up to 12 instances per pattern and reports the number of additional instances\. For*Evidence gap*, it simply states that the manuscript contains no concrete example\. Only detected problems are included, and a round is skipped if no problems are found\. The editor receives no scoring formulas or target scores\. Table[13](https://arxiv.org/html/2610.00531#A4.T13)gives the definition and repair text of each pattern\.
Slop\-aware revisionSystem\.You are revising a LaTeX research paper\. Source files \(main\.tex, sections/\*\.tex, math\_commands\.tex, a \.bib file\) are in your working directory\. Use Read/Edit/Write to modify the \.tex files in place\. Do not create files outside the paper directory\. When done, finish your turn; do not run latex and do not summarize\.User\.An automated check of this manuscript found the following issues\. Revise the paper to address them\. Do not fabricate experiments, numbers, or citations; if evidence is missing, narrow the claim instead\. Do not add citations that are not already in the paper, and do not change reported numbers\. Edit the \.tex files in place\.This is revision round⟨r⟩\\langle r\\rangle\.\#\# Issue 1:⟨\\langlepattern name⟩\\rangle⟨\\langledefinition⟩\\rangleWhat to do:⟨\\langlerepair direction⟩\\rangleFindings \(⟨k⟩\\langle k\\rangletotal;⟨\\langleunit label⟩\\rangle\):\-⟨\\langlelocated unit⟩\\rangle\(up to 12 lines\)\- \(and⟨k−12⟩\\langle k\-12\\ranglemore of the same kind; treat them the same way\)\#\# Issue 2: …
Table 13:Definitions and repair instructions for slop\-aware revision\. The first four rows are used in the four\-pattern setting\.
## Appendix EICLR Data and Measurement
Figure[4](https://arxiv.org/html/2610.00531#S5.F4)uses two corpora of human\-written ICLR submissions, 1,177 papers from 2017 to 2025 and 691 from ICLR 2026, scored by every system of Table[3](https://arxiv.org/html/2610.00531#S4.T3)with the same scripts, settings, and document views\. The measurers and the calibration were fixed onSciSlopBench, and every statistic is computed within a year or within ICLR 2026, because the ICLR rating scale changed over the years\.
##### Corpora\.
ICLR is the one venue where every submission, accepted or rejected, carries public review scores and a decision, and its arXivLaTeXe\-prints give the source that the measures read\. The multi\-year corpus therefore holds ICLR 2017 to 2025 submissions with both, nine years that span the arrival of LLM writing\. It was collected in strata of year by integer rating and of year by award, so that every year holds papers from rating 2 to rating 9 and its orals; at the ends of the scale the strata take every paper that exists\. Of 1,381 decided papers, 204 fail the extraction rules below, leaving 1,177, 590 rejected, 434 accepted, and 153 oral, with 78 to 240 papers a year\. The ICLR 2026 corpus serves the review dimensions\. It is the one year with a soundness, presentation, and contribution score for every paper on one rating scale, and the one year with an independent commercial AI judgment of the same full text, an analysis that Pangram published for every ICLR 2026 submission\. Its 691 submissions with an arXiv e\-print were sampled over six bands of that AI share crossed with the decision, so that the high\-AI tail, mostly rejected and rarely on arXiv, is present; 560 are accepted and 131 rejected\.
##### Exclusion\.
A paper is removed only when its source cannot be measured: a missing`\\input`file, a main file that is a supplement, a failed section parse, fewer than three body sections, fewer than five declared objects, a prose view under 1,500 words, or any measure below its minimum denominator \(Appendix[A](https://arxiv.org/html/2610.00531#A1)\)\. ICLR 2026 adds fewer than three reviews and a body length outside the 1st to 99th percentile\. No rule reads a score or a rating, and a removed paper is removed for every system\.
##### Systems and scores\.
Detectors read the prose view, reviewers the body source, and the measures the full source\. Every score is oriented so that a higher value means more AI\-like, which reverses Binoculars and the reviewers’ Overall rating\. The detectors score all 1,177 papers; the two reviewer models score a subset drawn by rating quintile within each year and filled at random by year and decision, 287 papers for CycleReviewer and 282 for AI Scientist\. The SciSlop line is the aggregate of Eq\.[7](https://arxiv.org/html/2610.00531#A1.E7)\(Appendix[A](https://arxiv.org/html/2610.00531#A1)\) scored on every paper of both corpora\.*Argument graph*was not run on ICLR 2026, and*Evidence gap*and*Figure exposition*apply to 869 and 662 of the multi\-year papers\.
##### Panel \(a\)\.
Each system is calibrated once on the FARS pairs ofSciSlopBenchby a logistic regression of the AI label on its raw score, and the fitted curve is applied unchanged to the ICLR papers, giving the probability the system assigns to AI authorship\. Each probability is adjusted for year by subtracting the mean of its own year and adding back the mean over all years, so the axis stays on the probability scale; a paper sits at its rounded mean rating\. Levels 2 to 9 hold 31 to 217 papers each; levels 1 and 10, with 11 and 1 papers, are not drawn\.
##### Panel \(b\)\.
A cell is the Spearman correlation between a system’s oriented score and the mean reviewer score on one dimension of ICLR 2026, over the papers the system scored, 686 for the detectors, 691 for Pangram, 680 for SciSlop, and 338 and 285 for CycleReviewer and AI Scientist\.
##### Panel \(c\)\.
Within each year, the value is the probability that a rejected paper scores more AI\-like than an accepted or oral paper of the same year\. A system is drawn when at least 15 papers stand on each side in at least five years; the reviewer models meet this in three years and are not drawn\.
## Appendix FAdditional Matched Revision Cases
Section[5\.4](https://arxiv.org/html/2610.00531#S5.SS4)compared slop\-aware revision andSciSlopHarnesson the same passages for two measures, Macro redundancy and Cross\-section references, and showed that the two lower the score by different means, one by deleting or scattering information and the other by repairing the located passage\. This appendix extends that comparison to the remaining four measures, Citation isolation, Evidence gap, Argument graph, and Figure exposition\.
##### The same distinction holds for the remaining four measures\.
Table[14](https://arxiv.org/html/2610.00531#A6.T14)repeats the format of Table[4](https://arxiv.org/html/2610.00531#S5.T4)for the four measures it does not cover and shows the same pattern\. Slop\-aware revision lowers the score by leaving, deleting, or inventing content, whereasSciSlopHarnessrepairs the located passage itself\.
- •Citation isolationshows whether a repair relates the cited work to anything else\. Slop\-aware revision rewrites the CALM sentence but still describes CALM alone, so the located sentence stays isolated; its lower manuscript score comes from changes elsewhere in the paper\.SciSlopHarnesskeeps the sentence and adds the contrast with InfMem, which stops at the chunk level and reports a 3\.3–5\.1×\\timesspeedup\. The isolation is removed by adding a relation, not by rewording the sentence\.
- •Evidence gapshows what a revision does when the record holds no concrete instance\. Slop\-aware revision reaches a zero score by inserting a worked example whose question, responses, and numbers appear nowhere in the manuscript or the project records\.SciSlopHarnessstates that the claims rest on aggregate results only, and its review reverts the one example the editor inserted as beyond the record\. The gap stays acknowledged at0\.50\.5instead of being filled with an invented instance\.
- •Argument graphshows whether a claim follows the method it depends on\. Slop\-aware revision leaves the design claim ahead of the method and adds a section pointer instead\.SciSlopHarnessmoves the claim directly after the method description without deleting either sentence, and the manuscript score falls to zero\. The repair reorders existing content rather than adding text\.
- •Figure expositionshows whether a revision can reach text inside an image\. Slop\-aware revision edits source text only, so all eight phrases that restate the paper remain in the diagram and the score stays at its original value\.SciSlopHarnesserases six of the eight phrases while leaving the components and arrows intact\. This comparison is not under equal tools, since the baseline has no image editor, but it shows how far the harness’s image edits reach\.
Table 14:One passage per remaining pattern before revision, after slop\-aware revision, and afterSciSlopHarness, markingthe located slop,what a revision lost or left, andwhat it kept or repaired\. Brackets replace citation and reference commands, and the arrow marks the order of two statements across intervening text\. Manuscript scores \(original / slop\-aware /SciSlopHarness\):0\.615 / 0\.385 / 0\.571,1\.000 / 0\.000 / 0\.500,0\.333 / 0\.167 / 0\.000, and0\.500 / 0\.500 / 0\.167, where the baseline edits no images and so keeps the original figure score\.Similar Articles
I benchmarked which of 18 AI models writes the least like "AI slop"
A developer built an open-source benchmark called The Slop Index to measure how much 18 AI models produce 'AI slop', using human baselines and five dimensions including conciseness, templating, and human preference.
I found a way to fight AI slop
The author proposes using AI as a research moderation and synthesis tool rather than a content generator to combat 'AI slop.' By building a pipeline that compares and ranks expert sources, the author argues the future human role in AI is curation and judgment.
The “AI slop” debate conflates three separate questions: authorship, productivity, and engineering quality
The article argues that the 'AI slop' debate conflates authorship, productivity, and engineering quality, and proposes treating generative coding systems as high-throughput, error-prone producers within an engineering control loop, shifting scarce skills toward specification, verification, and accountability.
slop-grader
Slop-grader is a rule-based CLI tool for evaluating documents against custom rulesets, generating scores and flags to assist AI agents in fixing issues like AI filler in text.
Training AI on AI slop
The article discusses the problem of training AI models on low-quality data generated by other AI systems, known as 'slop,' and its implications for model reliability and performance.