When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation
Summary
This paper studies the feedback loop where AI-generated reviews influence future training of AI reviewers, leading to reduced judgment diversity called 'scientific-judgment collapse,' and introduces TrustReviewer, an open-source system to mitigate this through curated training and activation steering.
View Cached Full Text
Cached at: 09/21/26, 03:20 AM
Paper page - When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation
Source: https://huggingface.co/papers/2609.20942
Abstract
Largelanguagemodels(LLMs)increasinglyparticipateinscientificevaluation,bothasautomatedreviewersandasassistantstohumanreviewers.Asmodel-generatedreviewsenterpublicdataandfuturetrainingcorpora,AIpeerreviewcanbecomerecursive:laterreviewerslearnfromjudgmentsproducedbyearliermodels.Westudyonestepofthisfeedbackloopinacontrolledsetting.StartingfromLlama3.18B,wefirstfine-tunearevieweronofficialICLRreviewsfrom2018--2023andthentrainfoursuccessormodelsonICLR2024datawithsystematicallyvariedmixturesofofficialandmodel-generatedreviews.Ourstudyshowsthatintroducingsyntheticreviewscompressesratingdistributionsandreducesbothsame-paperandcorpus-levelsemanticdiversity.Wecallthispatternscientific-judgmentcollapse.Tomitigatethisfailuremode,weintroduceTrustReviewer,anopen-sourceLLM-basedsystemforgeneratingpeerreviewsofAIandmachinelearningpapers.TrustReviewerintervenesattwocomplementarystages.Fortraining-timeprevention,wetrainthecorereviewerinasinglestageonacuratedcorpusdesignedtoreducelow-qualityandsemanticallydegeneratesupervision.Fortest-timecorrection,pairedactivationsteeringaimstofurthermitigateresidualtendenciestowardcollapsedjudgmentswithoutfurthertrainingoradditionalexpertannotation.Together,theseresultscharacterizeaconcreteriskofrecursivereviewertrainingandprovidepracticalinterventionsforpreservingjudgmentdiversityandimprovingrecommendationalignmentinAI-assistedscientificevaluation.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2609\.20942
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.20942 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.20942 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.20942 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Gaming AI-Assisted Peer Reviews Poses New Risks to the Scientific Community
A new study demonstrates that AI-assisted peer review is vulnerable to low-cost manipulation via superficial rephrasing of paper abstracts, significantly inflating AI-generated review scores and potentially biasing human editorial decisions, highlighting the need for safeguards.
On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists
A study evaluating AI reviewers (GPT-5.2, Claude Opus 4.5, Gemini 3.0 Pro) against 45 expert human reviewers on Nature-family papers found that AI reviewers can exceed top-rated humans in aggregate review quality, though they are less correct but raise more significant issues.
AI-written critiques help humans notice flaws
OpenAI trained language models to write critiques of text summaries, helping human evaluators spot flaws more effectively — a step toward scalable oversight of AI systems on difficult tasks. The work explores how AI-assisted feedback can improve human evaluation quality as a proof of concept for alignment research.
The Trust–Oversight Paradox: As AI Gets Better, Humans May Stop Really Overseeing It
A thought piece arguing that as AI becomes more accurate, human oversight may degrade into routine approval, creating a 'Trust–Oversight Paradox' where high-performing AI can still fail due to incomplete representation, stale data, or automation bias, suggesting a shift from human review to governing boundaries.
NeurIPS 2026 AI-generated reviews [D]
Discussion about the use of AI-generated reviews at NeurIPS 2026, including concerns over prompt injection and lack of consequences for reviewers using LLMs without proper oversight.