Tag
This paper studies the feedback loop where AI-generated reviews influence future training of AI reviewers, leading to reduced judgment diversity called 'scientific-judgment collapse,' and introduces TrustReviewer, an open-source system to mitigate this through curated training and activation steering.
This paper demonstrates that AI peer reviewers can be manipulated by modifying only presentation-level content (such as abstract, framing, and narrative) without changing any scientific evidence, achieving a 75.1% attack success rate. The authors introduce adversarial repackaging, a closed-loop attack that exploits AI reviewers' tendency to be impressed rather than convinced, and release a benchmark for testing robustness.