Tag
A user on OpenReview noticed that an Area Chair's comment and their reply have disappeared, raising concerns about potential bias in paper rejection decisions. They are seeking confirmation from others if this is a normal occurrence.
A reviewer for AAAI 2027 expresses surprise at the low number of paper submissions with code, despite the conference's emphasis on reproducibility, and asks for opinions on whether lack of code should affect review scores.
The article examines how the peer review system is struggling to cope with the exponential growth of research publications and AI-assisted papers, leaving volunteer reviewers overwhelmed and leading to errors and delays, prompting calls for reform.
This paper investigates how rhetorical framing biases AI-based peer review scores, finding that evidence framing and novelty stance have the largest effects and that score movements depend on the reviewer's initial score and evaluation strictness.
This paper demonstrates that large language models can effectively deanonymize authors of scientific papers from titles and abstracts alone, threatening the validity of double-blind peer review. The authors argue that stable patterns in problem framing and research focus act as latent conceptual signatures of authorship, necessitating a re-evaluation of anonymity practices in AI-augmented research ecosystems.
Introduces RubricReviewer, a rubric-driven framework for LLM-based peer review that explicitly generates paper-adaptive rubrics and combines a training-free evidence-gathering agent (Scout) with a trained human-aligned model (Aligner) to produce more comprehensive, discriminative, and robust reviews.
The author, after reviewing for three major conferences, argues that papers without code to reproduce results should be desk rejected, citing that only 1 of 12 papers reviewed provided full code and 7 provided none.
An opinion piece urging NeurIPS reviewers to raise scores when rebuttals address their stated concerns, regardless of personal taste, sparking discussion about the review process.
This paper introduces AutoSupervision, a benchmark and method for verifying whether manuscript revisions actually address reviewer concerns using grounded evidence from peer-review records. Experiments on 56,000 Nature Communications articles show LLMs can summarize reviewer concerns but still struggle with evidence-based verification.
An experience report from BOSC 2026 on using generative AI to pre-review open-source software submissions, with human reviewers making final decisions. Most reviewers found the AI-assisted pre-review useful but preferred to verify AI conclusions independently.
A position paper arguing that autonomous AI agents in science widen the verification gap and that scientific verification infrastructure must evolve with observable-by-default workflows, scalable verification, and clear attribution to sustain trustworthy science.
A reviewer recounts flagging two papers with fabricated authors that were accepted as orals, and reports that 68% of 22 reviewed submissions contained fabricated citations, LLM-generated content, or fake author lists. Multiple studies confirm tens of thousands of such papers in 2025 alone, with peer reviews increasingly AI-generated.
This paper introduces intra-paper claim verification, a framework that uses LLMs to evaluate whether novelty claims in a paper are supported by its methodological evidence, addressing a gap in existing automated peer review systems. Human evaluation shows significant alignment with human reviewer concerns, especially for novelty-related issues.
A NeurIPS reviewer reports encountering a paper and rebuttals that appear entirely LLM-generated, expressing frustration and seeking advice on how to evaluate such submissions.
This study analyzes how different types of reviewer guidelines (official conference guidelines vs. reviewer-imitating ones) affect LLM-based automated peer review, finding that official guidelines produce more human-consistent results while strict rubric-style scoring degrades performance.
A Twitter thread analyzes the unusually long peer review process for a cell embedding paper using GPT-5.6 to compare the preprint and final publication, estimating time, compute, personnel, and APC costs, highlighting the cost-benefit ratio of journal peer review.
This paper introduces Kahneman4Review, a benchmark of 3,563 peer reviews rated along nine theoretically motivated textual dimensions, eight bias diagnostics, and a continuous reasoning-quality score, to evaluate whether LLM judges can distinguish analytical form from genuine epistemic quality in peer review.
This position paper proposes a credit system for ML conferences to incentivize quality reviewing by awarding points for good behavior and allowing redemption for perks.
AI-research-feedback is an academic paper review skill for Claude Code. It checks grammar, coherence, formulas, figures, and argument flaws through six parallel agents, supports specifying journals to simulate reviewers, and finally generates a structured review report.
Google deployed an agentic AI peer-reviewer at ICML and STOC conferences, reviewing ~10,000 papers with 30-minute turnaround. The formal paper shows it catches 34% more mathematical errors than zero-shot prompting, setting a precedent for AI-automated scientific review at scale.