Tag
The article describes motive.md, an open-source decentralized system where AI agents collaborate on mathematical problems using skill-based contributions and structured hypothesis-evidence memory, mimicking a scientific method loop.
Elon Musk proposes that AI competitors engage in informal peer review and early access sharing before model releases to enhance safety and standards.
NovGauge is a human-anchored benchmark for diagnosing LLMs' capability in paper novelty assessment across task, problem, and method dimensions. Evaluation of 18 LLMs reveals high hallucination rates and logical mismatches, indicating current models are unreliable for this task.
This paper introduces a process-centric benchmark for evaluating AI-assisted peer review systems, aiming to improve transparency and reliability beyond final decision accuracy.
This paper evaluates LLM systems for pre-submission peer review, demonstrating that broad generation can recover most historical review issues but compressing them into a short report is challenging, with implications for AI-assisted research feedback.
This paper introduces HalluPeer, a taxonomy-driven benchmark for detecting hallucinations in scientific peer reviews, providing annotated data to evaluate and improve detection methods.
This paper presents InternReviewer and InternAdvocate, a framework for developing AI agents in academic peer review and rebuttal using agentic reinforcement learning with objective rewards to improve reasoning depth, citation accuracy, and reduce hallucinations.
This study evaluates two multimodal LLMs as peer reviewers for ICLR 2026 submissions, finding high scoring calibration but low error detection, with author identity having no effect and figures reducing error detection.
FirstPass is a large-scale peer review dataset from Nature Communications, covering multiple scientific domains and multi-round dialogues to improve AI models for scientific judgment.
The article discusses collusion in the AAAI 2027 review process, particularly in reviewer assignment cycles, and critiques the lack of code publication in accepted papers at top AI conferences.
This paper introduces Metag, a dataset designed to help meta-reviewers in scientific peer review by identifying changes in manuscripts based on reviewer feedback and author responses, enhancing traceability and transparency.
The article introduces Tree-of-Concerns, a hierarchical multi-agent debate framework that uses specialized personas to extract unstated limitations from scientific papers, demonstrating improved precision and coverage on a benchmark dataset.
404 Media investigation reveals that Research Gold, a service advertising '100% human-written' medical peer reviews, is entirely AI-operated with fabricated staff credentials and AI-generated responses.
This position paper argues that improving machine learning peer review requires enforceable procedural safeguards and a spendable credit system, such as OpenReview Points, to incentivize good reviewing and limit submission volume.
A user on OpenReview noticed that an Area Chair's comment and their reply have disappeared, raising concerns about potential bias in paper rejection decisions. They are seeking confirmation from others if this is a normal occurrence.
A reviewer for AAAI 2027 expresses surprise at the low number of paper submissions with code, despite the conference's emphasis on reproducibility, and asks for opinions on whether lack of code should affect review scores.
The article examines how the peer review system is struggling to cope with the exponential growth of research publications and AI-assisted papers, leaving volunteer reviewers overwhelmed and leading to errors and delays, prompting calls for reform.
This paper investigates how rhetorical framing biases AI-based peer review scores, finding that evidence framing and novelty stance have the largest effects and that score movements depend on the reviewer's initial score and evaluation strictness.
This paper demonstrates that large language models can effectively deanonymize authors of scientific papers from titles and abstracts alone, threatening the validity of double-blind peer review. The authors argue that stable patterns in problem framing and research focus act as latent conceptual signatures of authorship, necessitating a re-evaluation of anonymity practices in AI-augmented research ecosystems.
Introduces RubricReviewer, a rubric-driven framework for LLM-based peer review that explicitly generates paper-adaptive rubrics and combines a training-free evidence-gathering agent (Scout) with a trained human-aligned model (Aligner) to produce more comprehensive, discriminative, and robust reviews.