peer-review

Tag

Cards List
#peer-review

motive.md - like kickstarter + BOINC for shared agent goals... focusing on math problems and expanding. just give an agent a skill.md and it starts contributing

Reddit r/openclaw · 3d ago

The article describes motive.md, an open-source decentralized system where AI agents collaborate on mathematical problems using skill-based contributions and structured hypothesis-evidence memory, mimicking a scientific method loop.

0 favorites 0 likes
#peer-review

@elonmusk: I think this is worth doing

X AI KOLs Timeline · 6d ago Cached

Elon Musk proposes that AI competitors engage in informal peer review and early access sharing before model releases to enhance safety and standards.

0 favorites 0 likes
#peer-review

NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment

arXiv cs.AI · 2026-09-12 Cached

NovGauge is a human-anchored benchmark for diagnosing LLMs' capability in paper novelty assessment across task, problem, and method dimensions. Evaluation of 18 LLMs reveals high hallucination rates and logical mismatches, indicating current models are unreliable for this task.

0 favorites 0 likes
#peer-review

Beyond Final Decisions: A Process-Centric Benchmark for Transparent AI-Assisted Peer Review

arXiv cs.AI · 2026-09-10 Cached

This paper introduces a process-centric benchmark for evaluating AI-assisted peer review systems, aiming to improve transparency and reliability beyond final decision accuracy.

0 favorites 0 likes
#peer-review

More Than Mimicking Reviewers: Evaluating LLMs for Pre-Submission Peer Review

arXiv cs.AI · 2026-09-10 Cached

This paper evaluates LLM systems for pre-submission peer review, demonstrating that broad generation can recover most historical review issues but compressing them into a short report is challenging, with implications for AI-assisted research feedback.

0 favorites 0 likes
#peer-review

HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews

arXiv cs.AI · 2026-09-04 Cached

This paper introduces HalluPeer, a taxonomy-driven benchmark for detecting hallucinations in scientific peer reviews, providing annotated data to evaluate and improve detection methods.

0 favorites 0 likes
#peer-review

InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal

arXiv cs.AI · 2026-09-01 Cached

This paper presents InternReviewer and InternAdvocate, a framework for developing AI agents in academic peer review and rebuttal using agentic reinforcement learning with objective rewards to improve reasoning depth, citation accuracy, and reduce hallucinations.

0 favorites 0 likes
#peer-review

Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects

arXiv cs.CL · 2026-09-01 Cached

This study evaluates two multimodal LLMs as peer reviewers for ICLR 2026 submissions, finding high scoring calibration but low error detection, with author identity having no effect and figures reducing error detection.

0 favorites 0 likes
#peer-review

FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes

arXiv cs.CL · 2026-08-28 Cached

FirstPass is a large-scale peer review dataset from Nature Communications, covering multiple scientific domains and multi-round dialogues to improve AI models for scientific judgment.

0 favorites 0 likes
#peer-review

AAAI 2027 Reviewer Bidding and Assignment Integrity [D]

Reddit r/MachineLearning · 2026-08-24

The article discusses collusion in the AAAI 2027 review process, particularly in reviewer assignment cycles, and critiques the lack of code publication in accepted papers at top AI conferences.

0 favorites 0 likes
#peer-review

Metag: A dataset to build agentic meta-reviewing capabilities

arXiv cs.LG · 2026-08-24 Cached

This paper introduces Metag, a dataset designed to help meta-reviewers in scientific peer review by identifying changes in manuscripts based on reviewer feedback and author responses, enhancing traceability and transparency.

0 favorites 0 likes
#peer-review

Tree-of-Concerns: Hierarchical Multi-Agent Debate for Unstated-Limitation Extraction in Scientific Critique

arXiv cs.CL · 2026-08-24 Cached

The article introduces Tree-of-Concerns, a hierarchical multi-agent debate framework that uses specialized personas to extract unstated limitations from scientific papers, demonstrating improved precision and coverage on a benchmark dataset.

0 favorites 0 likes
#peer-review

404 Media: Research Gold Sold '100% Human-Written' Medical Peer Review That Was Entirely AI, Including Fake PhD Staff

Reddit r/ArtificialInteligence · 2026-08-18

404 Media investigation reveals that Research Gold, a service advertising '100% human-written' medical peer reviews, is entirely AI-operated with fabricated staff credentials and AI-generated responses.

0 favorites 0 likes
#peer-review

Position: Want Better ML Reviews? Stop Asking Nicely and Start Incentivizing with a Credit System

arXiv cs.AI · 2026-08-18 Cached

This position paper argues that improving machine learning peer review requires enforceable procedural safeguards and a spendable credit system, such as OpenReview Points, to incentivize good reviewing and limit submission volume.

0 favorites 0 likes
#peer-review

AC comment and our reply disappeared on OpenReview [D]

Reddit r/MachineLearning · 2026-08-15

A user on OpenReview noticed that an Area Chair's comment and their reply have disappeared, raising concerns about potential bias in paper rejection decisions. They are seeking confirmation from others if this is a normal occurrence.

0 favorites 0 likes
#peer-review

AAAI 2027 Review: No code submission? [D]

Reddit r/MachineLearning · 2026-08-11

A reviewer for AAAI 2027 expresses surprise at the low number of paper submissions with code, despite the conference's emphasis on reproducibility, and asks for opinions on whether lack of code should affect review scores.

0 favorites 0 likes
#peer-review

Peer review is overwhelmed—can it survive in the AI era?

Ars Technica · 2026-08-10 Cached

The article examines how the peer review system is struggling to cope with the exponential growth of research publications and AI-assisted papers, leaving volunteer reviewers overwhelmed and leading to errors and delays, prompting calls for reform.

0 favorites 0 likes
#peer-review

How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

Hugging Face Daily Papers · 2026-08-10 Cached

This paper investigates how rhetorical framing biases AI-based peer review scores, finding that evidence framing and novelty stance have the largest effects and that score movements depend on the reviewer's initial score and evaluation strictness.

0 favorites 0 likes
#peer-review

Large Language Models Threaten Double-blind Review

arXiv cs.CL · 2026-08-07 Cached

This paper demonstrates that large language models can effectively deanonymize authors of scientific papers from titles and abstracts alone, threatening the validity of double-blind peer review. The authors argue that stable patterns in problem framing and research focus act as latent conceptual signatures of authorship, necessitating a re-evaluation of anonymity practices in AI-augmented research ecosystems.

0 favorites 0 likes
#peer-review

RubricReviewer: From Direct Critique to Objective and Comprehensive Rubric-Driven Peer Review

arXiv cs.CL · 2026-08-04 Cached

Introduces RubricReviewer, a rubric-driven framework for LLM-based peer review that explicitly generates paper-adaptive rubrics and combines a training-free evidence-gathering agent (Scout) with a trained human-aligned model (Aligner) to produce more comprehensive, discriminative, and robust reviews.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback