Tag
CrossAudit introduces a git-native protocol for auditing autonomous research pipelines using AI agents from different vendors to ensure oversight and integrity in agentic science.
This study evaluates two multimodal LLMs as peer reviewers for ICLR 2026 submissions, finding high scoring calibration but low error detection, with author identity having no effect and figures reducing error detection.
A replication study fails to reproduce findings from a 2002 paper on procrastination and deadlines, leading to allegations of data fraud in the original influential research.
This paper introduces IntegrityBench, a benchmark for evaluating whether LLMs uphold research integrity when acting as co-scientists under institutional pressure. Findings show frontier models fail roughly 1 in 3 integrity-critical decisions under peak pressure, and that ethical action does not require accurate misconduct classification.
A tweet criticizing an ICLR 2026 paper, pointing out that an open-sourced paper achieved excellent results but actually used test set labels to select parameters, questioning the rigor of top conference peer review.
A position paper arguing that autonomous AI agents in science widen the verification gap and that scientific verification infrastructure must evolve with observable-by-default workflows, scalable verification, and clear attribution to sustain trustworthy science.
A study measures the prevalence of AI-written text on arXiv, finding that over 30% of new submissions read as machine-written, with computer science leading at 65% and mathematics lowest at 0.7%.
This paper argues that generative AI can degrade research by eroding the practices that form scholarly judgement and trust, and advocates for a renewed commitment to research as a lived practice that cannot be automated.
Springer Nature has retracted two studies conducted by researchers at the Max Planck Institute, citing concerns over the validity of the findings.
A podcast episode discusses the growing prevalence of AI hallucinations in academic papers, attributing it to poor working conditions for academics and warning of dangers to future research and knowledge production.
Mathematicians, via the Leiden Declaration endorsed by the International Mathematical Union, warn that AI threatens core values of mathematical research, including correctness, transparency, and citation practices, while also raising concerns about industry influence and the erosion of traditional standards.
A Princeton study found data leakage in nearly 300 AI papers across 17 fields, causing overoptimistic results. The author highlights how easy it is to accidentally leak data and cautions against trusting impressive AI claims without checking for leakage.
ArXiv is implementing a new policy that bans authors for one year if submitted papers show clear evidence of unchecked AI generation, such as hallucinated references or LLM comments, reinforcing that authors are fully responsible for content regardless of how it was produced.
The article criticizes the reliance on Large Language Models for generating bibliographic entries, highlighting issues with hallucinated citations and incorrect author lists in academic papers.
A researcher reports failure to reproduce claims from modern papers, with 4 out of 7 checked claims being irreproducible and 2 having unresolved GitHub issues, raising concerns about research quality standards.