Tag
This paper proposes a method for deterministic replay in AI agent systems, enabling reproducible debugging and analysis.
A researcher questions the reproducibility of MLA outperforming GQA under same KV cache, sharing early small-scale ablation results and plans for scaling experiments to decide on architecture for next large-scale run.
Alphaxiv and Hugging Face launch a community challenge to test the reproducibility of AI research papers from ICML 2026, offering $4500 in GPU credits and an autoresearch agent to help participants.
This paper presents the first systematic meta-evaluation of LLM-generated rubrics for reproducing experiments from research papers. It reformulates rubrics into a checklist format and evaluates generation settings both intrinsically (semantic similarity) and extrinsically (score alignment), finding that augmented settings improve downstream evaluation alignment but generated rubrics are often overly fine-grained and biased toward high scores.
This paper analyzes how many tasks are needed in partial evaluations of LLM agent benchmarks to reach the same pairwise conclusions as full benchmarks. It finds that required task fractions vary sharply across benchmarks and suggests reporting standards for partial evaluations.
This paper audits sparse autoencoder features using causal interventions, finding that up to 77% of correlationally recovered features in degraded SAEs and 9% in well-trained ones are causally inert. The authors introduce the sae-causal-audit tool for reproducible evaluation.
This paper presents Constitutional Meta-STPA, a self-validating LLM-assisted hazard analysis tool that applies STPA to itself to derive governance principles. It demonstrates that a frontier model ensemble recovers most principles and improves safety scores on adversarial probes.
The article explores how AI agent workflows are reintroducing software engineering challenges around reproducibility, auditability, and state management that were previously solved with version control, CI/CD, and static code practices, while noting emerging solutions like GitHub's Agentic Workflows and git-native approaches.
This paper identifies five failure modes in perturbation-based benchmark-validity audits used for AI governance, demonstrating that implementation details can silently manufacture conclusions. It proposes a due-diligence gate to improve the reliability of evaluation evidence.
Veritas is a general-purpose replication tool that uses CLI coding agents to automatically replicate scientific research from papers and code, extracting claims and evaluating them, achieving state-of-the-art performance on benchmarks.
OpenScience is an open-source alternative to Claude Science, supporting multiple AI models and over 250 research skills, with native Atlas integration for reproducible research graphs and user-controlled infrastructure.
An auditor finds that 0 of 32 published "discoveries" by an autonomous research agent were novel as framed, with the dominant failure being over-labeling rather than bad measurement.
Claude Science is a new application for the complete scientific research workflow, supporting code traceability, environment management, and connectivity to over 60 databases. Beta testing is now open.
This paper introduces MediaRef, a publicly available knowledge store of web-sourced documents for reproducible, low-cost evaluation of media background check generation using LLMs, addressing the limitation of costly proprietary APIs.
This paper introduces EPC, a standardized protocol for measuring evaluator preference coupling in LLM agent systems, including a reference snapshot and versioning convention to address reproducibility and measurement decay.
Anthropic launched Claude Science, an AI workbench that provides scientists with a unified environment for computational research, including connections to over 60 databases and prebuilt toolkits, without introducing a new AI model.
This paper audits license provenance of over twenty African NLP corpus families, identifies compatibility failures like the JW300 violation and hidden NoDerivs clauses, and provides a due diligence checklist for legally clean dataset creation.
This paper introduces the Agentic Publication Protocol (APP), a format that packages scientific papers with code, data, and agent instructions to improve reproducibility and allow AI agents to interact with and build upon published research.
Proposes a protocol to mitigate p-hacking in LLM-based research by preregistering experiments and running them on the first eligible model released after preregistration, demonstrating substantial mitigation across multiple models.
The author predicts that the next phase of the AI era will shift from model capabilities to infrastructure capabilities, emphasizing infra abilities such as reproducibility, observability, recoverability, and security isolation, believing that stably carrying AI behavior will be the key to competition.