reproducibility

Tag

Cards List
#reproducibility

Deterministic Replay for AI Agent Systems

arXiv cs.AI · 2026-07-21 Cached

This paper proposes a method for deterministic replay in AI agent systems, enabling reproducible debugging and analysis.

0 favorites 0 likes
#reproducibility

@classiclarryd: Question for the LLM Research Community: Is anyone aware of fully reproducible experimental results showing that MLA be…

X AI KOLs Following · 2026-07-20 Cached

A researcher questions the reproducibility of MLA outperforming GQA under same KV cache, sharing early small-scale ablation results and plans for scaling experiments to decide on architecture for next large-scale run.

0 favorites 0 likes
#reproducibility

@askalphaxiv: 70% of AI research isn’t reproducible. With ICML 2026 happening last week, over 6000+ research papers have dropped, but…

X AI KOLs Following · 2026-07-15 Cached

Alphaxiv and Hugging Face launch a community challenge to test the reproducibility of AI research papers from ICML 2026, offering $4500 in GPU credits and an autoresearch agent to help participants.

0 favorites 0 likes
#reproducibility

Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction

arXiv cs.CL · 2026-07-15 Cached

This paper presents the first systematic meta-evaluation of LLM-generated rubrics for reproducing experiments from research papers. It reformulates rubrics into a checklist format and evaluates generation settings both intrinsically (semantic similarity) and extrinsically (score alignment), finding that augmented settings improve downstream evaluation alignment but generated rubrics are often overly fine-grained and biased toward high scores.

0 favorites 0 likes
#reproducibility

How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks

arXiv cs.AI · 2026-07-15 Cached

This paper analyzes how many tasks are needed in partial evaluations of LLM agent benchmarks to reach the same pairwise conclusions as full benchmarks. It finds that required task fractions vary sharply across benchmarks and suggests reporting standards for partial evaluations.

0 favorites 0 likes
#reproducibility

From Geometric Recovery to Causal Validation: A Reproducible Audit of Sparse Autoencoder Features, from Superposition Geometry to Causal Inertness

arXiv cs.LG · 2026-07-15 Cached

This paper audits sparse autoencoder features using causal interventions, finding that up to 77% of correlationally recovered features in degraded SAEs and 9% in well-trained ones are causally inert. The authors introduce the sae-causal-audit tool for reproducible evaluation.

0 favorites 0 likes
#reproducibility

Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA

arXiv cs.LG · 2026-07-10 Cached

This paper presents Constitutional Meta-STPA, a self-validating LLM-assisted hazard analysis tool that applies STPA to itself to derive governance principles. It demonstrates that a frontier model ensemble recovers most principles and improves safety scores on adversarial probes.

0 favorites 0 likes
#reproducibility

Are AI agents reintroducing problems software engineering already solved?

Reddit r/ArtificialInteligence · 2026-07-07

The article explores how AI agent workflows are reintroducing software engineering challenges around reproducibility, auditability, and state management that were previously solved with version control, CI/CD, and static code practices, while noting emerging solutions like GitHub's Agentic Workflows and git-native approaches.

0 favorites 0 likes
#reproducibility

Auditing the Audit: Five Failure Modes in Benchmark-Validity Audits

arXiv cs.LG · 2026-07-07 Cached

This paper identifies five failure modes in perturbation-based benchmark-validity audits used for AI governance, demonstrating that implementation details can silently manufacture conclusions. It proposes a due-diligence gate to improve the reliability of evaluation evidence.

0 favorites 0 likes
#reproducibility

VERITAS: Towards a General-Purpose Replication Tool for Scientific Research

arXiv cs.AI · 2026-07-07 Cached

Veritas is a general-purpose replication tool that uses CLI coding agents to automatically replicate scientific research from papers and code, extracting claims and evaluating them, achieving state-of-the-art performance on benchmarks.

0 favorites 0 likes
#reproducibility

@SynScience: Introducing OpenScience. A better, open-source Claude Science. • Any model: GLM, Kimi, DeepSeek, Claude, GPT, your own …

X AI KOLs Following · 2026-07-05 Cached

OpenScience is an open-source alternative to Claude Science, supporting multiple AI models and over 250 research skills, with native Atlas integration for reproducible research graphs and user-controlled infrastructure.

0 favorites 0 likes
#reproducibility

I audited our autonomous research agent's 32 published "findings." 0 were novel as framed — the labels failed way more than the measurements did.

Reddit r/AI_Agents · 2026-07-05

An auditor finds that 0 of 32 published "discoveries" by an autonomous research agent were novel as framed, with the dominant failure being over-labeling rather than bad measurement.

0 favorites 0 likes
#reproducibility

@FinanceYF5: Claude Science: A new application built for the complete scientific research workflow. Research outputs can be traced back to the corresponding code, runtime environments can be managed on demand, and it can connect to over 60 optional scientific databases. Beta testing is now open.

X AI KOLs Timeline · 2026-07-03 Cached

Claude Science is a new application for the complete scientific research workflow, supporting code traceability, environment management, and connectivity to over 60 databases. Beta testing is now open.

0 favorites 0 likes
#reproducibility

Know Your Source: A Public Knowledge Store for Media Background Checks

arXiv cs.CL · 2026-07-03 Cached

This paper introduces MediaRef, a publicly available knowledge store of web-sourced documents for reproducible, low-cost evaluation of media background check generation using LLMs, addressing the limitation of costly proprietary APIs.

0 favorites 0 likes
#reproducibility

EPC: A Standardized Protocol for Measuring Evaluator Preference Dynamics in LLM Agent Systems

arXiv cs.LG · 2026-07-02 Cached

This paper introduces EPC, a standardized protocol for measuring evaluator preference coupling in LLM agent systems, including a reference snapshot and versioning convention to address reproducibility and measurement decay.

0 favorites 0 likes
#reproducibility

Anthropic’s Claude Science bets on workflow, not a new model, to win over scientists

TechCrunch AI · 2026-06-30 Cached

Anthropic launched Claude Science, an AI workbench that provides scientists with a unified environment for computational research, including connections to over 60 databases and prebuilt toolkits, without introducing a new AI model.

0 favorites 0 likes
#reproducibility

Open but Incompatible: A License Compatibility Analysis of Corpora for Low-Resource African Languages

arXiv cs.CL · 2026-06-30 Cached

This paper audits license provenance of over twenty African NLP corpus families, identifies compatibility failures like the JW300 violation and hidden NoDerivs clauses, and provides a due diligence checklist for legally clean dataset creation.

0 favorites 0 likes
#reproducibility

Agentic Publication Protocol: An Attempt to Modernize Scientific Publication

arXiv cs.AI · 2026-06-29 Cached

This paper introduces the Agentic Publication Protocol (APP), a format that packages scientific papers with code, data, and agent instructions to improve reproducibility and allow AI agents to interact with and build upon published research.

0 favorites 0 likes
#reproducibility

Mitigating LLM-based p-Hacking by Preregistering for the Next LLM

arXiv cs.CL · 2026-06-29 Cached

Proposes a protocol to mitigate p-hacking in LLM-based research by preregistering experiments and running them on the first eligible model released after preregistration, demonstrating substantial mitigation across multiple models.

0 favorites 0 likes
#reproducibility

@jakevin7: Let me make a prediction: The next phase of the AI era will become "Infra is all you need". AI-generated code is already very powerful, but it's still far from adequate in terms of usability and stability. Recently, OpenAI's subscription system had a huge bug, and the membership system completely broke down. The system…

X AI KOLs Following · 2026-06-26 Cached

The author predicts that the next phase of the AI era will shift from model capabilities to infrastructure capabilities, emphasizing infra abilities such as reproducibility, observability, recoverability, and security isolation, believing that stably carrying AI behavior will be the key to competition.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback