reproducibility

Tag

Cards List
#reproducibility

What should remain inspectable when the agent runtime is managed?

Reddit r/AI_Agents · 4h ago

OpenAI's Agents API introduces a managed agent runtime, raising questions about which artifacts should remain inspectable to ensure reproducibility and trust in such systems.

0 favorites 0 likes
#reproducibility

Pacing the Frontier – Tahuna: AI Training Infrastructure, Now Open Source [P]

Reddit r/MachineLearning · 14h ago

Tahuna is an open-source AI training infrastructure designed for small teams to train models, run inference, orchestrate GPUs, and experiment with autonomous research, featuring reproducible runs and tools like Hillclimb.

0 favorites 0 likes
#reproducibility

XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?

arXiv cs.AI · 3d ago Cached

This paper introduces XAI-Arena, an LLM-as-a-judge framework for scalable and reproducible evaluation of explainable AI explanation quality, showing strong correlation with human judgments.

0 favorites 0 likes
#reproducibility

Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers

arXiv cs.CL · 3d ago Cached

The article introduces AgentActionBench, a process-oriented benchmark for evaluating LLM agents in reproducing experiments from scientific papers across ML and AI4Science domains.

0 favorites 0 likes
#reproducibility

Announcing the first Guix-Science release

Lobsters Hottest · 4d ago Cached

The first release of the Guix-Science channel provides a dedicated, community-driven scientific software catalog for the Guix package manager, enhancing reproducibility and collaboration in scientific computing.

0 favorites 0 likes
#reproducibility

Reproducibility seems to be headed towards irrelevance in ML research. Is it too late? [D]

Reddit r/MachineLearning · 2026-09-06

An opinion piece arguing that reproducibility in machine learning research is becoming a lost cause due to the rise of physical AI requiring expensive hardware, unverifiable performance claims from big tech companies, and competitive incentives that discourage authors from sharing code.

0 favorites 0 likes
#reproducibility

6.5% of the Neuro-Symbolic Literature Can Be Reproduced from Its Published Artifacts, a Six-Stage Audit Framework and First Instantiation

arXiv cs.AI · 2026-08-28 Cached

The paper presents a six-stage audit framework for assessing reproducibility in neuro-symbolic AI literature, finding only 6.5% of studies with published artifacts can be reproduced, highlighting a crisis in research reproducibility.

0 favorites 0 likes
#reproducibility

I made an LLM test you can clone and break

Reddit r/artificial · 2026-08-27 Cached

A reproducible test system for language models that evaluates continuation gating based on risk thresholds, demonstrating consistent behavior across frontier models like GPT-5.4 and GPT-5.6-sol.

0 favorites 0 likes
#reproducibility

Flower Hub: A Reproducible Benchmarking Platform for Federated Learning in Simulation and Deployment

arXiv cs.LG · 2026-08-27 Cached

Flower Hub is a reproducible benchmarking platform for federated learning that enables execution and evaluation across both simulation and deployment runtimes.

0 favorites 0 likes
#reproducibility

Provenance Before Prose: Claim-Locked Reporting

arXiv cs.CL · 2026-08-27 Cached

This paper proposes claim-locked reporting, a provenance-before-prose protocol that fixes statistical evidence before LLM generation to improve reproducibility and accuracy in scientific reports.

0 favorites 0 likes
#reproducibility

A Reproducible, License-Aware Distillation Recipe for CPUDeployable Safety Classification

arXiv cs.AI · 2026-08-25 Cached

This paper presents a reproducible, license-aware knowledge distillation method for creating efficient safety classifiers for large language models that can run on CPU hardware, achieving performance comparable to larger models while reducing false alarms.

0 favorites 0 likes
#reproducibility

Software Frameworks for Explainable AI in Time Series Classification: A Systematic Review

arXiv cs.AI · 2026-08-25 Cached

This paper presents a systematic review of software frameworks for explainable AI in time series classification, comparing their features, evaluation practices, and limitations.

0 favorites 0 likes
#reproducibility

AAAI 2027 Reviewer Bidding and Assignment Integrity [D]

Reddit r/MachineLearning · 2026-08-24

The article discusses collusion in the AAAI 2027 review process, particularly in reviewer assignment cycles, and critiques the lack of code publication in accepted papers at top AI conferences.

0 favorites 0 likes
#reproducibility

Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA

Hugging Face Daily Papers · 2026-08-24 Cached

This paper introduces accuracy-blind answer churn in retrieval-augmented QA systems and proposes the Snapshot Compatibility Audit to detect hidden answer changes when the corpus is updated, even if overall accuracy appears stable.

0 favorites 0 likes
#reproducibility

Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis

arXiv cs.AI · 2026-08-21 Cached

Brain Researcher is an agentic platform that enhances AI-driven neuroimaging data analysis by enforcing analytic rigor, improving tool selection accuracy, and ensuring reproducibility. The study demonstrates substantial performance gains and integrates methodological judgment into the workflow.

0 favorites 0 likes
#reproducibility

AQuA's "self-improvement" updates research state, not the agent LM. What should a local port freeze?

Reddit r/LocalLLaMA · 2026-08-20

The article analyzes AQuA's preprint on recursive self-improvement in AI agents, clarifying that the agent LM remains fixed while research state updates, and advocates for detailed ablation studies and artifact sharing to enable credible local model ports.

0 favorites 0 likes
#reproducibility

Visual-Aware Representation of Web Pages for Machine Learning Applications

arXiv cs.LG · 2026-08-20 Cached

This paper presents a platform based on FitLayout for creating visual-aware representations of web pages to support machine learning applications, demonstrating its use with graph neural networks for recognizing key content elements.

0 favorites 0 likes
#reproducibility

What would make an uploader-run refusal table independently reproducible?

Reddit r/LocalLLaMA · 2026-08-18

The article discusses the requirements for making an uploader-run refusal table in AI models independently reproducible, highlighting the need for detailed replication packets including raw generations, decoding settings, per-prompt labels, and scoring code.

0 favorites 0 likes
#reproducibility

A Reproducibility Study of Partial Residual Ablations in Pre-LN Transformers

arXiv cs.LG · 2026-08-18 Cached

A reproducibility study reveals asymmetric effects when removing residual connections in Pre-LN transformers: attention-skip removal leads to collapse, while FFN-skip removal allows partial recovery at smaller scales.

0 favorites 0 likes
#reproducibility

Agent Lightning v1.0: Towards Harnessed Agentic RL

Hugging Face Daily Papers · 2026-08-18 Cached

Agent Lightning v1.0 is a lightweight framework that enables reproducible reinforcement learning for agent harnesses, significantly boosting coding-agent performance on benchmarks like SWE-bench Verified.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback