Tag
This paper introduces 'Paper-replication', a workflow for coding agents that systematically replicates scientific machine learning papers by turning each claim into a target with recorded evidence, and demonstrates across twelve runs that all workspaces and targets are completed.
This paper introduces AI agents that reproduce human analytical variation across datasets, finding that different personas lead to divergent conclusions. It proposes the m-value and Agentic Bootstrap to assess the credibility of reported analyses.