Tag
This paper designs a four-tier experimental teaching system for multimodal medical image intelligent diagnosis, translating research into undergraduate labs to address gaps in education for clinical AI.
This paper adapts the Music Lab experiment to study how social influence affects AI agents' selection of scientific papers, showing that social information reduces attention volume and breadth while increasing between-community variation.
The paper introduces Odds-Ratio Thompson Sampling (OR-TS), a method for batched multi-armed bandits that uses contrast-based updates to handle varying common levels, showing improved regret over absolute-rate memory in simulations and real experiments.
This paper introduces Target-Weighted Neyman Allocation (TWNA), a two-stage stratified experimental design that optimizes sample allocation across groups and treatment arms to improve precision of target-weighted group average treatment effects under population shift.
Tweet promoting the third edition of the statistics textbook 'Designing Experiments and Analyzing Data: A Model Comparison Perspective' by Maxwell, Delaney, and Kelley, highlighting its pedagogical features and the authors' academic credentials.
This paper presents PC-MCMC-CIGP, a gray-box workflow that combines spike-and-slab topology sampling with physical constraints and a Chemical-Informed Gaussian Process for reaction network discovery. The method demonstrates improved yield on styrene epoxidation and distinguishes elementary pathways from deceptive fits on a hydrogen-bromine benchmark.
This paper studies a staged promotion protocol for micro-pretraining, using escalating budgets from minutes to hours to filter configurations. It finds that early screens are useful but unstable, and that a staged approach can retain a long-horizon reference while identifying alternatives that fail continuation thresholds.
This paper introduces Cartograph, a verification layer for AI scientists that couples subspace experiment steering, ambiguity resolution, and library inadequacy detection. The framework outperforms baselines in autonomous discovery testbeds and retrospectively flags inconclusive claims in the A-Lab materials system.
This paper proposes a staged factorial screening workflow for budget-constrained micro-pretraining, demonstrating that short designed experiments can identify stable hyperparameter penalty directions and support a screen-then-refine strategy.
ScienceClaw is an AI assistant framework integrating 285 research skills, modularizing the entire research workflow into Skills. It supports connecting to databases such as PubMed, Semantic Scholar, and ArXiv, providing functions like literature search, paper deep reading, citation analysis, experimental design assistance, and writing assistance. It is suitable for advanced users who want deep customization.
This paper evaluates whether off-the-shelf LLMs can reliably simulate human responses in controlled behavioral experiments by comparing LLM-generated data with human survey responses on accuracy perception. The findings show that while LLMs capture directional effects and aggregate belief-updating patterns, they do not consistently match human-scale effect magnitudes, clarifying when synthetic LLM data can serve as behavioral proxies.