Tag
The AI Guide partnership updates to v7 with more honest reflections on unresolved contradictions and methodology, including re-evaluating a key term's valence and a metaphor's effectiveness, plus an outside review clarifying model sentiment vs. training data bias.
This paper introduces AI agents that reproduce human analytical variation across datasets, finding that different personas lead to divergent conclusions. It proposes the m-value and Agentic Bootstrap to assess the credibility of reported analyses.
A detailed guide on designing effective ML experiments, emphasizing starting with a clear research question, developing research taste, and scaling results. Based on the author's experience running ~100 experiments weekly at Poolside.
Proposes a protocol to mitigate p-hacking in LLM-based research by preregistering experiments and running them on the first eligible model released after preregistration, demonstrating substantial mitigation across multiple models.
This article shares a piece written by Anthropic researcher Vivek Nair on how to conduct good research, emphasizing that choosing the right problem is more important than solving it, cultivating research taste, upgrading information sources, and using writing as a thinking tool. These insights are not only applicable to academic research but also provide inspiration for career development and investment decisions.
This study examines how LLMs suggest research methods (datasets, models, metrics) when prompted only with a research question, finding that LLMs exhibit a strong provider bias and propose a much narrower range of methods compared to actual papers, potentially narrowing researchers' methodological search space.
Medical students are reportedly using a popular research tool to produce misleading studies, raising concerns about academic integrity and research quality.
This blog post argues for a return to rigorous full-system timing simulation in computer architecture to overcome the 'timing simulation wall' and accurately capture modern system behaviors, advocating for measuring the right execution intervals with statistically sound methods rather than simulating everything in detail.
A methodology article on how to excel at AI research, emphasizing problem selection, literature reading, writing notes, and other skills, suitable for researchers.
This paper introduces PEEL (Protocols for Epistemically Engaged Literacy in AI), a framework combining deterministic text analysis via Voyant Tools with LLM interpretation via Claude, grounded in Peircean semiotics, to expose systematic distortions in AI-generated research summaries and promote epistemic accountability.
yibie shared three new entries from the awesome-autoresearch list, covering automated quantitative trading, universal skill optimization, and a Claude Code plugin.
Yann LeCun contrasts the engineering mindset (focused on product innovation and shipping good-enough solutions) with the scientific mindset (focused on asking new questions, proposing solutions, and rigorous methodology), noting both are activities rather than identities and that product innovation builds on earlier scientific advances.
This paper from Google DeepMind and Carnegie Mellon argues that LLM-simulated experiments are actually observational studies due to user drift, where the simulated population shifts with interventions. The authors propose using negative control outcomes to diagnose confounding and show that eliciting setting-relevant confounders can reduce bias.
An experiment running the same research prompt about LENR and superconductivity through six AI systems in five languages reveals significant linguistic bias, with non-English queries surfacing information about real industrial commitments that English-only searches miss.
This position paper argues that machine learning research should prioritize ideas over benchmarks and theoretical guarantees, proposing an 'Ideas First' framework that values behavioral signatures and tailored experiments to promote equity and scientific understanding.