Tag
CMU PhD student Emmy Liu reflects on the evolution of ML/NLP research, from manual work to AI-assisted to the Agent era, discussing the value of researchers in the new age.
This paper explores using human–LLM disagreement to enhance checklist-based quality appraisal in systematic reviews, showing that analyzing disagreements can identify ambiguous items and improve checklist design for better agreement and study ranking preservation.
The article analyzes AQuA's preprint on recursive self-improvement in AI agents, clarifying that the agent LM remains fixed while research state updates, and advocates for detailed ablation studies and artifact sharing to enable credible local model ports.
The article critiques common practices in AI research to make sparse attention and KV cache compression techniques appear more effective than they might be, by exploiting benchmark flaws and implementation details.
Tweet promoting the third edition of the statistics textbook 'Designing Experiments and Analyzing Data: A Model Comparison Perspective' by Maxwell, Delaney, and Kelley, highlighting its pedagogical features and the authors' academic credentials.
The AI Guide partnership updates to v7 with more honest reflections on unresolved contradictions and methodology, including re-evaluating a key term's valence and a metaphor's effectiveness, plus an outside review clarifying model sentiment vs. training data bias.
This paper introduces AI agents that reproduce human analytical variation across datasets, finding that different personas lead to divergent conclusions. It proposes the m-value and Agentic Bootstrap to assess the credibility of reported analyses.
A detailed guide on designing effective ML experiments, emphasizing starting with a clear research question, developing research taste, and scaling results. Based on the author's experience running ~100 experiments weekly at Poolside.
Proposes a protocol to mitigate p-hacking in LLM-based research by preregistering experiments and running them on the first eligible model released after preregistration, demonstrating substantial mitigation across multiple models.
This article shares a piece written by Anthropic researcher Vivek Nair on how to conduct good research, emphasizing that choosing the right problem is more important than solving it, cultivating research taste, upgrading information sources, and using writing as a thinking tool. These insights are not only applicable to academic research but also provide inspiration for career development and investment decisions.
This study examines how LLMs suggest research methods (datasets, models, metrics) when prompted only with a research question, finding that LLMs exhibit a strong provider bias and propose a much narrower range of methods compared to actual papers, potentially narrowing researchers' methodological search space.
Medical students are reportedly using a popular research tool to produce misleading studies, raising concerns about academic integrity and research quality.
This blog post argues for a return to rigorous full-system timing simulation in computer architecture to overcome the 'timing simulation wall' and accurately capture modern system behaviors, advocating for measuring the right execution intervals with statistically sound methods rather than simulating everything in detail.
A methodology article on how to excel at AI research, emphasizing problem selection, literature reading, writing notes, and other skills, suitable for researchers.
This paper introduces PEEL (Protocols for Epistemically Engaged Literacy in AI), a framework combining deterministic text analysis via Voyant Tools with LLM interpretation via Claude, grounded in Peircean semiotics, to expose systematic distortions in AI-generated research summaries and promote epistemic accountability.
yibie shared three new entries from the awesome-autoresearch list, covering automated quantitative trading, universal skill optimization, and a Claude Code plugin.
Yann LeCun contrasts the engineering mindset (focused on product innovation and shipping good-enough solutions) with the scientific mindset (focused on asking new questions, proposing solutions, and rigorous methodology), noting both are activities rather than identities and that product innovation builds on earlier scientific advances.
This paper from Google DeepMind and Carnegie Mellon argues that LLM-simulated experiments are actually observational studies due to user drift, where the simulated population shifts with interventions. The authors propose using negative control outcomes to diagnose confounding and show that eliciting setting-relevant confounders can reduce bias.
An experiment running the same research prompt about LENR and superconductivity through six AI systems in five languages reveals significant linguistic bias, with non-English queries surfacing information about real industrial commitments that English-only searches miss.
This position paper argues that machine learning research should prioritize ideas over benchmarks and theoretical guarantees, proposing an 'Ideas First' framework that values behavioral signatures and tailored experiments to promote equity and scientific understanding.