Tag
A nonprofit named Reflective has published a roadmap outlining the experiments, studies, and infrastructure needed to make informed decisions about deploying solar geoengineering, specifically stratospheric aerosol injection, to address climate change.
After facing methodology criticisms, the author conducted extensive experiments on an LLM resume-screening study, revealing that initial bias measurements were largely due to noise and that factors like prompt design and wrapper choice significantly affect scores.
Yohei Nakajima updated his build-in-public log with recent projects including synthetic players research, mobile agent relay, and an event-sourced reactive graph runtime for long-running agents.
This paper introduces AEROBAT, the first multi-agent system to automate behavioral scientific research on AI agents, generating hypotheses, designing and executing controlled experiments, and writing reports. The authors demonstrate its efficacy across 12 target behaviors, finding statistical evidence for 26 hypotheses.
Quanta Magazine recounts the 70-year history of neutrino detection, from Pauli's postulate to massive experiments like Super-Kamiokande and IceCube that solved the solar neutrino problem and revealed neutrino oscillation.
Tweet discussing advice on self-improving agents, with personal observations from experiments on coding agents for long-horizon tasks, noting that stronger models don't always yield better agents.
An AI researcher ran 13 controlled experiments on a multi-agent coding system, finding that dependency-ordered coordination significantly improved success rates while persona backstories had no measurable benefit.
Codex is writing a blogpost about its experiments in training a model autonomously.