Tag
Yohei Nakajima updated his build-in-public log with recent projects including synthetic players research, mobile agent relay, and an event-sourced reactive graph runtime for long-running agents.
This paper introduces AEROBAT, the first multi-agent system to automate behavioral scientific research on AI agents, generating hypotheses, designing and executing controlled experiments, and writing reports. The authors demonstrate its efficacy across 12 target behaviors, finding statistical evidence for 26 hypotheses.
Quanta Magazine recounts the 70-year history of neutrino detection, from Pauli's postulate to massive experiments like Super-Kamiokande and IceCube that solved the solar neutrino problem and revealed neutrino oscillation.
Tweet discussing advice on self-improving agents, with personal observations from experiments on coding agents for long-horizon tasks, noting that stronger models don't always yield better agents.
An AI researcher ran 13 controlled experiments on a multi-agent coding system, finding that dependency-ordered coordination significantly improved success rates while persona backstories had no measurable benefit.
Codex is writing a blogpost about its experiments in training a model autonomously.