Tag
TraceML is a tool for trajectory-level analysis of ML research agents, comparing Codex agent runs with Kaggle Grandmaster notebooks and finding that human experts explore broader strategy sets while agents collapse into narrow loops. Distilling human research skills into a ~1,000-token planning prompt lifted Codex scores on five of seven competitions; the work will be presented at NeurIPS 2026.
Maxime Rivest shares tips for optimizing Pi coding agent usage with custom system prompts, and Matei Zaharia discusses surprising results from benchmarking coding agents at Databricks.
Introduces AgentOdyssey, a procedural text game generation framework designed to evaluate agents on test-time continual learning abilities including exploration, episodic memory, world knowledge acquisition, skill learning, and long-horizon planning. The framework highlights significant gaps between current agents and human performance.
The paper introduces RealUserSim, a framework that grounds LLM-based user simulation in real human behavioral data from 14,000+ authentic conversations to bridge the reality gap in agent benchmarking. It shows that grounded simulation raises behavioral match rates from 24.2% to 45.3% and reveals failure mechanisms invisible to cooperative simulators.