agent-benchmarking

Tag

Cards List
#agent-benchmarking

@sunweiwei12: What do auto-research agents actually do differently from top human researchers? Despite all the excitement around auto…

X AI KOLs Timeline ↗ · 4d ago Cached

TraceML is a tool for trajectory-level analysis of ML research agents, comparing Codex agent runs with Kaggle Grandmaster notebooks and finding that human experts explore broader strategy sets while agents collapse into narrow loops. Distilling human research skills into a ~1,000-token planning prompt lifted Codex scores on five of seven competitions; the work will be presented at NeurIPS 2026.

0 favorites 0 likes
#agent-benchmarking

@MaximeRivest: Pi coding agent is very often the best for its price. Pi's system prompt is very short and has only 4 tools. My Pi syst…

X AI KOLs Timeline ↗ · 2026-07-09 Cached

Maxime Rivest shares tips for optimizing Pi coding agent usage with custom system prompts, and Matei Zaharia discusses surprising results from benchmarking coding agents at Databricks.

0 favorites 0 likes
#agent-benchmarking

AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents

arXiv cs.CL ↗ · 2026-06-25 Cached

Introduces AgentOdyssey, a procedural text game generation framework designed to evaluate agents on test-time continual learning abilities including exploration, episodic memory, world knowledge acquisition, skill learning, and long-horizon planning. The framework highlights significant gaps between current agents and human performance.

0 favorites 0 likes
#agent-benchmarking

RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation

arXiv cs.AI ↗ · 2026-05-22 Cached

The paper introduces RealUserSim, a framework that grounds LLM-based user simulation in real human behavioral data from 14,000+ authentic conversations to bridge the reality gap in agent benchmarking. It shows that grounded simulation raises behavioral match rates from 24.2% to 45.3% and reveals failure mechanisms invisible to cooperative simulators.

0 favorites 0 likes
← Back to home

Submit Feedback