replay-benchmark

Tag

Cards List
#replay-benchmark

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World

arXiv cs.LG · 2d ago Cached

This paper demonstrates that replay-based static evaluation of model switching in LLM agents is fundamentally flawed: when swapping models mid-trajectory, the environment and subsequent actions diverge dramatically from logged trajectories, invalidating most benchmark results. The authors propose branching rollouts as a more faithful evaluation method and release their harness and trajectories.

0 favorites 0 likes
← Back to home

Submit Feedback