Tag
This paper argues that evaluating personal LLM agents requires replaying temporal interventions across different user-conditioned states and identifies a gap in current benchmarks. It proposes a minimal benchmark design and reporting metrics for user-conditioned adaptation.