Tag
Introduces Supersede, an environment to diagnose and train the memory-update gap in LLM agents, showing that standard models fail to maintain current facts as conversations grow, and that GRPO fine-tuning can improve performance.