Tag
This paper introduces invalidation contracts as a protocol for LLM agents to manage cached recovery suggestions across episodes, addressing server-side data drift and improving token efficiency by evicting stale entries while maintaining high compliance rates.
Introduces Self-Review Reinforcement Learning (SRRL), a training framework that embeds a self-review step into RL episodes to transform sparse environmental feedback into behavioral improvements, using cross-episode memory and selective policy distillation. Evaluated on GSM8K, SRRL outperforms standard RLVR with GRPO across Qwen 3-4B and OLMo-3-7B models.