group-relative-optimization

Tag

Cards List
#group-relative-optimization

Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents

arXiv cs.LG · 2026-05-22 Cached

Memory-R2 introduces LoGo-GRPO, a training framework that combines local and global group-relative optimization to provide fairer credit assignment for long-horizon memory-augmented LLM agents, improving accuracy and inference latency across backbones.

0 favorites 0 likes
← Back to home

Submit Feedback