Tag
This paper proposes GRAFT, an off-policy-aware framework for reinforcement learning with verifiable rewards (RLVR) that replaces all-fail groups with peer trajectories to improve performance in mathematical reasoning benchmarks.