Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR

Hugging Face Daily Papers Papers

Summary

This paper proposes GRAFT, an off-policy-aware framework for reinforcement learning with verifiable rewards (RLVR) that replaces all-fail groups with peer trajectories to improve performance in mathematical reasoning benchmarks.

Reinforcement Learning with Verifiable Rewards (RLVR) methods such as GRPO rely on successful self-generated trajectories, but finite rollout budgets can produce all-fail groups with no reward-based policy-gradient signal. While additional rollouts improve the chance of success at higher cost, successful trajectories missing from one model's rollouts may already have been discovered by another. Indeed, we observe that heterogeneous models often succeed on complementary prompts, creating opportunities for mutual learning without a designated stronger teacher. To exploit this complementarity, we propose GRAFT (Gated Replacement of Answer-Failed groups with peer Trajectories), an off-policy-aware framework that replaces all-fail groups with informative peer groups. GRAFT transfers both successful and unsuccessful peer responses with peer-computed advantages, while controlling cross-model mismatch through sequence-level compatibility weighting and token-level importance ratio clipping. Across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT consistently improves both models over GRPO with the same per-model rollout budget, gaining 2.1 points on average and up to 4.5 points in model-level average performance. Stored peer trajectories preserve most of the gains, improving over GRPO by 1.8 points on average without simultaneous co-training.
Original Article
View Cached Full Text

Cached at: 09/30/26, 04:22 AM

Paper page - Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR

Source: https://huggingface.co/papers/2609.37868

Abstract

ReinforcementLearningwithVerifiableRewards(RLVR)methodssuchasGRPOrelyonsuccessfulself-generatedtrajectories,butfiniterolloutbudgetscanproduceall-failgroupswithnoreward-basedpolicy-gradientsignal.Whileadditionalrolloutsimprovethechanceofsuccessathighercost,successfultrajectoriesmissingfromonemodel’srolloutsmayalreadyhavebeendiscoveredbyanother.Indeed,weobservethatheterogeneousmodelsoftensucceedoncomplementaryprompts,creatingopportunitiesformutuallearningwithoutadesignatedstrongerteacher.Toexploitthiscomplementarity,weproposeGRAFT(GatedReplacementofAnswer-FailedgroupswithpeerTrajectories),anoff-policy-awareframeworkthatreplacesall-failgroupswithinformativepeergroups.GRAFTtransfersbothsuccessfulandunsuccessfulpeerresponseswithpeer-computedadvantages,whilecontrollingcross-modelmismatchthroughsequence-levelcompatibilityweightingandtoken-levelimportanceratioclipping.Acrossthreeheterogeneousmodelpairsandfivemathematicalreasoningbenchmarks,GRAFTconsistentlyimprovesbothmodelsoverGRPOwiththesameper-modelrolloutbudget,gaining2.1pointsonaverageandupto4.5pointsinmodel-levelaverageperformance.Storedpeertrajectoriespreservemostofthegains,improvingoverGRPOby1.8pointsonaveragewithoutsimultaneousco-training.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2609\.37868

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.37868 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.37868 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.37868 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR

arXiv cs.AI

The paper proposes OM-GRPO, a label-free RLVR framework that decouples reward estimation from policy optimization by masking gradients on answer spans, and introduces Contrast-Augmented Reward to refine reward estimation without extra rollouts. It consistently outperforms existing label-free methods and matches supervised ground-truth reward training across reasoning benchmarks.

Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR

Hugging Face Daily Papers

This paper proposes Transfer-Aware Curriculum (TAC), a bandit-style online curriculum for multi-domain RLVR that prioritizes domains whose updates benefit other domains using gradient-geometry alignment. TAC improves macro-averaged accuracy on Qwen3-1.7B and Llama3.2-3B over fixed and learnability-only curricula.

Tandem Reinforcement Learning with Verifiable Rewards

arXiv cs.AI

Proposes Tandem Reinforcement Learning (TRL), extending the tandem training paradigm to RLVR to improve reasoning compatibility and legibility for weaker models and humans, showing that TRL matches solo performance while enhancing handoff robustness and reducing distributional drift.