Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR
Summary
This paper proposes GRAFT, an off-policy-aware framework for reinforcement learning with verifiable rewards (RLVR) that replaces all-fail groups with peer trajectories to improve performance in mathematical reasoning benchmarks.
View Cached Full Text
Cached at: 09/30/26, 04:22 AM
Paper page - Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR
Source: https://huggingface.co/papers/2609.37868
Abstract
ReinforcementLearningwithVerifiableRewards(RLVR)methodssuchasGRPOrelyonsuccessfulself-generatedtrajectories,butfiniterolloutbudgetscanproduceall-failgroupswithnoreward-basedpolicy-gradientsignal.Whileadditionalrolloutsimprovethechanceofsuccessathighercost,successfultrajectoriesmissingfromonemodel’srolloutsmayalreadyhavebeendiscoveredbyanother.Indeed,weobservethatheterogeneousmodelsoftensucceedoncomplementaryprompts,creatingopportunitiesformutuallearningwithoutadesignatedstrongerteacher.Toexploitthiscomplementarity,weproposeGRAFT(GatedReplacementofAnswer-FailedgroupswithpeerTrajectories),anoff-policy-awareframeworkthatreplacesall-failgroupswithinformativepeergroups.GRAFTtransfersbothsuccessfulandunsuccessfulpeerresponseswithpeer-computedadvantages,whilecontrollingcross-modelmismatchthroughsequence-levelcompatibilityweightingandtoken-levelimportanceratioclipping.Acrossthreeheterogeneousmodelpairsandfivemathematicalreasoningbenchmarks,GRAFTconsistentlyimprovesbothmodelsoverGRPOwiththesameper-modelrolloutbudget,gaining2.1pointsonaverageandupto4.5pointsinmodel-levelaverageperformance.Storedpeertrajectoriespreservemostofthegains,improvingoverGRPOby1.8pointsonaveragewithoutsimultaneousco-training.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2609\.37868
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.37868 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.37868 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.37868 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR
The paper proposes OM-GRPO, a label-free RLVR framework that decouples reward estimation from policy optimization by masking gradients on answer spans, and introduces Contrast-Augmented Reward to refine reward estimation without extra rollouts. It consistently outperforms existing label-free methods and matches supervised ground-truth reward training across reasoning benchmarks.
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
This paper proposes RLSVR, a task-transformation paradigm that extends reinforcement learning with verifiable rewards to open-ended LLM tasks by creating self-verifiable proxy environments, instantiated via the SpyRL multi-agent self-play framework, showing gains on summarization, creative writing, and math reasoning.
StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction
StraTA proposes strategic trajectory abstraction for long-horizon LLM agents, using hierarchical GRPO-style rollout with diverse strategy sampling and critical self-judgment to improve sample efficiency and final performance over frontier models and prior RL baselines.
Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR
This paper proposes Transfer-Aware Curriculum (TAC), a bandit-style online curriculum for multi-domain RLVR that prioritizes domains whose updates benefit other domains using gradient-geometry alignment. TAC improves macro-averaged accuracy on Qwen3-1.7B and Llama3.2-3B over fixed and learnability-only curricula.
Tandem Reinforcement Learning with Verifiable Rewards
Proposes Tandem Reinforcement Learning (TRL), extending the tandem training paradigm to RLVR to improve reasoning compatibility and legibility for weaker models and humans, showing that TRL matches solo performance while enhancing handoff robustness and reducing distributional drift.