cross-model-learning

Tag

Cards List
#cross-model-learning

Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR

Hugging Face Daily Papers ↗ · yesterday Cached

This paper proposes GRAFT, an off-policy-aware framework for reinforcement learning with verifiable rewards (RLVR) that replaces all-fail groups with peer trajectories to improve performance in mathematical reasoning benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback