Tag
Proposes W2SPO, an off-policy RL method that injects short auxiliary segments from a weaker model into target-model trajectories to diversify exploration and improve reasoning, achieving better performance and training speedup on math benchmarks.