off-policy-grpo

Tag

Cards List
#off-policy-grpo

On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers

arXiv cs.LG · 2026-09-03 Cached

This paper introduces a reinforcement learning-based distillation framework for training compact instruction-following rerankers, using off-policy GRPO for teacher enhancement and on-policy distillation for student learning, demonstrating superior performance under distribution shift.

0 favorites 0 likes
← Back to home

Submit Feedback