reward-filtering

Tag

Cards List
#reward-filtering

Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress

arXiv cs.AI · 2026-08-21 Cached

The paper proposes Reasoning-Progress-Aware Reward Filtering for On-Policy Distillation (R2-OPD) to address the mismatch between teacher-derived rewards and genuine reasoning progress, improving reasoning performances in language model training.

0 favorites 0 likes
← Back to home

Submit Feedback