training-stabilization

Tag

Cards List
#training-stabilization

PowerOPD: Stabilizing On-Policy Distillation with Bounded Power Transformation

arXiv cs.LG · 2026-06-17 Cached

PowerOPD introduces a bounded power transformation to stabilize on-policy distillation for large language models, achieving significant gains in accuracy and sample efficiency while reducing computational cost.

0 favorites 0 likes
← Back to home

Submit Feedback