verifier-advantages

Tag

Cards List
#verifier-advantages

Self-Distilled Policy Gradient

Hugging Face Daily Papers · 2026-06-02 Cached

This paper proposes SDPG, a self-distilled policy-gradient framework that combines on-policy self-distillation with verifier advantages and KL regularization to improve reinforcement learning stability and performance.

0 favorites 0 likes
← Back to home

Submit Feedback