Tag
This paper introduces a game-theoretic approach to fine-tuning language models that optimizes the trade-off between reward and deviating from a reference policy, providing a principled method for setting the KL regularization coefficient.
Proposes LARA, a framework for safe inference-time alignment that uses Lagrangian dualization to derive an augmented reward from separate reward and cost models, improving the helpfulness-harmlessness tradeoff without retraining.
This paper proposes SDPG, a self-distilled policy-gradient framework that combines on-policy self-distillation with verifier advantages and KL regularization to improve reinforcement learning stability and performance.