kl-regularization

Tag

Cards List
#kl-regularization

Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning

arXiv cs.LG · yesterday Cached

This paper introduces a game-theoretic approach to fine-tuning language models that optimizes the trade-off between reward and deviating from a reference policy, providing a principled method for setting the KL regularization coefficient.

0 favorites 0 likes
#kl-regularization

Safe Inference-Time Alignment via Lagrangian Reward Augmentation

arXiv cs.LG · 2026-07-07 Cached

Proposes LARA, a framework for safe inference-time alignment that uses Lagrangian dualization to derive an augmented reward from separate reward and cost models, improving the helpfulness-harmlessness tradeoff without retraining.

0 favorites 0 likes
#kl-regularization

Self-Distilled Policy Gradient

Hugging Face Daily Papers · 2026-06-02 Cached

This paper proposes SDPG, a self-distilled policy-gradient framework that combines on-policy self-distillation with verifier advantages and KL regularization to improve reinforcement learning stability and performance.

0 favorites 0 likes
← Back to home

Submit Feedback