Tag
This paper defines and quantifies sentiment drift in RLHF-trained summarization models, proposes a Policy Attribution framework to identify causes, and introduces a Sentiment-Aware KL Regularization method to reduce drift.