Tag
The paper introduces ERPO, a method that moves regularization from the action-side to the input-side by controlling query distribution, addressing the stability-exploration dilemma in LLM policy optimization, and showing improvements on mathematical reasoning benchmarks.