rubric-guided-policy-optimization

Tag

Cards List
#rubric-guided-policy-optimization

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

Hugging Face Daily Papers · 3d ago Cached

CoRT proposes a token-level credit weighting method for GRPO that uses counterfactual replay to compute token-wise log-likelihood contrasts, redistributing the signed advantage across tokens without an auxiliary scorer, achieving average gains of 4.4 percentage points over response-level GRPO.

0 favorites 0 likes
← Back to home

Submit Feedback