Tag
This paper introduces Contrastive Policy Optimization (CPO), which uses token-level contrastive disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping in reinforcement learning with verifiable rewards. CPO outperforms entropy-based RLVR methods on both in-domain and out-of-domain benchmarks.
PASS is a middleware that fixes three pathologies in process-supervised RL for LLM reasoners, improving GRPO by independently standardizing streams, chunking by value, and using average value density. It shows consistent gains in math reasoning and multi-hop QA.