advantage-shaping

Tag

Cards List
#advantage-shaping

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

Hugging Face Daily Papers · 2026-07-16 Cached

This paper introduces Contrastive Policy Optimization (CPO), which uses token-level contrastive disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping in reinforcement learning with verifiable rewards. CPO outperforms entropy-based RLVR methods on both in-domain and out-of-domain benchmarks.

0 favorites 0 likes
#advantage-shaping

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners

arXiv cs.AI · 2026-06-30 Cached

PASS is a middleware that fixes three pathologies in process-supervised RL for LLM reasoners, improving GRPO by independently standardizing streams, chunking by value, and using average value density. It shows consistent gains in math reasoning and multi-hop QA.

0 favorites 0 likes
← Back to home

Submit Feedback