Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization
Abstract
Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, entropy cannot distinguish useful uncertainty from detrimental confusion, limiting its effectiveness as a correctness signal. We propose Contrastive Policy Optimization (CPO), which uses token-level contrastive disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping. Both theoretical and empirical results show that this disagreement reliably indicates token-level correctness. We further show that On-policy Distillation is a special case of CPO, where the posterior distribution is instantiated by an external teacher model. CPO also resolves the zero-advantage problem. Experiments on in-domain and out-of-domain benchmarks demonstrate that CPO substantially outperforms entropy-based RLVR methods while maintaining strong generalization. Further analysis shows that correct and incorrect responses naturally support exploration and exploitation respectively, and balancing both leads to the best performance.
Community
We propose Contrastive Policy Optimization (CPO), a framework that uses the disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping. We provide a theoretical explanation for why this disagreement reliably indicates token-level correctness. We further show that on-policy self-distillation can be viewed as a special instantiation of CPO.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- GeneralThinker: Domain-General Reasoning through Likelihood-Guided Answer-Conditioned Optimization (2026)
- Learning from Own Solutions: Self-Conditioned Credit Assignment for Reinforcement Learning with Verifiable Rewards (2026)
- CAST: Non-Privileged Clipped Asymmetric Self-Teaching with Advantage Flipping for GRPO (2026)
- Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning (2026)
- Stabilizing On-Policy Distillation for MLLM Reasoning with Global Normalization (2026)
- RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation (2026)
- PAEC: Position-Aware Entropy Calibration for LLM Reasoning in RLVR (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.14614 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper