value-calibration

Tag

Cards List
#value-calibration

Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search

arXiv cs.LG · 2026-07-31 Cached

This paper analyzes contrastive critics used as value-like objectives in reinforcement learning, showing that good ranking accuracy does not make them safe to maximize due to off-support norm inflation and misranking, and demonstrates that value-calibrated scalar critics like TD-Q succeed where contrastive critics fail.

0 favorites 0 likes
← Back to home

Submit Feedback