标签
This paper analyzes contrastive critics used as value-like objectives in reinforcement learning, showing that good ranking accuracy does not make them safe to maximize due to off-support norm inflation and misranking, and demonstrates that value-calibrated scalar critics like TD-Q succeed where contrastive critics fail.