value-calibration

标签

Cards List
#value-calibration

Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search

arXiv cs.LG · 2026-07-31 缓存

This paper analyzes contrastive critics used as value-like objectives in reinforcement learning, showing that good ranking accuracy does not make them safe to maximize due to off-support norm inflation and misranking, and demonstrates that value-calibrated scalar critics like TD-Q succeed where contrastive critics fail.

0 人收藏 0 人点赞
← 返回首页

提交意见反馈