Tag
This paper investigates the calibration gap that arises when temperature scaling, a common post-hoc calibration method relying on one-hot labels, is applied to models trained with soft label distributions reflecting genuine human disagreement. Experiments across vision and language domains show that temperature scaling calibrated on hard labels consistently underperforms direct soft-label calibration, with larger gaps in language tasks, highlighting risks for safety-critical deployments.