Tag
This paper proposes a sample-wise adaptive temperature scaling method for Transformed Teacher Matching in knowledge distillation, improving performance on image classification benchmarks by locally minimizing KL divergence between teacher and student distributions.
This paper proposes Meta-Persona Anchoring and Filtered Temperature Scaling to reduce semantic convergence in LLMs, showing a drop in pairwise cosine similarity from ~0.85 to ~0.65 on the INFINITY-CHAT dataset with sub-20B open-weight models.
This paper investigates the calibration gap that arises when temperature scaling, a common post-hoc calibration method relying on one-hot labels, is applied to models trained with soft label distributions reflecting genuine human disagreement. Experiments across vision and language domains show that temperature scaling calibrated on hard labels consistently underperforms direct soft-label calibration, with larger gaps in language tasks, highlighting risks for safety-critical deployments.