Tag
Introduces the Sys1Cal-v1 dataset for evaluating probability calibration in AI models, showing that models like Jev may suppress uncertainty in binary outputs.
RLCD is explained as a schema-conditioned Plackett–Luce objective that advances reward modeling from scalar rewards to pairwise preferences to multiway calibrated decisions, simplifying the understanding of Jev.
This paper presents the first study of probability calibration as a mitigation for evaluator preference coupling in LLM agent feedback loops, showing that calibrated evaluator judgments reduce coupling coefficients by 20-49% and divergence by 45-67%.
Introduces Simplex-Constrained Sparse Bagging (SCSB), a post-training framework that optimizes estimator weights over the probability simplex using out-of-bag samples, achieving up to 96% ensemble compression and improved calibration.