Tag
This preregistered replication tests whether the monotonicity effect on label agreement in NLI generalizes from selected low-agreement items to unselected populations, finding that the effect reverses and is small, suggesting the earlier finding was conditional on selection.
This paper introduces Heckman-corrected epistemic uncertainty to address selection on unobservables in machine learning, demonstrating that importance weighting fails when selection depends on unobservables correlated with outcomes. The method restores calibration in controlled experiments and real data, outperforming standard UQ baselines.
This paper systematically evaluates reject inference methods in credit scoring and identifies a failure mode where accuracy improves while recall collapses, creating an illusion of improvement while rejection quality deteriorates. It proposes a controlled exploration strategy that breaks the feedback loop and shows that even minimal exploration rates are sufficient to diagnose the problem.