Tag
This paper identifies spectral drift in internal activations of neural networks during misclassifications and introduces Self-Detecting Neural Networks (SDNN) that monitors spectral dynamics to detect failures, achieving 79% AUROC on CIFAR-10, outperforming confidence-based methods by 25-30 percentage points.
This paper introduces a method to calibrate uncertainty in language models by extracting eleven scale-invariant geometric features from per-layer MLP update trajectories and feeding them to a sparse linear probe, outperforming MSP under selective abstention by up to 21 AURC points.