Tag
The paper investigates neurons in frozen BERT that drive AI-text detection using sparse probing and activation patching on the RAID benchmark, identifying a small set of causally relevant neurons that generalize across generator families.
Introduces SPOT, a method for on-policy distillation that uses sparse probing and outcome calibration to improve reasoning performance in smaller student models while balancing solution quality and coverage.