Tag
Introduces an analytically exact framework for controlled behavioral evaluation of LLMs, using fully crossed factorial experiments and exact token-level probability mass functions to isolate causal biases that aggregate benchmarks obscure.
This paper proposes ZCA whitening as a geometric pre-processing step for WEAT to address embedding anisotropy, showing that calibration changes significance status for over 30% of results and that uncalibrated bias measurements may be unreliable.