Tag
This paper introduces a site-asymmetry audit to separate baseline effects from interaction in activation-space order-swap interventions, showing that single-intervention baselines explain most variance in language models and proposing a corrected residual for detecting geometric structure.
This paper proposes Neural-Bayesian Structure Learning, a framework that integrates differentiable structure learning with discrete choice modeling to predict choice behavior and evaluate interventions, achieving comparable performance while recovering coherent dependency structures.