Tag
This preprint uses adversarial training as a controlled instrument on GPT-2 Small to test whether representational simplicity (SAE decomposability, concentrated attribution) implies causal circuit simplicity. It finds that robust models are more SAE-decomposable, while circuit size is regime-dependent: robustness helps at high faithfulness levels (90-95%) but not below 85%.