Tag
This paper investigates whether side effects of activation steering in language models can be predicted before intervention, constructing a cross-effect matrix across 67 behaviors and finding that side effects are systematic and forecastable from unsteered representations.
This paper introduces a new problem setting called side-effect introspection, where the goal is to detect alignment degradation in fine-tuned LLMs that occurs as an unintended side effect rather than explicitly implanted behavior. The authors propose a novel Delta-Aware Introspection Adapter (DAIA) that outperforms existing introspection adapters in generalizing to unseen models and safety categories.
This paper investigates preprocessing-based stereotype mitigation methods in NLP and finds that while they reduce targeted stereotypes, they can inadvertently increase stereotyping or counter-stereotyping for other demographic groups, including across unrelated categories. The authors demonstrate these side effects across model families and preprocessing strategies, and discuss implications for evaluation and mitigation practices.