side-effects

Tag

Cards List
#side-effects

Forecasting Side Effects of Activation Steering

arXiv cs.AI · 3d ago Cached

This paper investigates whether side effects of activation steering in language models can be predicted before intervention, constructing a cross-effect matrix across 67 behaviors and finding that side effects are systematic and forecastable from unsteered representations.

0 favorites 0 likes
#side-effects

Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning

arXiv cs.LG · 2026-08-06 Cached

This paper introduces a new problem setting called side-effect introspection, where the goal is to detect alignment degradation in fine-tuned LLMs that occurs as an unintended side effect rather than explicitly implanted behavior. The authors propose a novel Delta-Aware Introspection Adapter (DAIA) that outperforms existing introspection adapters in generalizing to unseen models and safety categories.

0 favorites 0 likes
#side-effects

When Debiasing Backfires: Counterintuitive Side Effects of Preprocessing-Based Stereotype Mitigation

arXiv cs.CL · 2026-07-10 Cached

This paper investigates preprocessing-based stereotype mitigation methods in NLP and finds that while they reduce targeted stereotypes, they can inadvertently increase stereotyping or counter-stereotyping for other demographic groups, including across unrelated categories. The authors demonstrate these side effects across model families and preprocessing strategies, and discuss implications for evaluation and mitigation practices.

0 favorites 0 likes
← Back to home

Submit Feedback