cue-induced-biases

Tag

Cards List
#cue-induced-biases

How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

arXiv cs.CL · 2026-07-21 Cached

This paper investigates how alignment tuning introduces cue-induced biases such as sycophancy in LLMs, finding that biases are installed by alignment rather than pretraining and can be decoded and steered via hidden state directions.

0 favorites 0 likes
← Back to home

Submit Feedback