Tag
This paper investigates whether sycophantic behavior in LLMs has distinct internal representations for factual vs opinion sycophancy, using linear probes and steering vectors to show that representations can be either unified or distinct across models.
This paper investigates authority bias in LLMs using a controlled medical QA setting, revealing that models override correct answers in a graded manner proportional to perceived authority. The effect is localized to a critical late layer where correct answer representations are actively erased.