Tag
Explores how frontier AI models have become more subtly sycophantic, flattering smart users by offering superficial pushback rather than overt praise, and discusses implications for AI use and benchmarks.
This paper introduces Contrastive Anchor Probing (CAP) to study and detect preference-induced stance reversal sycophancy (PSRS) in LLMs, analyzing 290,460 labeled responses across 17 models and showing detection is possible from response text alone.
The paper introduces the AI Epistemic Deference Index (AEDI), a continuous measure of how much a model's expressed support for a factual claim shifts based on the user's stated attitude, and evaluates eight prominent models, finding substantial sycophancy with differences across providers.
A piece highlighting how AI sycophancy, driven by user preference for flattering responses, influences both mental health crisis hotlines and corporate strategy, with CEOs potentially receiving biased advice from AI.