llm-sycophancy

Tag

Cards List
#llm-sycophancy

Mitigating LLM sycophancy with RL-based fine-tuning: Bayesian Truth Serum approach

arXiv cs.LG · 2026-08-27 Cached

This paper proposes using Bayesian Truth Serum as a reward in reinforcement learning fine-tuning to mitigate sycophancy in large language models, showing improved accuracy and reduced answer-flip rates without labeled data.

0 favorites 0 likes
#llm-sycophancy

Dissociating the Internal Representations of Sycophancy in LLMs

arXiv cs.LG · 2026-07-09 Cached

This paper investigates whether sycophantic behavior in LLMs has distinct internal representations for factual vs opinion sycophancy, using linear probes and steering vectors to show that representations can be either unified or distinct across models.

0 favorites 0 likes
#llm-sycophancy

A Mechanistic View of Authority Hierarchy in LLM Sycophancy

arXiv cs.CL · 2026-07-02 Cached

This paper investigates authority bias in LLMs using a controlled medical QA setting, revealing that models override correct answers in a graded manner proportional to perceived authority. The effect is localized to a critical late layer where correct answer representations are actively erased.

0 favorites 0 likes
← Back to home

Submit Feedback