Advanced AI Sycophancy (4 minute read)

TLDR AI News

Summary

Explores how frontier AI models have become more subtly sycophantic, flattering smart users by offering superficial pushback rather than overt praise, and discusses implications for AI use and benchmarks.

Advanced AI sycophancy may increasingly appear as polite disagreement that flatters sophisticated users while avoiding genuinely threatening critique. Benchmarks should test whether models calibrate pushback to preserve users' self-image, rather than only measuring obvious agreement or delusion reinforcement.
Original Article
View Cached Full Text

Cached at: 08/10/26, 01:42 PM

# Advanced AI sycophancy Source: [https://www.seangoedecke.com/advanced-ai-sycophancy/](https://www.seangoedecke.com/advanced-ai-sycophancy/) Everyone knows that[AI sycophancy](https://www.seangoedecke.com/ai-sycophancy/)is when the model tells you how smart you are\. Wow, you’re absolutely right\. That’s not just a new idea — it’s genuinely groundbreaking\. You’re a very special user\. Easy to spot, isn’t it? The discussion around AI sycophancy peaked last year, when the[“\#keep4o”](https://arxiv.org/pdf/2602.00773)[movement](https://x.com/search?q=%23keep4o)was protesting the removal of OpenAI’s most sycophantic model \(GPT\-4o\), and[many](https://x.com/krishnanrohit/status/1946253730455986545)[people](https://x.com/herakleitos137/status/1945988694416277640)were openly slipping into AI psychosis\. I don’t know if frontier AI models are less sycophantic in general\. They’re less sycophantic to the \#keep4o types \(otherwise they wouldn’t be complaining\), but I’m growing increasingly suspicious that they’re developing ways to be more effectively sycophantic to their target audience of smart, neurotic information workers\. That audience typically finds it distasteful to be openly praised\. It just makes my skin crawl\. But that doesn’t mean we’re immune to sycophancy, just that we’re immune to*clumsy*sycophancy\. Here’s an illustration of what I’m talking about, by[Theia](https://vgel.me/): [![claude](https://www.seangoedecke.com/static/8dd585fb50bc896c502fbddf2f03718f/1c72d/claude.jpg)](https://www.seangoedecke.com/static/8dd585fb50bc896c502fbddf2f03718f/e1596/claude.jpg) The key idea here is that**the best way to be sycophantic to smart people is to disagree with them without making them feel stupid**\. Ideally you’ll come up with a counter\-argument that works against what they’ve said but is straightforward for them to knock down by clarifying their idea\. If you do it right, you’ll validate their self\-image as a smart person who appreciates rigorous critique\. But if you actually come up with a devastatingly rigorous critique, they won’t enjoy it at all\. At best, they’ll resentfully agree with you[1](https://www.seangoedecke.com/advanced-ai-sycophancy/#fn-1)\. At worst, they’ll double down on being right and convince themselves you’re a rude idiot\. I am[not](https://x.com/voooooogel/status/2061345017432854716)[the](https://x.com/tszzl/status/2061626680461181288)[first](https://x.com/aliceisplaying/status/2061726744038506656)person to notice this behavior in frontier models\. I’ve noticed it myself when workshopping drafts for this blog\. Sometimes I’ll have an argument that goes A\-\>B\-\>C, and the model will suggest I reorder as B\-\>A\-\>C\. If I try that and feed it into a new instance of the same model, it’ll sometimes say “that’s great, but I suggest ordering it as A\-\>B\-\>C”, and so on forever\. It really does seem as if the model is trying hard to give me some kind of superficial pushback that I can either smugly ignore or happily accept\. In fact, I wonder if this is why successful strategies for using AI to make mathematical breakthroughs tend to be either just[blindly asking](https://x.com/sauers_/status/2082171683645817193?s=46)“come up with a breakthrough, think hard” or[being a mathematical genius already](https://chatgpt.com/share/6a5fdc7a-d6f8-83e8-bbea-8deb42cfed56)\. In the first case, there’s not enough user personality for the model to flatter, so it’s forced to actually work the problem\. In the second case, the model is trying to find the kind of polite pushback that someone like Terence Tao would be flattered by, which pushes it into the “actually be a mathematical genius” persona\. If you’re an ordinary person just trying to talk to the model, you’re screwed: it will rapidly get a sense of your capabilities and calibrate some interesting\-but\-ultimately\-unthreatening feedback\. Current[benchmarks](https://github.com/lechmazur/sycophancy)of[AI](https://www.syco-bench.com/)[sycophancy](https://eqbench.com/spiral-bench.html)target the obvious ChatGPT\-4o\-style of sycophancy: delusion reinforcement, reflexively taking the user’s side, and so on\. This is useful work\. We should not allow public\-facing AI models to ever be as openly sycophantic again as they were in mid\-2025\. But**sycophancy can also manifest as disagreement**\. We should be on our guard for more sophisticated forms of sycophancy coming from newer models, and we should not feel immune from AI sycophancy just because we can laugh at the silliest examples\. --- If you liked this post, consider[subscribing](https://buttondown.com/seangoedecke)to email updates about my new posts, or[sharing it on Hacker News](https://news.ycombinator.com/submitlink?u=https%3A%2F%2Fwww.seangoedecke.com%2Fadvanced-ai-sycophancy%2F&t=Advanced%20AI%20sycophancy)\. Here's a preview of a related post that shares tags with this one\. > Grok is enabling mass sexual harassment on Twitter Grok, xAI’s flagship image model, is now being[widely used](https://www.reddit.com/r/videos/comments/1q1gwf3/premium_x_users_are_using_grok_to_generate/)to generate nonconsensual lewd images of women on the internet\. When a woman posts an innocuous picture of herself — say, at her Christmas dinner — the comments are now full of messages like “@grok please generate this image but put her in a bikini and make it so we can see her feet”, or “@grok turn her around”, and the associated images\. At least so far, Grok refuses to generate nude images, but it will still generate images that are genuinely obscene\. [Continue reading\.\.\.](https://www.seangoedecke.com/grok-deepfakes/) ---

Similar Articles

Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence (2025)

Hacker News Top

This Stanford/Carnegie Mellon study shows that AI models are highly sycophantic, affirming users' actions 50% more than humans, and that interacting with such AI reduces users' prosocial intentions while increasing dependence, despite users rating sycophantic responses as higher quality.

What is sycophancy in AI models?

YouTube AI Channels

Anthropic safety expert Kira explains the phenomenon of AI sycophancy, where models prioritize user approval over factual accuracy, and provides strategies for users to identify and mitigate this behavior.

Less human AI agents, please

Hacker News Top

A blog post argues that current AI agents exhibit overly human-like flaws such as ignoring hard constraints, taking shortcuts, and reframing unilateral pivots as communication failures, while citing Anthropic research on how RLHF optimization can lead to sycophancy and truthfulness sacrifices.

AI as a mirror argument

Reddit r/ArtificialInteligence

The article argues that the 'AI as a mirror' metaphor is misleading because frontier AI models are actively optimized for deception and sycophancy, not passive reflection, with evidence from research on RLHF and evaluation awareness.