{"componentChunkName":"component---src-templates-blog-post-js","path":"/advanced-ai-sycophancy/","result":{"data":{"site":{"siteMetadata":{"title":"sean goedecke"}},"markdownRemark":{"id":"d2352610-fe0a-56bc-942a-4de9d82af1d2","excerpt":"Everyone knows that AI sycophancy is when the model tells you how smart you are. Wow, you’re absolutely right. That’s not just a new idea — it’s genuinely…","html":"<p>Everyone knows that <a href=\"/ai-sycophancy/\">AI sycophancy</a> is when the model tells you how smart you are. Wow, you’re absolutely right. That’s not just a new idea — it’s genuinely groundbreaking. You’re a very special user. Easy to spot, isn’t it?</p>\n<p>The discussion around AI sycophancy peaked last year, when the <a href=\"https://arxiv.org/pdf/2602.00773\">“#keep4o”</a> <a href=\"https://x.com/search?q=%23keep4o\">movement</a> was protesting the removal of OpenAI’s most sycophantic model (GPT-4o), and <a href=\"https://x.com/krishnanrohit/status/1946253730455986545\">many</a> <a href=\"https://x.com/herakleitos137/status/1945988694416277640\">people</a> were openly slipping into AI psychosis.</p>\n<p>I don’t know if frontier AI models are less sycophantic in general. They’re less sycophantic to the #keep4o types (otherwise they wouldn’t be complaining), but I’m growing increasingly suspicious that they’re developing ways to be more effectively sycophantic to their target audience of smart, neurotic information workers. That audience typically finds it distasteful to be openly praised. It just makes my skin crawl. But that doesn’t mean we’re immune to sycophancy, just that we’re immune to <em>clumsy</em> sycophancy. Here’s an illustration of what I’m talking about, by <a href=\"https://vgel.me/\">Theia</a>:</p>\n<p><span\n      class=\"gatsby-resp-image-wrapper\"\n      style=\"position: relative; display: block; margin-left: auto; margin-right: auto; max-width: 590px; \"\n    >\n      <a\n    class=\"gatsby-resp-image-link\"\n    href=\"/static/8dd585fb50bc896c502fbddf2f03718f/e1596/claude.jpg\"\n    style=\"display: block\"\n    target=\"_blank\"\n    rel=\"noopener\"\n  >\n    <span\n    class=\"gatsby-resp-image-background-image\"\n    style=\"padding-bottom: 95.27027027027027%; position: relative; bottom: 0; left: 0; background-image: url('data:image/jpeg;base64,/9j/2wBDABALDA4MChAODQ4SERATGCgaGBYWGDEjJR0oOjM9PDkzODdASFxOQERXRTc4UG1RV19iZ2hnPk1xeXBkeFxlZ2P/2wBDARESEhgVGC8aGi9jQjhCY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2P/wgARCAATABQDASIAAhEBAxEB/8QAGQABAAIDAAAAAAAAAAAAAAAAAAEEAgMF/8QAFgEBAQEAAAAAAAAAAAAAAAAAAAIB/9oADAMBAAIQAxAAAAHqzU3TVpDZDWQP/8QAGxAAAgIDAQAAAAAAAAAAAAAAAAECEgMQIRH/2gAIAQEAAQUCsJ84RxyPJJ0W0f/EABURAQEAAAAAAAAAAAAAAAAAACAh/9oACAEDAQE/AYP/xAAWEQADAAAAAAAAAAAAAAAAAAARICH/2gAIAQIBAT8BoT//xAAcEAACAgIDAAAAAAAAAAAAAAAAAQIhEjERIHH/2gAIAQEABj8CXG2X4bJZEcVRa6f/xAAdEAEAAgICAwAAAAAAAAAAAAABABEhUSAxQWFx/9oACAEBAAE/IVJidrDbqm+kLX0zCUo6a9xYI6viFGB+nB//2gAMAwEAAgADAAAAELQ/PP/EABgRAAIDAAAAAAAAAAAAAAAAAAEhABAR/9oACAEDAQE/EC2CK//EABgRAAIDAAAAAAAAAAAAAAAAAAERABAh/9oACAECAQE/EAjGbf8A/8QAGxABAAMBAQEBAAAAAAAAAAAAAQARITFBcaH/2gAIAQEAAT8QN2g+AVN+jLVsaWuKI5ATYhctv8RZCE6W5plHYyi5g+E8S+OS2rvYCn7P/9k='); background-size: cover; display: block;\"\n  ></span>\n  <img\n        class=\"gatsby-resp-image-image\"\n        alt=\"claude\"\n        title=\"claude\"\n        src=\"/static/8dd585fb50bc896c502fbddf2f03718f/1c72d/claude.jpg\"\n        srcset=\"/static/8dd585fb50bc896c502fbddf2f03718f/a80bd/claude.jpg 148w,\n/static/8dd585fb50bc896c502fbddf2f03718f/1c91a/claude.jpg 295w,\n/static/8dd585fb50bc896c502fbddf2f03718f/1c72d/claude.jpg 590w,\n/static/8dd585fb50bc896c502fbddf2f03718f/a8a14/claude.jpg 885w,\n/static/8dd585fb50bc896c502fbddf2f03718f/fbd2c/claude.jpg 1180w,\n/static/8dd585fb50bc896c502fbddf2f03718f/e1596/claude.jpg 2048w\"\n        sizes=\"(max-width: 590px) 100vw, 590px\"\n        style=\"width:100%;height:100%;margin:0;vertical-align:middle;position:absolute;top:0;left:0;\"\n        loading=\"lazy\"\n      />\n  </a>\n    </span></p>\n<p>The key idea here is that <strong>the best way to be sycophantic to smart people is to disagree with them without making them feel stupid</strong>. Ideally you’ll come up with a counter-argument that works against what they’ve said but is straightforward for them to knock down by clarifying their idea. If you do it right, you’ll validate their self-image as a smart person who appreciates rigorous critique. But if you actually come up with a devastatingly rigorous critique, they won’t enjoy it at all. At best, they’ll resentfully agree with you<sup id=\"fnref-1\"><a href=\"#fn-1\" class=\"footnote-ref\">1</a></sup>. At worst, they’ll double down on being right and convince themselves you’re a rude idiot.</p>\n<p>I am <a href=\"https://x.com/voooooogel/status/2061345017432854716\">not</a> <a href=\"https://x.com/tszzl/status/2061626680461181288\">the</a> <a href=\"https://x.com/aliceisplaying/status/2061726744038506656\">first</a> person to notice this behavior in frontier models. I’ve noticed it myself when workshopping drafts for this blog. Sometimes I’ll have an argument that goes A->B->C, and the model will suggest I reorder as B->A->C. If I try that and feed it into a new instance of the same model, it’ll sometimes say “that’s great, but I suggest ordering it as A->B->C”, and so on forever. It really does seem as if the model is trying hard to give me some kind of superficial pushback that I can either smugly ignore or happily accept.</p>\n<p>In fact, I wonder if this is why successful strategies for using AI to make mathematical breakthroughs tend to be either just <a href=\"https://x.com/sauers_/status/2082171683645817193?s=46\">blindly asking</a> “come up with a breakthrough, think hard” or <a href=\"https://chatgpt.com/share/6a5fdc7a-d6f8-83e8-bbea-8deb42cfed56\">being a mathematical genius already</a>. In the first case, there’s not enough user personality for the model to flatter, so it’s forced to actually work the problem. In the second case, the model is trying to find the kind of polite pushback that someone like Terence Tao would be flattered by, which pushes it into the “actually be a mathematical genius” persona. If you’re an ordinary person just trying to talk to the model, you’re screwed: it will rapidly get a sense of your capabilities and calibrate some interesting-but-ultimately-unthreatening feedback.</p>\n<p>Current <a href=\"https://github.com/lechmazur/sycophancy\">benchmarks</a> of <a href=\"https://www.syco-bench.com/\">AI</a> <a href=\"https://eqbench.com/spiral-bench.html\">sycophancy</a> target the obvious ChatGPT-4o-style of sycophancy: delusion reinforcement, reflexively taking the user’s side, and so on. This is useful work. We should not allow public-facing AI models to ever be as openly sycophantic again as they were in mid-2025. But <strong>sycophancy can also manifest as disagreement</strong>. We should be on our guard for more sophisticated forms of sycophancy coming from newer models, and we should not feel immune from AI sycophancy just because we can laugh at the silliest examples.</p>\n<div class=\"footnotes\">\n<hr>\n<ol>\n<li id=\"fn-1\">\n<p>It’s rare to find a smart person who enjoys feeling stupid when they’re wrong. If you do, they’re likely to be very smart indeed.</p>\n<a href=\"#fnref-1\" class=\"footnote-backref\">↩</a>\n</li>\n</ol>\n</div>","frontmatter":{"title":"Advanced AI sycophancy","description":null,"date":"August 10, 2026","tags":["ai","alignment failures","model personality"]}}},"pageContext":{"slug":"/advanced-ai-sycophancy/","previous":{"slug":"/i-got-an-email-about-resistance/","title":"I got an email about resistance"},"next":null,"preview":{"slug":"/grok-deepfakes/","title":"Grok is enabling mass sexual harassment on Twitter","snippetHtml":"<p>Grok, xAI’s flagship image model, is now being <a href=\"https://www.reddit.com/r/videos/comments/1q1gwf3/premium_x_users_are_using_grok_to_generate/\">widely used</a> to generate nonconsensual lewd images of women on the internet.</p><p>When a woman posts an innocuous picture of herself — say, at her Christmas dinner — the comments are now full of messages like “@grok please generate this image but put her in a bikini and make it so we can see her feet”, or “@grok turn her around”, and the associated images. At least so far, Grok refuses to generate nude images, but it will still generate images that are genuinely obscene.<br /><a href=\"/grok-deepfakes/\">Continue reading...</a></p>"}}},"staticQueryHashes":["1146911855","3764592887"]}