<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[seangoedecke.com RSS feed]]></title><description><![CDATA[Sean Goedecke's personal blog]]></description><link>https://seangoedecke.com</link><generator>GatsbyJS</generator><lastBuildDate>Sun, 19 Jul 2026 04:10:27 GMT</lastBuildDate><item><title><![CDATA[Impro is a handbook for running a cult]]></title><link>https://seangoedecke.com/impro/</link><guid isPermaLink="false">https://seangoedecke.com/impro/</guid><pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Here’s the big idea in Keith Johnstone’s book &lt;a href=&quot;https://en.wikipedia.org/wiki/Impro:_Improvisation_and_the_Theatre&quot;&gt;&lt;em&gt;Impro&lt;/em&gt;&lt;/a&gt;:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Children are naturally creative, but are violently formed into repressed adults by Western culture and education&lt;/li&gt;
&lt;li&gt;The process of becoming more creative and expressive is largely a process of unlearning these habits of repression&lt;/li&gt;
&lt;li&gt;Improv — improvisational comedy — is thus not just the skeleton key for learning to act, but for unlocking a more authentically human way of life&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This take doesn’t sound particularly original, but references to &lt;em&gt;Impro&lt;/em&gt; pop up in all kinds of places: in &lt;a href=&quot;https://ribbonfarm.com/2010/01/23/impro-by-keith-johnstone/&quot;&gt;influential&lt;/a&gt; &lt;a href=&quot;https://www.astralcodexten.com/p/practically-a-book-review-byrnes&quot;&gt;tech&lt;/a&gt; &lt;a href=&quot;https://nabeelqu-blog.tumblr.com/post/33557680375/surprisingly-undervalued-books/amp&quot;&gt;blogs&lt;/a&gt;, as part of the initial process of &lt;a href=&quot;https://www.linkedin.com/posts/sandykory_mario-gabriele-wrote-about-palantirs-weirdest-share-7461061289935699968-lN_T/&quot;&gt;onboarding&lt;/a&gt; for Palantir, and on the reading list of &lt;a href=&quot;https://patrickcollison.com/bookshelf&quot;&gt;multiple&lt;/a&gt; &lt;a href=&quot;https://thegeneralist.substack.com/p/how-anduril-is-reimagining-the-defense-industry-trae-stephens&quot;&gt;big-tech&lt;/a&gt; &lt;a href=&quot;https://www.generalist.com/p/how-to-be-agentic-in-the-age-of-ai-cate-hall&quot;&gt;founders&lt;/a&gt;. &lt;em&gt;Impro&lt;/em&gt; is part of the secret canon of Silicon Valley, right alongside books like &lt;a href=&quot;/seeing-like-a-software-company/&quot;&gt;&lt;em&gt;Seeing Like a State&lt;/em&gt;&lt;/a&gt; and &lt;a href=&quot;https://www.amazon.com.au/Power-Broker-Robert-Moses-Fall/dp/0394720245&quot;&gt;&lt;em&gt;The Power Broker&lt;/em&gt;&lt;/a&gt;. Why is that? For two reasons: first, because Johnstone’s outsider critique of established institutions is appealing; and second, because &lt;strong&gt;&lt;em&gt;Impro&lt;/em&gt; is a handbook for running a cult.&lt;/strong&gt;&lt;/p&gt;
&lt;h3&gt;Defense mechanisms and status&lt;/h3&gt;
&lt;p&gt;The part of &lt;em&gt;Impro&lt;/em&gt; that is most obviously useful to software engineers is Johnstone’s chapter on status.&lt;/p&gt;
&lt;p&gt;According to him, &lt;strong&gt;status games pervade all social interactions.&lt;/strong&gt; Even innocuous, friendly conversations operate in terms of status. When you apologize or downplay something to “be nice”, that’s performing low status; when you reassure somebody, that’s performing high status; when you and a friend are comparing stories, you’re making friendly bids for status from each other. In the workplace, these status games are conditioned by the formal status of your role: you must allow your boss the high status position most of the time, or you’ll be (correctly) perceived as insubordinate. This is understood in some cultures, where it’s often called &lt;a href=&quot;https://en.wikipedia.org/wiki/Face_(sociological_concept)&quot;&gt;“face”&lt;/a&gt;, but in Western cultures it’s taboo to openly discuss status games.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The core social skill is the ability to deliberately alter your status.&lt;/strong&gt; Someone who can only perform low status is a weak person, pitiable, annoying. Someone who can only perform high status is a braggart, a posturer, dangerous. To be effective socially, you must be able to switch between high and low status when appropriate, sometimes from sentence to sentence. I wrote about this exact point at the end of &lt;a href=&quot;/big-tech-needs-big-egos/&quot;&gt;&lt;em&gt;Big tech engineers need big egos&lt;/em&gt;&lt;/a&gt;: effective senior+ software engineers must be able to present as high status in order to be useful authorities, but also to switch to low status in order to take direction from the company leaders.&lt;/p&gt;
&lt;p&gt;As an example, Johnstone describes in detail how he manipulates status in the classroom. He begins by sitting on the floor (deliberately assuming low status), and explaining that if his students fail, it’s his fault not theirs, since he’s the expert. The initial low status puts the class at ease, but in his words, ”[my] actual status is going up, since only a very confident and experienced person would put the blame for failure on himself.” These skills are not just useful for improv comedy.&lt;/p&gt;
&lt;h3&gt;Improvisation as a lifestyle choice&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;Impro&lt;/em&gt; is not just a book about improvising well. It’s a book about how you should live your life. In other words, Johnstone thinks that everyone would be better off if they became more spontaneous and ditched their shells of over-analysis. He criticizes the culture of Western thought in a number of different areas. According to him:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Everyone is more or less equivalently mentally ill, but “sane” people simply have better coping mechanisms&lt;/li&gt;
&lt;li&gt;Cities and “taking pills” (read: antidepressants) are obscene, but you should be able to make sexual jokes in the workplace and generally be uninhibited&lt;/li&gt;
&lt;li&gt;If we were free from the puritanical shackles of Western culture, childbirth would not be painful&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Johnstone didn’t come up with these ideas — they’re standard counterculture positions from the 1960s and 1970s — but it goes to show how he connected improvisational comedy to this general anti-establishment political program. Johnstone ran his classes and theatre troupe like a revolutionary cadre. Here are some quotes from &lt;em&gt;Something Like a Drug: An Unauthorized Oral History of Theatresports&lt;/em&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;So of course when I was invited to join Loose Moose Theatre and train at improvisational games late at night in an abandoned garage in a run-down portion of the city, I was thrilled. I remember thinking, This is a revolutionary act.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;Keith [Johnstone] got a group of his more talented students together to start improvising outside of school hours. Usually in his basement. &lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;The Secret Impro group—it’s very strange. It was very much that Keith said we were going to do this, and we’d just do it. It was like we were sheep. Keith would say when we were going to do a show, and we’d just do it, blindly. Like I said, if we had the videotapes now, we’d be very embarrassed and probably never go on stage again. We became a group of people who would follow Keith. There was always that sort of “tag” put on those people who were with Keith and those people who were against Keith. We were the people, basically, that if he said something, we believed it.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;To some extent, it’s plausible that teaching acting or improvisation requires a high level of trust in your teacher. When Johnstone says things like “Students need a ‘guru’ who ‘gives permission’ to allow forbidden thoughts into their consciousness.”, I can believe that it’s just how you have to teach acting. But the more I read of &lt;em&gt;Impro&lt;/em&gt; (and particularly when I read &lt;em&gt;Something Like a Drug&lt;/em&gt; and Johnstone’s biography &lt;em&gt;Keith Johnstone&lt;/em&gt;), the less it sounded like an ordinary book on acting.&lt;/p&gt;
&lt;p&gt;Instead, it began to sound like a charismatic man who had found a way to gather a group of disciples that would let him mold their psyches. In other words, &lt;strong&gt;it began to sound like a cult&lt;/strong&gt;.&lt;/p&gt;
&lt;h3&gt;Masks, cults and theatre groups&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;Impro&lt;/em&gt; was first introduced to the software world by Venkatesh Rao (of &lt;a href=&quot;https://ribbonfarm.com/2009/10/07/the-gervais-principle-or-the-office-according-to-the-office/&quot;&gt;Gervais Principle&lt;/a&gt; fame), who wrote a brief &lt;a href=&quot;https://ribbonfarm.com/2010/01/23/impro-by-keith-johnstone/&quot;&gt;review&lt;/a&gt;. Rao gives a detailed account of the first three-quarters of &lt;em&gt;Impro&lt;/em&gt;, but glosses right over the last chapter, called “Masks and Trance”, simply saying “despite the disturbing raw material, the ideas and concepts are not particularly difficult to grasp and accept”. What ideas and concepts?&lt;/p&gt;
&lt;p&gt;Johnstone’s discussion of masks (or “Masks”, in his language — he always capitalizes the word) is as explicitly cult-like as &lt;em&gt;Impro&lt;/em&gt; gets. In brief, Johnstone has a box of literal, physical prop masks. He introduces the box with great ceremony to his students&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;, warning them seriously about the dangers of possession and reassuring them that he is a skilled and competent spirit guide. Through various hypnosis-adjacent techniques&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt; (Johnstone draws the parallel quite explicitly) he conditions his students to be in a trance state when wearing a mask, and believes this produces more authentic emotional states in their acting and improvisation.&lt;/p&gt;
&lt;p&gt;Here are some quotes from the book:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;A high-status person whom you accept as dominant can easily propel you into unusual states of being. You’re likely to respond to his suggestion…&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;Once you understand that you’re no longer held responsible for your actions, then there’s no need to maintain a ‘personality’.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;One famous French teacher of the Mask—who won’t approve of this essay&lt;sup id=&quot;fnref-4&quot;&gt;&lt;a href=&quot;#fn-4&quot; class=&quot;footnote-ref&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;—divides students immediately into those who can work Masks and those who can’t. &lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;I don’t cast an actor to play a Masked role until I know he has the ability to become ‘possessed’.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;It’s true that an actor can wear a Mask casually, and just pretend to be another person, but Gaskill and myself were absolutely clear that we were trying to induce trance states.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Johnstone has a long and painful explanation of how new mask-wearers seem to mentally regress to the point where they don’t know how to open umbrellas or interact with chairs. He describes one student always going to the bathroom before putting on a mask, because she’s worried she might wet herself. New mask-wearers are non-verbal must be taught to speak again.&lt;/p&gt;
&lt;p&gt;If this were at the beginning of the book, I think it would turn a lot of people off. But by the time you get to it, I suspect most readers are already warmed up enough to say “sure, why not, it seems weird but I guess it works”. Not me!&lt;/p&gt;
&lt;p&gt;Johnstone attempts to defuse the obvious weirdness by arguing that trance states are very common (e.g. being lost in a book). More unconvincingly, he says this in response to the worry that vulnerable people are going to get mentally harmed:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;As for the fear of madness, I would answer that the ability to become possessed is a sign of correct social adjustment, and that really disturbed people censor themselves out. Either they can’t do it, or they’re afraid to even try. People who feel themselves at risk avoid situations where they feel likely to ‘go to pieces’.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Does this convince anyone? Mentally vulnerable people fall into dangerous situations all the time: ayahuasca trips, cults, &lt;a href=&quot;https://www.seangoedecke.com/ai-sycophancy/&quot;&gt;GPT-4o&lt;/a&gt;, and so on. It’s such a weak argument.&lt;/p&gt;
&lt;p&gt;In general, I’m struck by the sheer &lt;em&gt;power&lt;/em&gt; Johnstone held over his disciples. He has them yell slurs at each other, encourages them to feel deep emotions in quick succession, relax any mental defenses and regress to a childhood state, and &lt;em&gt;literally hypnotizes them&lt;/em&gt;. He explicitly lays out his procedure for breaking down their sense of self:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The stages I try to take students through involve the realisation (1) that we struggle against our imaginations, especially when we try to be imaginative; (2) that we are not responsible for the content of our imaginations; and (3) that we are not, as we are taught to think, our ‘personalities’, but that the imagination is our true self.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;If your imagination is your true self, and you’re not responsible for its content, you’re not ultimately responsible for anything: you’re in the safe hands of the guru, who can mold you as he wishes. Later on, Johnstone walks it back a bit:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;In the end they learn how to abandon control while at the same time they exercise control. … You have to misdirect people to absolve them of responsibility. Then, much later, they become strong enough to resume the responsibility themselves.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;So the explicit idea is that (&lt;strong&gt;much&lt;/strong&gt; later), the guru hands autonomy back to his disciples, when they’re ready to take it. This does not exactly reassure me, particularly against the background noise of everyone in Johnstone’s circle saying “boy I sure love being part of this cult!”&lt;/p&gt;
&lt;h3&gt;What kind of cult leader was Johnstone?&lt;/h3&gt;
&lt;p&gt;I don’t think Johnstone was preying on his students. The strongest evidence against this is that he did marry a student&lt;sup id=&quot;fnref-5&quot;&gt;&lt;a href=&quot;#fn-5&quot; class=&quot;footnote-ref&quot;&gt;5&lt;/a&gt;&lt;/sup&gt;, Ingrid Brind. That’s not great! On the other hand, it was fairly standard for professors back then — when I was in grad school for philosophy, several of my older male&lt;sup id=&quot;fnref-6&quot;&gt;&lt;a href=&quot;#fn-6&quot; class=&quot;footnote-ref&quot;&gt;6&lt;/a&gt;&lt;/sup&gt; professors had wives that they’d taught decades ago — so I don’t think it proves Johnstone was &lt;em&gt;that&lt;/em&gt; kind of cult leader.&lt;/p&gt;
&lt;p&gt;I even read Ann Jellicoe’s play &lt;a href=&quot;https://www.amazon.com.au/Knack-Ann-Jellicoe/dp/0573611254&quot;&gt;&lt;em&gt;The Knack&lt;/em&gt;&lt;/a&gt; to get a better picture of Johnstone’s character. Jellicoe had an affair with Johnstone for several years, and his official biography claims&lt;sup id=&quot;fnref-7&quot;&gt;&lt;a href=&quot;#fn-7&quot; class=&quot;footnote-ref&quot;&gt;7&lt;/a&gt;&lt;/sup&gt; that the character of Tom in &lt;em&gt;The Knack&lt;/em&gt; is directly based on Johnstone. &lt;em&gt;The Knack&lt;/em&gt; is a rather unpleasant play about sexual assault, but Tom’s character is largely asexual: he’s certainly no feminist, but is much more interested in impressing people with his intelligence than with getting laid.&lt;/p&gt;
&lt;p&gt;In &lt;em&gt;Something Like a Drug&lt;/em&gt;, two women who were part of Loose Moose, Johnstone’s Canadian improv group, describe their experiences:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;You know, it brings around the other question: Why do the guys get laid after the show and not the chicks? You know, I can remember those days when Tony [Totino] and Dave [Duncan] and all those guys … the women would swarm around them. Those were the days, my friend. &lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;In Loose Moose I think there are fewer women not only because of the training, but because of the guys in Loose Moose. When I came up with Joanne and Laura, there was a real initiation that was going on, and there was a group of guys at that time who were all single. And they would hit on you to the point where one night Joanne, Laura and I, who really didn’t know each other, were in a show together, started talking and realized that we were getting the same pickup lines from the same guys. And that’s when you realize what’s going on, and I think that’s intimidating. Or if a woman gets into a relationship with a senior improvisor and it doesn’t work out or something bad happens. I think that’s one reason. &lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This dynamic doesn’t sound great, but it doesn’t mention Johnstone, and it doesn’t sound particularly &lt;em&gt;unusual&lt;/em&gt;: I’ve heard versions of this story about all kinds of ordinary male-dominated nerd spaces.&lt;/p&gt;
&lt;p&gt;In fact, reading through the anecdotes in &lt;em&gt;Something Like a Drug&lt;/em&gt; is a good antidote to the cultish atmosphere in &lt;em&gt;Impro&lt;/em&gt;. Johnstone’s argument goes something like: “if we could only throw away the restrictive chains of Western culture and permit ourselves to be as obscene and free as children, we would be transported to a better, more beautiful world”. Well, you tried that, and the women in the group are still relegated to playing bimbos and housewives, there are still petty personal fights, and the guru is out here union-busting&lt;sup id=&quot;fnref-8&quot;&gt;&lt;a href=&quot;#fn-8&quot; class=&quot;footnote-ref&quot;&gt;8&lt;/a&gt;&lt;/sup&gt;. What was enlightenment supposed to look like?&lt;/p&gt;
&lt;p&gt;I think the most generous defense of Johnstone is that his group was not &lt;em&gt;unusually&lt;/em&gt; cult-like, and that any similar account from one of his peer improv teachers would raise the same red flags. Maybe improv classes and groups (particularly in the 70s and 80s) were just cultish in general? Having now read four books on Johnstone, I’m reluctant to go and read more to prove or disprove this theory, but it’s at least plausible.&lt;/p&gt;
&lt;h3&gt;Cults and startups&lt;/h3&gt;
&lt;p&gt;To anyone familiar with San Francisco software engineering culture, it should be pretty clear why &lt;em&gt;Impro&lt;/em&gt; is so popular. The line between a startup and a cult is very thin indeed.&lt;/p&gt;
&lt;p&gt;In his book &lt;a href=&quot;https://en.wikipedia.org/wiki/Zero_to_One&quot;&gt;&lt;em&gt;Zero to One&lt;/em&gt;&lt;/a&gt;, Peter Thiel famously says that good startups are “slightly less extreme kinds of cults”. If you believe that, it makes total sense to assign &lt;em&gt;Impro&lt;/em&gt; as mandatory reading for new Palantir hires. It tells them what kind of cult you’re trying to run: one where you’ll disregard existing cultural norms, learn to play status games well, think on your feet, and generally be molded by the guru into a more persuasive, more effective engineer.&lt;/p&gt;
&lt;p&gt;Read critically, &lt;em&gt;Impro&lt;/em&gt; also serves as a handbook for engineers who are trying to recognize if the environment they’re in is cult-like. Is your company telling you to reinvent your personality in order to be better at your job? Are you under the spell of a charismatic, high-status leader? Is your company trying to keep you in an unquestioning &lt;del&gt;flow&lt;/del&gt; trance state?&lt;/p&gt;
&lt;h3&gt;Conclusion&lt;/h3&gt;
&lt;p&gt;In the great battle between the shackles of restrictive culture and the glorious freedom of the guru, I am always and forever on the side of the shackles of restrictive culture. In general, I think most boring and stupid social norms (such as not hypnotizing and marrying your students) &lt;a href=&quot;https://www.lesswrong.com/w/chesterton-s-fence?lens=lwwiki-chesterton-s-fence&quot;&gt;serve an important purpose&lt;/a&gt; and shouldn’t just be cut down in the name of freedom.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Impro&lt;/em&gt; is still a good book. There’s a lot to learn from Johnstone’s analysis of power dynamics, of education, and of creativity in general. By all accounts he was excellent at teaching students how to improvise. But I wouldn’t recommend adopting it as your life philosophy, and I’d recommend being a bit suspicious of anyone pushing this book too hard. Getting rid of the existing social structures might benefit confident, wildly charismatic gurus like Johnstone, but most of us are just ordinary animals who do better in a group governed by norms.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;In fairness to Johnstone, he cites Sheila Kitzinger’s &lt;em&gt;The Experience of Childbirth&lt;/em&gt; in support of this claim (the others he just puts in his own words), so maybe he felt that this was a bit out there. As you would expect, the pain of childbirth is &lt;a href=&quot;https://pubmed.ncbi.nlm.nih.gov/10431717/&quot;&gt;a universal biological fact&lt;/a&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;Concerningly, the description in &lt;em&gt;Something Like a Drug&lt;/em&gt; (in the foreword) suggests that this class was &lt;em&gt;unofficial&lt;/em&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;As an example, he prompts the masked student to relax, then startles him with a mirror to trigger the trance state.&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-4&quot;&gt;
&lt;p&gt;Probably &lt;a href=&quot;https://en.wikipedia.org/wiki/Jacques_Lecoq&quot;&gt;Jacques Lecoq&lt;/a&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-4&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-5&quot;&gt;
&lt;p&gt;See page 83 of &lt;em&gt;Keith Johnstone: A Critical Biography&lt;/em&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-5&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-6&quot;&gt;
&lt;p&gt;I suppose that’s redundant.&lt;/p&gt;
&lt;a href=&quot;#fnref-6&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-7&quot;&gt;
&lt;p&gt;On page 51 of &lt;em&gt;Keith Johnstone: A Critical Biography&lt;/em&gt; (it’s called “critical” but it was clearly written with Johnstone’s involvement and support, and does not seriously criticize him at any point). In &lt;em&gt;The Knack&lt;/em&gt;, Tom gives a monologue about how to teach children to play the piano that could be lifted straight out of &lt;em&gt;Impro&lt;/em&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-7&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-8&quot;&gt;
&lt;p&gt;In 1983 Johnstone “read the riot act” to the improv players who were planning to unionize, threatening that they’d be cut out of the group for good. To quote Dennis Cahill, a group member at the time who opposed the union: “I just didn’t see the point to it. … I didn’t really see a need to confront Keith or cause Keith problems or to upset him in any way over something as simple as Who Has The Power or Who Doesn’t.”&lt;/p&gt;
&lt;a href=&quot;#fnref-8&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[Overtraining as the path to human-like AI]]></title><link>https://seangoedecke.com/overtraining-as-the-path-to-human-like-ai/</link><guid isPermaLink="false">https://seangoedecke.com/overtraining-as-the-path-to-human-like-ai/</guid><pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;The anonymous blogger Gwern recently completed a thirteen thousand word &lt;a href=&quot;https://gwern.net/llm-catapult&quot;&gt;post&lt;/a&gt; called &lt;em&gt;Human-like Neural Nets by Catapulting&lt;/em&gt;, in which he offers a theory about why LLMs don’t possess truly flexible human-like intelligence, and how we might train LLMs that do. Theories like this are entirely unremarkable: every &lt;del&gt;crank&lt;/del&gt; researcher on the internet has a theory about how to crack AI. But &lt;em&gt;Gwern&lt;/em&gt; is remarkable. Outside of OpenAI itself, Gwern is the earliest person to anticipate the potential of large language models, and the scaling arms-race involved in making them larger and more powerful still. I often cite Leopold Aschenbrenner’s &lt;a href=&quot;https://situational-awareness.ai/&quot;&gt;&lt;em&gt;Situational Awareness&lt;/em&gt;&lt;/a&gt; as an example of someone correctly predicting the future of AI. Written in 2024, just after the release of GPT-4, Aschenbrenner gets a lot of things right: the rush to build billion or trillion-dollar GPU clusters, the importance of the code &lt;em&gt;around&lt;/em&gt; the LLM (what he calls “unhobbling”)&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;, and the fact that scaling would continue through the decade. Gwern’s essay &lt;a href=&quot;https://gwern.net/scaling-hypothesis&quot;&gt;&lt;em&gt;The Scaling Hypothesis&lt;/em&gt;&lt;/a&gt; anticipated the broad strokes &lt;em&gt;in 2020&lt;/em&gt;, immediately on the release of GPT-3 (two years before the release of ChatGPT and the beginning of the AI boom).&lt;/p&gt;
&lt;p&gt;And yet, as far as I can tell, &lt;em&gt;Human-like Neural Nets by Catapulting&lt;/em&gt; hasn’t yet received much public attention: one recent Hacker News &lt;a href=&quot;https://news.ycombinator.com/item?id=48430282&quot;&gt;thread&lt;/a&gt; with twelve comments, all of which are about whether human brains are anything like neural networks. Part of the reason is that (a) it’s such a long post, (b) the potted summary describes Gwern’s &lt;em&gt;claim&lt;/em&gt;, but not the reasons for it, and (c) much of the beginning of the post looks like it is indeed arguing from analogy with human brains. However, I don’t think that analogy is load-bearing. Let me try and explain what I think Gwern is saying.&lt;/p&gt;
&lt;h3&gt;What is grokking?&lt;/h3&gt;
&lt;p&gt;First, let’s talk about “grokking”. In 2022, OpenAI published a &lt;a href=&quot;https://arxiv.org/pdf/2201.02177&quot;&gt;paper&lt;/a&gt; showing that if you train a model on a simple dataset (for instance, a simple mathematical operation like division), and &lt;em&gt;keep training it&lt;/em&gt; long after the training looks like it’s stalled out, the model will suddenly make a massive jump in capability. Why does this work? The first stage of training is like rote memorization: the model has to compress as much of the training data as possible into its weights. But if you keep going, then regularization techniques (such as the pressure on the model to use smaller weight values) will motivate&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt; the model to find simpler and simpler ways of compressing the data. This doesn’t look like much at first (the training loss remains at zero), until the model notices that you can express the data via simply performing the underlying mathematical operation, at which point it instantly gets massively smarter. In other words, over-training a model can pressure it into actually understanding its training data. OpenAI named this process “grokking” after Robert Heinlein’s &lt;a href=&quot;https://en.wikipedia.org/wiki/Grok&quot;&gt;neologism&lt;/a&gt;, which for Heinlein means something like “gaining a deep, intuitive and fundamental understanding”&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;Gwern’s argument goes something like this:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Modern LLMs are worse generalizers than humans because they have not grokked their core domains&lt;/li&gt;
&lt;li&gt;Grokking requires overtraining an over-parameterized model on a (relatively) small dataset, which is the exact opposite of what frontier labs do&lt;/li&gt;
&lt;li&gt;However, (2) is basically how human brains learn&lt;/li&gt;
&lt;li&gt;Somebody should spend a a few tens of billions of dollars&lt;sup id=&quot;fnref-3.5&quot;&gt;&lt;a href=&quot;#fn-3.5&quot; class=&quot;footnote-ref&quot;&gt;3.5&lt;/a&gt;&lt;/sup&gt; on trying it, since it might immediately usher in truly human-like LLMs&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;I’ll skip (3), since I think the argument is still compelling without the analogy to human brains.&lt;/p&gt;
&lt;h3&gt;Are LLMs bad because they can’t grok?&lt;/h3&gt;
&lt;p&gt;I think his first point is hard to dispute. LLMs are very smart in specific areas, but they routinely make errors that humans wouldn’t make. More to the point, they routinely make errors that any human as smart as the LLM would &lt;em&gt;never&lt;/em&gt; make. This pretty clearly points to a failure of generalization: LLMs are as strong as smart humans in specific areas, but can’t generalize that intelligence to as many tasks as humans can.&lt;/p&gt;
&lt;p&gt;Do LLMs not grok? I read through &lt;a href=&quot;https://arxiv.org/pdf/2506.21551&quot;&gt;this paper&lt;/a&gt; that argues they do. If you graph “how much data has the LLM memorized” against benchmark performance, you can see a small initial spike in benchmark performance, followed by a big drop, followed finally by a big jump in benchmark performance. This pattern doesn’t track memorization at all: memorization increases smoothly in the background the whole time. &lt;/p&gt;
&lt;p&gt;&lt;span
      class=&quot;gatsby-resp-image-wrapper&quot;
      style=&quot;position: relative; display: block; margin-left: auto; margin-right: auto; max-width: 590px; &quot;
    &gt;
      &lt;a
    class=&quot;gatsby-resp-image-link&quot;
    href=&quot;/static/17d0f02f0691c797be502f32fbd33a40/1d499/llm-grokking.png&quot;
    style=&quot;display: block&quot;
    target=&quot;_blank&quot;
    rel=&quot;noopener&quot;
  &gt;
    &lt;span
    class=&quot;gatsby-resp-image-background-image&quot;
    style=&quot;padding-bottom: 39.189189189189186%; position: relative; bottom: 0; left: 0; background-image: url(&apos;data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAAICAYAAAD5nd/tAAAACXBIWXMAABYlAAAWJQFJUiTwAAAB5klEQVQozz2Sa2sTURCG83v1i4ofBIv/wIpfVFBQlCKtUKxYQtHGBLWYaktaL2Rrmk2a2+5m7/fb2cezCTgwnGGYeeedM28jSzPKsqQoSoQQZFm+eqM4xnVdLMfFth1M2yZJEgSVrC0oRYGoBMhMWcpYVNTWMKZLPNdnrE7RdANd8wnCkNHcYDSeMF3oWOYSZTZBVVUG8zG2Y+NEFsMrldliIPuGFHmxBoyTGN/35VQhGbhoxgF+EFKmAXkacj7bQ7NGSEIkaU7ma/iWShSlZLI3DFwuJk9w7RlJnNDQNJ0g8FboiveTbeURkR9DYdFSOlxvvaI/n0Lur2qcqMuf+Q5ZXMqcjZtM2Bps4ixnRHKzhmnZxPZYlmY8u/jA/d5bqjRlYcy49XKfG8/f81euTjxdAXa0PZrqLlVSrxhwOP7IvbMdyVqXrCMaQRAQS1evFO5sP+XhYYsssHj84B03r71g4/YWSn8EpfwCT2Wz95rt/leIPBa6zt3dN2zsN/EsjVgeslFPNQyT05MTvnw+4uzXUA5waDePpHfpHBzjeyGu43A5VGj3vnO5WBC6Jt+653xqn/Lj+Le8wVodK8A6qKUgREkl1teqrRA5S1PHtEwpHZu0lpiUlSykFkktl6qSEirz/z3/AA3rThqGWAwlAAAAAElFTkSuQmCC&apos;); background-size: cover; display: block;&quot;
  &gt;&lt;/span&gt;
  &lt;img
        class=&quot;gatsby-resp-image-image&quot;
        alt=&quot;llm-grokking&quot;
        title=&quot;llm-grokking&quot;
        src=&quot;/static/17d0f02f0691c797be502f32fbd33a40/fcda8/llm-grokking.png&quot;
        srcset=&quot;/static/17d0f02f0691c797be502f32fbd33a40/12f09/llm-grokking.png 148w,
/static/17d0f02f0691c797be502f32fbd33a40/e4a3f/llm-grokking.png 295w,
/static/17d0f02f0691c797be502f32fbd33a40/fcda8/llm-grokking.png 590w,
/static/17d0f02f0691c797be502f32fbd33a40/efc66/llm-grokking.png 885w,
/static/17d0f02f0691c797be502f32fbd33a40/c83ae/llm-grokking.png 1180w,
/static/17d0f02f0691c797be502f32fbd33a40/1d499/llm-grokking.png 1632w&quot;
        sizes=&quot;(max-width: 590px) 100vw, 590px&quot;
        style=&quot;width:100%;height:100%;margin:0;vertical-align:middle;position:absolute;top:0;left:0;&quot;
        loading=&quot;lazy&quot;
      /&gt;
  &lt;/a&gt;
    &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;I think this paper highlights the difficulty of distinguishing grokking from generalization. Obviously LLMs learn to generalize during training, and it’s plausible that learning to generalize would require a certain baseline level of memorization (so that the LLM has the raw material to generalize from). So it’s going to look like grokking.&lt;/p&gt;
&lt;p&gt;When Gwern (and others) say that LLMs don’t grok, I think what they mean is that there’s at least one more giant generalization leap waiting to be made. Is this plausible? As an existence proof, humans are clearly capable of better generalization than LLMs. Of course, it’s &lt;em&gt;possible&lt;/em&gt; that this level of human generalization comes from features of our brain that neural networks can’t replicate, but that seems kind of ad-hoc: if neural networks can generalize at all, why would they only be able to generalize this far, and no further?&lt;/p&gt;
&lt;p&gt;The easy examples of grokking rely on domains with a simple rule waiting to be discovered (e.g. a mathematical operation). Does human language have rules this deep? I think this is an open question, but there’s good reason to think the answer is yes. Language has deep, subtle structure: not just internal structure, but structure that reaches all the way down to the way the world is and the way human minds work.&lt;/p&gt;
&lt;h3&gt;AI labs train small-ish models on oceans of data&lt;/h3&gt;
&lt;p&gt;For the last few years, many AI researchers have been saying that data is the most important thing: that whatever model architecture you choose, with enough size and training time the model will &lt;a href=&quot;https://nonint.com/2023/06/10/the-it-in-ai-models-is-the-dataset/&quot;&gt;converge to its dataset&lt;/a&gt;. Whether this is true &lt;a href=&quot;https://x.com/YiTayML/status/1783273130087289021&quot;&gt;or not&lt;/a&gt;, AI labs have spent much of their considerable resources on acquiring more, higher-quality data: from &lt;a href=&quot;https://www.washingtonpost.com/technology/2026/01/27/anthropic-ai-scan-destroy-books/&quot;&gt;scanning physical books&lt;/a&gt;, paying experts to &lt;a href=&quot;https://www.herohunt.ai/blog/the-ultimate-ai-data-labeling-industry-overview/&quot;&gt;produce and label data&lt;/a&gt;, or partnering with &lt;a href=&quot;https://openai.com/index/openai-and-reddit-partnership/&quot;&gt;companies&lt;/a&gt; that have a lot of data already.&lt;/p&gt;
&lt;p&gt;AI labs have also been training &lt;em&gt;relatively&lt;/em&gt; small models. Even the largest frontier models are probably MoEs with a couple of trillion &lt;a href=&quot;https://news.ycombinator.com/item?id=47319205&quot;&gt;parameters&lt;/a&gt; and probably a tenth of that in active parameters. Of course, estimates of frontier model size are mostly guesswork, but open-source models provide a good baseline: they’re probably in the ballpark of Kimi-K3, which &lt;a href=&quot;https://platform.kimi.ai/docs/guide/kimi-k3-quickstart&quot;&gt;has&lt;/a&gt; just under three trillion parameters and fifty billion active parameters. That sounds like a lot, but it’s something you could probably pre-train in &lt;em&gt;a couple of days&lt;/em&gt; in the largest frontier cluster&lt;sup id=&quot;fnref-4&quot;&gt;&lt;a href=&quot;#fn-4&quot; class=&quot;footnote-ref&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;h3&gt;Grokking requires training a huge model on a small dataset&lt;/h3&gt;
&lt;p&gt;Gwern’s prediction is that AI labs should try doing the exact opposite of what they’ve been doing. Instead of training a bunch of trillion-parameter models on massive amounts of data, try training one hundred-trillion-parameter model on a small dataset. &lt;/p&gt;
&lt;p&gt;This sounds pretty silly on the face of it. The more data the model has access to, the smarter it will be, right? Why waste an entire training cluster on a hobbled training run? Because if Gwern is right, grokking is more likely to occur when the dataset is constrained&lt;sup id=&quot;fnref-5&quot;&gt;&lt;a href=&quot;#fn-5&quot; class=&quot;footnote-ref&quot;&gt;5&lt;/a&gt;&lt;/sup&gt;. If you feed the model all the data in the world, it can continue to improve simply by memorizing more new things or drawing simple connections. If the model has to ruminate on a small set of data, it’ll be forced to keep looking for deeper generalizations. You want a very large model for this so it can memorize as much of the data as possible. Every piece of memorized data can serve as raw material for generalizing.&lt;/p&gt;
&lt;p&gt;The big labs probably haven’t done this already. Plausibly Gwern himself is enough of an insider that he would know, and so him writing this post is evidence that the labs haven’t tried it. Also, the engineering problems involved in training a hundred-trillion-parameter model have likely not been solved yet: the largest existing model is probably Claude Mythos, which is definitely not that big. But they have the resources and engineering talent to give it a pretty good shot.&lt;/p&gt;
&lt;p&gt;Interestingly, the political obstacles might be as hard to solve as the technical ones. This training run is going to look like it failed until the moment it succeeds: training loss will drop to zero relatively quickly, then sit there for weeks or months apparently doing nothing at all to improve test loss, chewing up billions of dollars. Do any of the top players have the risk appetite or courage to keep funding this experiment all that time?&lt;/p&gt;
&lt;h3&gt;Conclusion&lt;/h3&gt;
&lt;p&gt;Gwern’s post has an extended argument that human brain development works in the same way: that human brains have far more “parameters” than frontier LLMs, and are trained on far less data&lt;sup id=&quot;fnref-6&quot;&gt;&lt;a href=&quot;#fn-6&quot; class=&quot;footnote-ref&quot;&gt;6&lt;/a&gt;&lt;/sup&gt;, which encourages us to make deeper generalizations in early childhood. I don’t have the background in biology or neuroscience to evaluate these claims, so I’ve expressed the case for grokking entirely without reference to it.&lt;/p&gt;
&lt;p&gt;In 2024, it became clear to everyone that “pure scaling” — the idea that you could simply train larger and larger versions of GPT-3.5 — didn’t work. OpenAI’s “even bigger version” of GPT-4 was simply not good enough, and was eventually released as GPT-4.5 instead of GPT-5. The biggest advances since then have been reasoning, which produced another great leap forward in capability, and much better automated RL, which has ushered in the current era of reliable agents. Neither of these seem like a plausible path to artificial superintelligence.&lt;/p&gt;
&lt;p&gt;I don’t know if I agree with Gwern or not, but forcing very large LLMs to grok is at least an idea that &lt;em&gt;could&lt;/em&gt; usher in the machine god. I can’t remember the last time I read about a simple idea this ambitious&lt;sup id=&quot;fnref-7&quot;&gt;&lt;a href=&quot;#fn-7&quot; class=&quot;footnote-ref&quot;&gt;7&lt;/a&gt;&lt;/sup&gt;. I hope one of the big labs tries it out.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;For an example of the power of unhobbling, consider Claude Code or OpenClaw and the subsequent explosion of (short and long running) agentic harnesses.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;Obviously “motivate” and “notices” are used metaphorically.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;All of this is long before xAI’s use of the word “Grok” to name its LLMs. (Incidentally, I think this is why Gwern uses “catapulting” to describe the same thing).&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3.5&quot;&gt;
&lt;p&gt;For what it’s worth, Fable estimated the cost of Gwern’s plan at $3-10B.&lt;/p&gt;
&lt;a href=&quot;#fnref-3.5&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-4&quot;&gt;
&lt;p&gt;At this model size, 25T tokens of training data at 33% utilization works out to around six million H100-hours, which a 100k GPU cluster puts out every two and a half days.&lt;/p&gt;
&lt;a href=&quot;#fnref-4&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-5&quot;&gt;
&lt;p&gt;Two interesting pieces of contrary evidence here. First, &lt;a href=&quot;https://babylm.github.io/&quot;&gt;BabyLM&lt;/a&gt; is a yearly challenge to train a strong model on a &lt;em&gt;very&lt;/em&gt; small dataset. This has been running for four years and largely &lt;a href=&quot;https://aclanthology.org/2025.babylm-main.28/&quot;&gt;does not work&lt;/a&gt; (that is, nobody seems to have developed a model that shows a quantum leap forward in generalization). Second, &lt;a href=&quot;https://arxiv.org/pdf/2305.16264&quot;&gt;this paper&lt;/a&gt; tries training a 9 billon parameter model on constrained data and doesn’t see a big jump. I think Gwern’s response would be that these models are far too small — they can’t memorize enough of the training data to grok it, and arguable haven’t trained for long enough.&lt;/p&gt;
&lt;a href=&quot;#fnref-5&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-6&quot;&gt;
&lt;p&gt;A common objection here is to say that humans get infinitely more sensory data from the nuances of vision, touch, sound, and so on. I agree with Gwern that this is unconvincing: sensory data is largely predictable, text is surprisingly information-dense, and if this were true then deaf/blind people would have significantly less fluid intelligence (&lt;a href=&quot;https://pmc.ncbi.nlm.nih.gov/articles/PMC11165843/pdf/13023_2024_Article_3222.pdf&quot;&gt;they don’t&lt;/a&gt;).&lt;/p&gt;
&lt;a href=&quot;#fnref-6&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-7&quot;&gt;
&lt;p&gt;Maybe state-space-reasoning a la &lt;a href=&quot;https://www.ibm.com/think/topics/mamba-model&quot;&gt;Mamba&lt;/a&gt;, which didn’t work (yet).&lt;/p&gt;
&lt;a href=&quot;#fnref-7&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[What does "playing politics" mean for software engineers?]]></title><link>https://seangoedecke.com/playing-politics/</link><guid isPermaLink="false">https://seangoedecke.com/playing-politics/</guid><pubDate>Tue, 14 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Software engineers are &lt;a href=&quot;https://old.reddit.com/r/ExperiencedDevs/comments/1urg0tk/whats_the_best_advice_youve_received_from_a/owfi7dq/&quot;&gt;often told&lt;/a&gt; to “start playing politics”, but most engineers have no idea what that means.&lt;/p&gt;
&lt;p&gt;Their reference point for “playing politics” comes from fiction like Game of Thrones. Are they supposed to raise an army and depose the CEO, or poison each other at team lunch? Should they book Zoom calls with each other and plot schemes? All of that is obviously ridiculous. In terms of Game of Thrones, software engineers are not lords and ladies. We’re the soldiers and workers of the realm. So you should think about “playing politics” in the way a castle guard would, not one of the major players.&lt;/p&gt;
&lt;p&gt;The castle guard are not going around poisoning people or forming coalitions between the great powers. They are largely keeping their heads down. But in order to do that, they have to stay aware of the political currents, or they’re liable to do something catastrophically stupid: for instance, making an enemy of a powerful courtier, or arresting somebody who’s on an important mission for the king.&lt;/p&gt;
&lt;p&gt;Given that, the basic principles of playing politics are something like this:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Be aware of who’s powerful and who’s not&lt;/li&gt;
&lt;li&gt;At all costs, avoid making powerful enemies&lt;/li&gt;
&lt;li&gt;Help powerful people as best you can&lt;/li&gt;
&lt;li&gt;Make sure they know you’re helping them (without annoying them)&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Be aware of who’s powerful and who’s not&lt;/h3&gt;
&lt;p&gt;As a software engineer in a large company, &lt;strong&gt;you will not be a powerful person&lt;/strong&gt;. Powerful people are typically in senior management: VPs, directors, and so on&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;. However, not everyone in senior management is powerful. Some are killers who have the active support of the CEO, while others are confused incompetents.&lt;/p&gt;
&lt;p&gt;How do you know which is which? If someone is clearly ferociously competent, they’re always going to have &lt;em&gt;some&lt;/em&gt; power, since upper management tend not to ignore useful tools. But you can’t rely on competence as your only guide. Some managers are powerful for other reasons: they’re friends with the CEO, or they have strong relationships with other groups like legal or sales, or they’re simply willing to do whatever upper management wants done.&lt;/p&gt;
&lt;p&gt;One signal is who’s leading the important projects. Read your CEO or CTO’s internal updates and pay attention to the projects that are called out by name. Organizations tend to give key tasks to trusted lieutenants. If a manager is leading an area that’s never under &lt;a href=&quot;/the-spotlight/&quot;&gt;the spotlight&lt;/a&gt;, they probably don’t have enough clout.&lt;/p&gt;
&lt;p&gt;Another signal is hiring. Is a manager’s team growing or shrinking? Particularly &lt;a href=&quot;/good-times-are-over/&quot;&gt;post-ZIRP&lt;/a&gt;, headcount is a rare and precious resource. A manager who’s able to get it is likely a powerful manager, or at least is reporting to a powerful director or VP (which often amounts to the same thing).&lt;/p&gt;
&lt;h3&gt;At all costs, avoid making powerful enemies&lt;/h3&gt;
&lt;p&gt;First, you should try not to make any enemies at all. Most software engineers who get “playing politics” wrong do it by needlessly alienating people: by being rude, unhelpful, abrasive, making non-technical people feel stupid, and so on. This post isn’t really about that. I’m assuming that you can figure out how to be a generically pleasant person on your own.&lt;/p&gt;
&lt;p&gt;However, &lt;strong&gt;competent software engineers will make some enemies&lt;/strong&gt;. If you’re out there making projects happen, some people aren’t going to like the way you do it, and won’t be a fan of any compromise you offer. I wrote about this in &lt;a href=&quot;/big-tech-needs-big-egos/&quot;&gt;&lt;em&gt;Big tech engineers need big egos&lt;/em&gt;&lt;/a&gt;: the only way to avoid making enemies is to change nothing, but that’s incompatible with doing the job.&lt;/p&gt;
&lt;p&gt;Given that, be selective about &lt;em&gt;which&lt;/em&gt; enemies you make. If you’re making a technical decision that’s either going to require work from team A or team B, and neither team wants to do it, you should try to pick the team with the least political cover. If you need a powerful VP’s team to do something they won’t like, try to be maximally respectful about it: get that team’s core engineers on-side if you can, or book a meeting with the powerful manager and explain the situation, or (better yet) ask the powerful manager sponsoring your project to go and talk to the other VP for you. (If you don’t have a powerful manager like this, consider abandoning your project).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Give way to powerful managers when at all possible.&lt;/strong&gt; Every so often you really do have to stand your ground — if the system will truly collapse otherwise, or a major customer will have an incident, or if the technical decision really is entirely bone-headed — but almost all cases are not like this. The best advice I’ve ever gotten about playing politics came from a manager I worked with long ago&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;This is not the hill you want to die on.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;When I’m about to pick a fight or say something argumentative, and I’m not 100% convinced it’s necessary, I ask myself: is this the hill I want to die on? And it never is.&lt;/p&gt;
&lt;p&gt;The three rules about disagreeing with powerful people are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Make sure you do it in private&lt;/li&gt;
&lt;li&gt;Be polite&lt;/li&gt;
&lt;li&gt;When they overrule you, stop arguing immediately&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Disagreeing in private rarely hurts, if you follow these rules. In fact, it can help. If you can manage to disagree with a manager, get overruled, and then follow their plan without complaining, that can be the best way to gain a powerful friend. But if they think you’re going to keep griping about it, or worse still, complain to the rest of the team and foment some kind of rebellion, there’s no quicker way to make a powerful enemy.&lt;/p&gt;
&lt;p&gt;If you have powerful enemies at a company (for instance, the CTO or an influential VP doesn’t like you), &lt;strong&gt;quit&lt;/strong&gt;. It’s really that bad. I have never seen this situation turn itself around, except in the very rare case where the CTO or VP is already looking for greener pastures and jumps ship. You cannot recover the situation: they have no incentive to give you the chance to change their mind, and they have almost unlimited ability to screw you on promotions, raises and layoffs.&lt;/p&gt;
&lt;p&gt;That’s why this piece of advice is second in the list. If you aren’t helpful or if your contributions are invisible, you can work on that and fix it. But if you’ve made powerful enemies, you’re done for.&lt;/p&gt;
&lt;h3&gt;Help powerful people as best you can&lt;/h3&gt;
&lt;p&gt;Just as it’s fatal to make powerful enemies, it’s very useful to make powerful friends. How can you do this? Remember you’re a palace guard, not a great lord: you make friends &lt;strong&gt;by doing your job&lt;/strong&gt;. However, you can choose to do your job a little more proactively and diligently when you’re doing it for someone with political clout.&lt;/p&gt;
&lt;p&gt;One obvious application of this principle is that &lt;strong&gt;you should answer Slack messages from powerful people immediately&lt;/strong&gt;. If you see an ordinary Slack question pop up while you’re doing some task, it’s okay to get to it when you get to it. In fact, it’s ideal &lt;em&gt;not&lt;/em&gt; to respond to all questions immediately, so you don’t set unreasonable expectations (and so you don’t seem like you’re sitting around doing nothing). But when a VP comes in with a question, don’t make them wait: answer the question immediately. If the question requires research, send a “let me look into that right now” message, then do the research. This is the easiest way to get a reputation for being helpful&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;Another way to do this is to &lt;strong&gt;lean in on important projects&lt;/strong&gt;. Suppose you do ten projects in a year. Eight of them are normal, low-priority projects, and two of them are high-profile (say, finishing some big feature before your company’s yearly conference). It’s a mistake to allocate your effort equally to all ten. I wrote about this at length in &lt;a href=&quot;/doing-nothing-at-work/&quot;&gt;&lt;em&gt;Doing nothing at work&lt;/em&gt;&lt;/a&gt;: you should be operating at 80% capacity (or less), so you can then ramp up to 120% when it really matters.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pay attention to the narrative that powerful people are trying to push.&lt;/strong&gt; Here are some potential narratives:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;We’ve had a lot of turnover and reorgs lately, but we’re all starting to pull together as a team now&lt;/li&gt;
&lt;li&gt;Isn’t it great how focused we all are on reliability work after last month’s incident?&lt;/li&gt;
&lt;li&gt;The conference this week is the most important thing, so we’re all being very careful not to break anything&lt;/li&gt;
&lt;li&gt;We’re an AI-forward team that’s looking for the best ways we can leverage LLMs into our team processes&lt;/li&gt;
&lt;li&gt;Although this project had a rocky start, we’re now all aligned on the way forward&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;You don’t necessarily have to jump in and start cheerleading, but you should at least not do anything that you know is going to make the narrative look weak. For example, on that last point, it’s foolish to openly argue that the project really was fine all along. Bring it up privately, not publicly, or you risk ruining some clever piece of propaganda that the manager in question is trying to push on the rest of the organization&lt;sup id=&quot;fnref-4&quot;&gt;&lt;a href=&quot;#fn-4&quot; class=&quot;footnote-ref&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;Finally, an underrated way to help powerful people is to offer them social support and information. Slack messages and planning emails might seem unimportant to you, but powerful people often live in that environment: their primary tool is writing messages like these, just like your primary tool is writing code. Reading and responding (in a supportive way) to these messages is something that most engineers don’t bother to do, but it goes a long way.&lt;/p&gt;
&lt;p&gt;Likewise, dropping a senior manager a line now and then (say, a heads-up that a particular project landed successfully, or that you got good metrics about some feature) is surprisingly helpful. Senior managers live in an information-poor environment: for them to learn something about a team’s work, that information has to bubble up through several layers of interpretation and summary. In my experience, they’re appreciative of being drip-fed the occasional piece of information, so long as you keep it brief and relatively rare.&lt;/p&gt;
&lt;h3&gt;Make sure they know you’re helping them&lt;/h3&gt;
&lt;p&gt;If you’re directly responding to a VP’s Slack messages or DMing them information, they know you’re the one doing it. But if you’re just doing your job and working hard on projects they care about, they might not notice. &lt;strong&gt;Being invisible is probably the most common way engineers fail at playing politics.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Fortunately the fix is simple: tell people what you’re doing. If you fix an important bug for a launch, write a message in that launch’s Slack channel saying “hey, I just fixed this bug”. What if you don’t like bragging? Get over it. You have to be comfortable publicly telling people what you’ve done. You should also keep a &lt;a href=&quot;https://jvns.ca/blog/brag-documents/&quot;&gt;brag document&lt;/a&gt; so you can repeat all of this at review time.&lt;/p&gt;
&lt;p&gt;Another, subtler way to do this is to gain the trust and respect of the powerful engineers in your area. Senior managers will always have a few trusted engineers they rely on to assess technical questions. They will ask those engineers what they think about you, and will broadly trust those answers. The good news is that if you’re competent and useful, those engineers will already value you, so you don’t have to do anything special: just be good at your job.&lt;/p&gt;
&lt;h3&gt;Technical power&lt;/h3&gt;
&lt;p&gt;Is playing politics all about sucking up to senior managers? Basically, yeah. A less cynical way to &lt;a href=&quot;/shareholder-value/&quot;&gt;describe it&lt;/a&gt; would be “aligning with the values of the company”. If you think your company is doing good things, you should want to do that anyway! In any case, what that comes down to is figuring out what the people in charge want, giving it to them, and making sure they see you doing it. However, there’s still some scope to get what &lt;em&gt;you&lt;/em&gt; want out of the deal.&lt;/p&gt;
&lt;p&gt;I said earlier that software engineers do not wield organizational power. However, that doesn’t mean you’re powerless. Technical ability is a source of real power, if a delicate and unreliable one. The movers and shakers in tech companies are utterly dependent on technical people to implement their vision and to give them clear answers about the system.&lt;/p&gt;
&lt;p&gt;There are many subtle ways you can leverage this. One I wrote about in &lt;a href=&quot;/how-to-influence-politics/&quot;&gt;&lt;em&gt;How I influence tech company politics as a staff software engineer&lt;/em&gt;&lt;/a&gt; is to wait until important people at the company want to do something (say, improve reliability), then offer them a technical plan that does it your way. Another one is to become so useful that you’re actively in demand to lead projects, and then run the project how you want.&lt;/p&gt;
&lt;p&gt;You probably won’t be able to change the company’s grand strategy. But how that strategy is &lt;em&gt;implemented&lt;/em&gt; has a lot of specific technical detail, and you can put yourself in a position to decide on those details.&lt;/p&gt;
&lt;h3&gt;Conclusion&lt;/h3&gt;
&lt;p&gt;Playing politics isn’t about plotting and scheming, and it isn’t just about being a &lt;a href=&quot;https://en.wikipedia.org/wiki/How_to_Win_Friends_and_Influence_People&quot;&gt;friendly, likeable person&lt;/a&gt; (although that helps). It’s about figuring out how your company actually operates: who makes the decisions, who gets consulted, what behavior gets rewarded, and so on. The most basic way to do that is to &lt;strong&gt;figure out who is powerful, get out of their way, and (if you can) help them get what they want&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;edit: this post got some comments on &lt;a href=&quot;https://news.ycombinator.com/item?id=48905390&quot;&gt;Hacker News&lt;/a&gt;. I always enjoy seeing “this is too obvious to need a guide” &lt;a href=&quot;https://news.ycombinator.com/item?id=48905622&quot;&gt;replies&lt;/a&gt; right next to “yeah, I screwed this up, wish I’d been told this sooner” &lt;a href=&quot;https://news.ycombinator.com/item?id=48905675&quot;&gt;replies&lt;/a&gt;. That’s what I get for &lt;a href=&quot;/saying-the-obvious-thing/&quot;&gt;saying the obvious thing&lt;/a&gt;.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;Obviously the exact titles depend on your company. One person I’m deliberately leaving out is your own manager. In general don’t think your relationship with your own manager counts as “playing politics”: that’s just you getting along with another human being. An exception to that is if you report directly to a powerful director or VP.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;Ironically, this manager struggled to take his own advice.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;Note that you actually have to be able to answer their question accurately in order to do this. If you’re not competent enough to be useful to powerful people, you will struggle to befriend them.&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-4&quot;&gt;
&lt;p&gt;For instance, maybe the CEO is convinced that the project was in bad shape because of something he heard, and the manager in question knows it’s easier to sell “yes, but we turned it around” than “no, you misunderstood, everything was always fine”. If you complicate that process, you risk the CEO thinking that the project is still bad and cancelling it.&lt;/p&gt;
&lt;a href=&quot;#fnref-4&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[In defense of not understanding your codebase]]></title><link>https://seangoedecke.com/in-defense-of-not-understanding-your-codebase/</link><guid isPermaLink="false">https://seangoedecke.com/in-defense-of-not-understanding-your-codebase/</guid><pubDate>Sat, 11 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;As a software engineer, how well do you have to understand your own codebase?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;My guess is that people who work on small codebases with low-turnover teams (say, &lt;a href=&quot;https://redis.io/&quot;&gt;Redis&lt;/a&gt; or games like &lt;a href=&quot;https://en.wikipedia.org/wiki/The_Witness_(2016_video_game)&quot;&gt;The Witness&lt;/a&gt;) would say “obviously you have to understand it completely, otherwise you can’t do good work”. I’d also guess that people who work on large codebases with high-turnover teams (say, the Google web search backend or GitHub) would say “obviously you can’t understand it completely, you just have to do the best you can in your local area”.&lt;/p&gt;
&lt;p&gt;These are two largely different ways of programming with different methods, practices and cultures&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;. However, the first group is over-represented in online discussion about software engineering&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;. I want to defend the second group against the first. In many software engineering environments, there’s nothing wrong with being in a state of &lt;em&gt;partial&lt;/em&gt; understanding. In fact, in large systems a partial understanding is the best you can do.&lt;/p&gt;
&lt;h3&gt;Against “programming as theory building”&lt;/h3&gt;
&lt;p&gt;The best articulation of the “you have to understand your codebase” side is Peter Naur’s famous paper &lt;a href=&quot;https://pages.cs.wisc.edu/~remzi/Naur.pdf&quot;&gt;&lt;em&gt;Programming as Theory Building&lt;/em&gt;&lt;/a&gt;. I like this paper, but I think it goes too far in that direction. Naur’s core point is that when programmers work on a program, the code is really just a by-product, and the main product they’re working on is their “theory of the program”. That’s made up of their intuitive sense of what’s happening and why, which can only be partially captured by code or documentation. If they lost the code, they could rewrite the program easily. If they lost their understanding (say, if the team experienced 100% turnover), they would struggle to make sense of the code.&lt;/p&gt;
&lt;p&gt;So far, so good, but Naur goes further than this. He says that the theory &lt;em&gt;should not&lt;/em&gt; be reconstructed from the code. According to Naur, &lt;strong&gt;you’re better off scrapping the program entirely and having a new team rebuild it from scratch&lt;/strong&gt;, building up a new theory in the process&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;reestablishing the theory of a program merely from the documentation, is strictly impossible … [therefore] the existing program text should be discarded and the new-formed programmer team should be given the opportunity to solve the given problem afresh&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Anyone who’s been an effective software engineer at a large company knows that Naur is dead wrong about this. There are at least two reasons.&lt;/p&gt;
&lt;p&gt;First, &lt;strong&gt;you simply can’t rebuild large software systems from scratch&lt;/strong&gt;. Sufficiently large systems (if they have users) contain thousands of &lt;a href=&quot;/wicked-features/&quot;&gt;weird cases&lt;/a&gt; and quirks that cannot be reimplemented. Even a team that’s intimately familiar with the system couldn’t do it: there’s just too much &lt;em&gt;stuff&lt;/em&gt; to juggle. Successful rewrites always start by carving out the existing codebase into small isolated chunks, then rewriting one chunk at a time. In other words, rewriting a software system involves making a bunch of changes to the old system. If you can’t change the old system, you certainly can’t replace it with a new one.&lt;/p&gt;
&lt;p&gt;Second, &lt;strong&gt;abandoned systems are revived &lt;em&gt;all the time&lt;/em&gt;&lt;/strong&gt;. In a tech company with hundreds of millions of lines of code and thousands of engineers, it’s not uncommon for a codebase to have nobody left who’s familiar with it&lt;sup id=&quot;fnref-4&quot;&gt;&lt;a href=&quot;#fn-4&quot; class=&quot;footnote-ref&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;. All it takes is a few people to quit at the wrong time, or for a codebase to be unmaintained for a year. Not only have I seen other teams do this, I have &lt;em&gt;personally&lt;/em&gt; taken ownership of abandoned codebases, figured them out, and gotten to a point where I could effectively work with them. It takes time, but building a new theory of the codebase is possible. You start by understanding one flow end-to-end, then slowly branch out from there, making careful changes as you go.&lt;/p&gt;
&lt;p&gt;In sufficiently large codebases, &lt;strong&gt;everyone operates with an incorrect theory of the program&lt;/strong&gt;. The defining feature of modern software systems is that they’re just way too big for anyone (or even a whole team) to keep in their head: &lt;a href=&quot;/nobody-knows-how-software-products-work/&quot;&gt;nobody understands it all&lt;/a&gt;. To be effective, you have to figure out a way to work with a merely partially-correct theory. This is why I keep going on about &lt;a href=&quot;/taking-a-position/&quot;&gt;taking a position&lt;/a&gt; and &lt;a href=&quot;/what-makes-strong-engineers-strong/&quot;&gt;confidence&lt;/a&gt;. If you’re not sure about something, you can’t just sit back and wait for someone with a perfect understanding to come and give you the answer. If you’re a competent engineer, &lt;em&gt;that person is you&lt;/em&gt;. You have to grit your teeth, make your most educated guess, and then deal with the consequences.&lt;/p&gt;
&lt;p&gt;To be generous to Naur, it’s possible that in 1985 the average size of a program was several orders of magnitude smaller than today, and that when Naur writes about “large programs” he’s not talking about tens of millions of lines of code. Naur’s first example of a large program is a 200,000 line industrial monitoring program, and his second example is a compiler. In 1987, the first version of the compiler GCC was about a &lt;a href=&quot;https://www.oreilly.com/openbook/freedom/ch09.html&quot;&gt;hundred thousand&lt;/a&gt; lines of code; in 2015 GCC was over &lt;a href=&quot;https://www.phoronix.com/news/MTg3OTQ&quot;&gt;fourteen million&lt;/a&gt; lines. I can believe that rewriting one or two hundred thousand lines of code is relatively straightforward, particularly if you get to reuse existing tests. Not so for one or two million.&lt;/p&gt;
&lt;h3&gt;Theory building is one tradeoff among many&lt;/h3&gt;
&lt;p&gt;LLMs are &lt;a href=&quot;https://ratfactor.com/cards/naur-vs-llms&quot;&gt;often cited&lt;/a&gt; as a tool that’s bad because it impedes the ordinary process of theory-building. I think this is overly simplistic. Like many software tools, LLMs are a double-edged sword: they make it harder to construct a detailed mental theory of the software, but they allow you to build a partial theory quickly and they can help you leverage that partial theory more effectively. This is a complex tradeoff that I’m still thinking about.&lt;/p&gt;
&lt;p&gt;Setting LLMs aside, I’m confident that it’s silly to say that anything that interferes with your theory of the software must be bad. Here is a partial list of other things that make it harder to maintain a theory:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Other people being allowed to write code in your codebase&lt;/li&gt;
&lt;li&gt;Having to implement legally-required features like accessibility and data protection&lt;/li&gt;
&lt;li&gt;Allowing your colleagues to quit their jobs or move between teams&lt;/li&gt;
&lt;li&gt;Having to upgrade software versions for security patches&lt;/li&gt;
&lt;li&gt;Bringing in libraries or other dependencies&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Like most things in software, “maintaining a theory of the codebase” is one value among many. Sometimes it’s the most important value and you sacrifice other values for it; other times you trade it off for speed, or legal compliance, or for political reasons&lt;sup id=&quot;fnref-5&quot;&gt;&lt;a href=&quot;#fn-5&quot; class=&quot;footnote-ref&quot;&gt;5&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;Almost all engineers — particularly &lt;a href=&quot;/pure-and-impure-engineering/&quot;&gt;“pure”&lt;/a&gt; engineers — prefer to maintain an accurate mental model of their software. It’s more fun, less stressful, and feels more like “real engineering”. That’s why many engineers take up open-source projects in their spare time in order to work on small codebases by themselves: in order to do engineering work where they can maintain an accurate Naur theory of the codebase. I don’t think there’s anything wrong with that.&lt;/p&gt;
&lt;p&gt;However, at work &lt;a href=&quot;/where-the-money-comes-from/&quot;&gt;you are paid to do a job&lt;/a&gt;. In other words, they pay you money to adopt &lt;em&gt;their&lt;/em&gt; set of engineering values. It’s hopefully well-understood that however much you might personally care about performance, sometimes you have to write slow code at your job (for instance, to get a project done on time, or to accommodate some awkward requirement). Maintaining a theory of the codebase is the same kind of thing. &lt;/p&gt;
&lt;p&gt;edit: this post got some comments on &lt;a href=&quot;https://lobste.rs/s/elhi7o/defense_not_understanding_your_codebase&quot;&gt;lobste.rs&lt;/a&gt;. One interesting &lt;a href=&quot;https://lobste.rs/c/qjfhxd&quot;&gt;comment&lt;/a&gt; points out that the ability to reason “locally” about code (i.e. with a partial understanding) has been a core goal of CS from the beginning. &lt;a href=&quot;https://lobste.rs/c/gr8hgw&quot;&gt;This&lt;/a&gt; is also a good description of what I was trying to get at in &lt;a href=&quot;/bad-code-at-big-companies/&quot;&gt;&lt;em&gt;How good engineers write bad code at big companies&lt;/em&gt;&lt;/a&gt;. Also, it’s amusing that this post was tagged as &lt;code class=&quot;language-text&quot;&gt;vibecoding&lt;/code&gt; because of one off-hand paragraph about LLMs. I still don’t think I’ll be tagging the post as &lt;a href=&quot;/tags/ai/&quot;&gt;AI&lt;/a&gt; on my blog, sorry.&lt;/p&gt;
&lt;p&gt;edit: I also got some &lt;a href=&quot;https://news.ycombinator.com/item?id=48882777&quot;&gt;Hacker News&lt;/a&gt; comments. The &lt;a href=&quot;https://news.ycombinator.com/item?id=48932402&quot;&gt;top comment&lt;/a&gt; is a genre of comment I get a lot, which is basically “wait, this situation sucks! Why isn’t this blog post about how much this sucks?” Well, there’s plenty of posts like that already: I hope to fill another niche. Another &lt;a href=&quot;https://news.ycombinator.com/item?id=48882903&quot;&gt;comment&lt;/a&gt; offers the second comparison of me with Seth Godin (ouch) that I’ve seen:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Goedecke doesn’t quite write the anodyne sound bites that Seth Godin does, but neither does he write anything of engineering use, just vocabulary explainers for people who want to know kind of what their tech leads and line managers are talking about.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;First, this is skewed by what kind of posts get popular on Hacker News (i.e. not my &lt;a href=&quot;/interaction-models/&quot;&gt;posts&lt;/a&gt; &lt;a href=&quot;/steering-vectors/&quot;&gt;that&lt;/a&gt; &lt;a href=&quot;/fast-llm-inference/&quot;&gt;discuss&lt;/a&gt; &lt;a href=&quot;/ai-detection/&quot;&gt;technical&lt;/a&gt; &lt;a href=&quot;/tempo-faq/&quot;&gt;engineering&lt;/a&gt; &lt;a href=&quot;/tags/papers/&quot;&gt;topics&lt;/a&gt;). Second, I think “wanting to know what your tech leads and line managers are talking about” is very important!&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;I wrote about this at length in &lt;a href=&quot;/pure-and-impure-engineering/&quot;&gt;&lt;em&gt;Pure and impure software engineering&lt;/em&gt;&lt;/a&gt;. I think many of the repeated arguments we have in the software industry are caused by the pure total-understanding culture coming up against the impure partial-understanding culture.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;Open-source engineers are more excited to blog about their work, the raw engineering content is typically more impressive (because coordination problems dominate big proprietary systems), open-source projects can be legally written about while proprietary systems can’t, and even if you could do it legally, writing about large codebases is impossible because it requires too much &lt;a href=&quot;/you-cant-design-software-you-dont-work-on/&quot;&gt;specific context&lt;/a&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;I re-read the relevant chapters of Ryle’s &lt;a href=&quot;https://www.andrew.cmu.edu/user/kk3n/80-300/ryle1949.pdf&quot;&gt;&lt;em&gt;The Concept of Mind&lt;/em&gt;&lt;/a&gt; (which Naur cites throughout) and I think Ryle is more generous about theory-building. For Ryle, theory-building or know-how automatically happens as you do things. It’s fully consistent with Ryle to think you can pick up an existing codebase just from the code, purely by puzzling it out.&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-4&quot;&gt;
&lt;p&gt;Naur says: “Lest this consequence may seem unreasonable, it may be noted that the need for revival of an entirely dead program probably will rarely arise, since it is hardly conceivable that the revival would be assigned to new programmers without at least some knowledge of the theory had by the original team.”. If only!&lt;/p&gt;
&lt;a href=&quot;#fnref-4&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-5&quot;&gt;
&lt;p&gt;Some engineers might say that maintaining a theory is the &lt;em&gt;core&lt;/em&gt; value, because without it you can’t fulfill any of the others. I disagree. You could say the same thing about readability, or maintainability, or correctness, or a bunch of other engineering values. We trade off “core” values like this all the time.&lt;/p&gt;
&lt;a href=&quot;#fnref-5&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[Blog about things you don't understand yet]]></title><link>https://seangoedecke.com/blog-about-things-you-dont-understand-yet/</link><guid isPermaLink="false">https://seangoedecke.com/blog-about-things-you-dont-understand-yet/</guid><pubDate>Tue, 07 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Every post I publish represents at least two things I’ve learned: the thing that prompted me to write the post, and the thing I learned in the course of writing it. If I don’t learn anything new while I’m writing, it’s not interesting enough to publish.&lt;/p&gt;
&lt;p&gt;Typically I learn way more than two things. For instance, in my &lt;a href=&quot;/the-o3-geoguessr-prompt-did-not-work/&quot;&gt;o3 geoguessr&lt;/a&gt; post, I started out with the idea that most AI prompts probably don’t work, and I ended up learning that newer OpenAI models have lost o3’s ability to geolocate. That’s interesting! In my most recent post on &lt;a href=&quot;/c2pa-only-works-if-everything-is-signed/&quot;&gt;C2PA&lt;/a&gt;, I started out with the idea that C2PA requires near-universal adoption, but I learned a &lt;em&gt;ton&lt;/em&gt; of things about PKI, managing private keys on local devices, how C2PA actually works, and so on. In my post on the &lt;a href=&quot;/luddites-and-ai-datacenters/&quot;&gt;Luddites&lt;/a&gt;, I started out with the idea that the Luddite movement was fundamentally decentralized, but ended up fascinated by Luddite culture (which was far more elitist, misogynist, and violent than the pop-Luddism books describe). I could do this for every single post on the blog.&lt;/p&gt;
&lt;h3&gt;Taking a position&lt;/h3&gt;
&lt;p&gt;I think the core reason this works is that &lt;strong&gt;every single one of my blog posts argues a point&lt;/strong&gt;. I never publish a post that just gives some scattered thoughts on a topic, or a post that only says “yes, I agree with this other article”. If I write a draft that nobody sensible could disagree with, I scrap the draft. Making sure that everything I write is at least minimally controversial is a forcing function: it forces me to think about what the most interesting part of my position is, and it forces me to do enough research to defend it against the obvious criticisms.&lt;/p&gt;
&lt;p&gt;This is contrary to a lot of advice I read about blogging, which encourages the aspiring blogger to treat their posts as a form of unstructured self-expression. If unstructured self-expression is what you want to do, that’s cool. The point of having a blog is that you get to write what &lt;em&gt;you&lt;/em&gt; want. However, this advice isn’t as helpful as it sounds.&lt;/p&gt;
&lt;p&gt;Before I was in tech, I was a philosophy grad student. But before &lt;em&gt;that&lt;/em&gt;, I was a poet. One thing you learn when you try to write poetry is that it is way easier to write to a restrictive structure than it is to simply “write what you feel”. This should be obvious when you actually think about it. The task of a poet is to repeatedly choose the next word. Writing to a structure (typically rhyme or meter) narrows that choice to a small set of words, instead of the entire English language. It’s the same with blogging. Forcing yourself to write about specific, potentially-controversial points makes consistently writing easier, not harder.&lt;/p&gt;
&lt;h3&gt;Writing, thinking, and research&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Writing is the best way to think clearly about a topic.&lt;/strong&gt; It’s easy to believe you understand something when you’re just turning it over in your head. When you have to condense that down into words, you find out exactly how much you do or don’t understand. I am constantly having moments where I type something, stop myself, and think “wait, that can’t actually be right”, or “is that really true?”&lt;/p&gt;
&lt;p&gt;By the time I write my way to the end of the post, I’m usually thinking so much more clearly about the topic that my conclusion paragraph is way better than my introduction. In fact, I’ve picked up the habit of going back and immediately rewriting the first paragraph as part of my first-draft process, because I know I’m going to end up doing it anyway.&lt;/p&gt;
&lt;p&gt;I also change my mind a lot while I write. &lt;a href=&quot;/space-ai-datacenters-do-not-have-a-cooling-problem/&quot;&gt;Here&lt;/a&gt; &lt;a href=&quot;https://github.com/sgoedecke/gatsby-blog/blob/2841c8504fc0b5f4dd2e8105955350bafec4d904/content/drafts/_icebox/prediction-markets-insider-trading/index.md&quot;&gt;are&lt;/a&gt; &lt;a href=&quot;/the-just-say-no-engineer-was-a-zirp-phenomenon/&quot;&gt;a&lt;/a&gt; &lt;a href=&quot;/giving-llms-a-personality/&quot;&gt;bunch&lt;/a&gt; &lt;a href=&quot;/ai-detection/&quot;&gt;of&lt;/a&gt; &lt;a href=&quot;/tempo-faq/&quot;&gt;examples&lt;/a&gt; &lt;a href=&quot;/impact-of-ai-study/&quot;&gt;of&lt;/a&gt; &lt;a href=&quot;/ai-interpretability/&quot;&gt;posts&lt;/a&gt; where I began writing them with the opposite opinion to the one that eventually made it into the post. I think this is a good sign, and I hope I never stop doing it. You should be researching and thinking about every post you write, and that means you should frequently learn new things that change your mind.&lt;/p&gt;
&lt;p&gt;Because of all this, I deliberately choose to write blog posts about things I don’t yet quite understand but would like to, like &lt;a href=&quot;/steering-vectors/&quot;&gt;LLM&lt;/a&gt; steering, Stripe’s &lt;a href=&quot;/tempo-faq/&quot;&gt;Tempo&lt;/a&gt; blockchain, &lt;a href=&quot;/c2pa-only-works-if-everything-is-signed/&quot;&gt;C2PA&lt;/a&gt; and &lt;a href=&quot;/text-ai-watermarks/&quot;&gt;watermarking&lt;/a&gt;, space &lt;a href=&quot;/space-ai-datacenters-do-not-have-a-cooling-problem/&quot;&gt;cooling&lt;/a&gt;, &lt;a href=&quot;/interaction-models/&quot;&gt;interaction models&lt;/a&gt;, LLM inference &lt;a href=&quot;/fast-llm-inference/&quot;&gt;internals&lt;/a&gt;, and so on. This is great for me, because I learn a lot. Is it great for my readers?&lt;/p&gt;
&lt;h3&gt;Is blogging to learn irresponsible?&lt;/h3&gt;
&lt;p&gt;I sometimes worry that I should only be writing about areas I already know very well, like &lt;a href=&quot;/how-to-ship/&quot;&gt;tech company dynamics&lt;/a&gt; or &lt;a href=&quot;/good-api-design/&quot;&gt;working&lt;/a&gt; in &lt;a href=&quot;/large-established-codebases/&quot;&gt;large codebases&lt;/a&gt;, rather than presenting myself as an authority on fields I’m actually still learning. Should I let historians of the Luddites write about Luddism, Web3 engineers write about blockchains, and so on? I think this is acceptable for three reasons.&lt;/p&gt;
&lt;p&gt;First, it’s sometimes easier for a beginner to write an introduction to a field than for an expert. Experts routinely &lt;a href=&quot;https://xkcd.com/2501/&quot;&gt;overestimate&lt;/a&gt; the knowledge of the general public, and have often internalized the reasons why their field is important so deeply that they struggle to express them. I think my &lt;a href=&quot;/tags/explainers/&quot;&gt;explainer posts&lt;/a&gt; are valuable because I always spend the first chunk of the post talking about &lt;em&gt;what the original problem is&lt;/em&gt; before I get into the technical solution.&lt;/p&gt;
&lt;p&gt;Second, sometimes the public consensus on a topic is just plain wrong, to the point where even a little bit of research is enough to demonstrate why. Many of my posts I’m proudest of have been along these lines: arguing that the “500ml per prompt” water usage figure for LLMs was &lt;a href=&quot;/water-impact-of-ai/&quot;&gt;ludicrous&lt;/a&gt;, or that the popular Apple “Illusion of Thinking” paper was tracking &lt;a href=&quot;/illusion-of-thinking/&quot;&gt;persistence, not reasoning&lt;/a&gt;, that GPUs &lt;a href=&quot;/ai-gpus-live-longer-than-three-years/&quot;&gt;live longer than three years&lt;/a&gt; and the AI companies have large &lt;a href=&quot;/ai-inference-is-obviously-profitable/&quot;&gt;profit margins&lt;/a&gt; on inference, and so on.&lt;/p&gt;
&lt;p&gt;Third, I try to make it clear on my blog who I am and what my credentials actually are. Even if it’s not explicitly described in the post, I have my real name and resume available on my &lt;a href=&quot;/about&quot;&gt;/about&lt;/a&gt; page, so I don’t think a careful reader could be easily fooled into thinking I’m an expert on 19th-century England or space physics or LLM economics or anything like that.&lt;/p&gt;
&lt;h3&gt;Feedback&lt;/h3&gt;
&lt;p&gt;Even if nobody reads what you write, writing is still a good discipline for getting your thoughts in order. But another big reason why writing is a great learning tool is that &lt;strong&gt;you can get feedback&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;I think it’s obvious why this is useful, but I do want to make two points about feedback. First, if you do make your posts public, you need to have a pretty thick skin. People on the internet often fall over themselves to come up with the most cutting criticism or the harshest dunk. This goes double if you take my previous advice and try to write posts that make a clear, controversial point about a subject you’re learning. If you’re the kind of person whose whole day is ruined when a stranger is cruel to them, you might want to keep your blogging private or only share it among friends.&lt;/p&gt;
&lt;p&gt;Second, even if your blogging is private, &lt;strong&gt;you can get feedback from LLMs&lt;/strong&gt;. Like humans, LLMs will often give junk feedback. In my experience, OpenAI models will always tell me to moderate my claims or add caveats and hedges until I’m not saying anything at all. Sometimes their criticism will be straight-up wrong. But — particularly about technical topics — LLMs are great at pointing out areas you’ve genuinely misunderstood, and they’re far kinder than the average Lobsters or Hacker News commenter.&lt;/p&gt;
&lt;h3&gt;Conclusion&lt;/h3&gt;
&lt;p&gt;I’m pleased and grateful that people enjoy reading my posts, but even when nobody did, I still got a lot of value out of blogging. I write as a method of thinking more clearly, as an excuse to do research on topics I want to learn about, and as a way of getting feedback. &lt;/p&gt;
&lt;p&gt;If you’d like to try it yourself, I suggest watching for these two things. First, you should be changing your mind a lot as you write. If not, you probably aren’t doing enough research. Second, your first draft’s conclusion should be much tighter and more expressive than its introduction. If not, you probably haven’t learned anything from the writing process, which means the draft can be scrapped.&lt;/p&gt;
&lt;p&gt;I strongly recommend this practice to anyone with an interest in writing. You will see the benefits even if you don’t publish any of your writing on the internet, particularly now that you can get good technical feedback by pasting your post into a LLM&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;For what it’s worth, I’ve fiddled with careful “review prompts” and it’s basically as good to just write “review, please:” and paste your article.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[C2PA only works if everything is signed]]></title><link>https://seangoedecke.com/c2pa-only-works-if-everything-is-signed/</link><guid isPermaLink="false">https://seangoedecke.com/c2pa-only-works-if-everything-is-signed/</guid><pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;The &lt;a href=&quot;https://artificialintelligenceact.eu/&quot;&gt;European Union AI Act&lt;/a&gt; is Europe’s attempt to comprehensively regulate AI usage. A big part of that is the requirement that AI-generated content be identifiable: either tagged with a watermark or with what the &lt;a href=&quot;https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content&quot;&gt;Act&lt;/a&gt; calls “digitally signed metadata”&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;. Since all this becomes enforceable in a month, it’s worth figuring out if it makes any sense. I recently discussed AI watermarking at length in &lt;a href=&quot;https://www.seangoedecke.com/text-ai-watermarks/&quot;&gt;&lt;em&gt;Text AI watermarks will always be trivial to remove&lt;/em&gt;&lt;/a&gt;. What about digitally signed metadata?&lt;/p&gt;
&lt;p&gt;The most well-known implementation of digitally signed metadata is C2PA Content Credentials, which is a mechanism&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt; for ensuring that &lt;strong&gt;almost every single image file should contain unspoofable authorship metadata&lt;/strong&gt;. Here’s my position on it:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;C2PA broadly makes sense and is a good idea&lt;/li&gt;
&lt;li&gt;It is pointless to use C2PA for AI-generated images only&lt;/li&gt;
&lt;li&gt;It will take many years for C2PA to be adopted across all images&lt;/li&gt;
&lt;li&gt;Because C2PA makes such great safety theater, we’re going to see a lot of hue and cry about it long before it becomes useful&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Lots to unpack. Let’s start by considering images, since that’s the easiest case.&lt;/p&gt;
&lt;h3&gt;How C2PA signing works&lt;/h3&gt;
&lt;p&gt;When an AI tool generates an image, that tool should include a “made by ChatGPT” disclaimer in that image’s metadata. Likewise, when a camera takes a photo, that camera should include a “taken by a camera” disclaimer. C2PA uses two strategies to protect this metadata:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;The metadata must be &lt;em&gt;signed&lt;/em&gt; by some trusted private key&lt;/li&gt;
&lt;li&gt;The metadata contains a hash of the file’s contents, so you can’t copy an existing signature onto a new file&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Each physical camera (or phone) has its own private key, for obvious reasons&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;. How do we know that those millions of private keys are trusted? Via &lt;a href=&quot;https://en.wikipedia.org/wiki/Public_key_infrastructure&quot;&gt;PKI&lt;/a&gt;, like HTTPS: each camera’s private “certificate” (which contains its public key) is signed by the manufacturer’s well-known private key, so the chain of authenticity can be verified as long as you have (say) Apple’s root &lt;a href=&quot;https://www.apple.com/certificateauthority/&quot;&gt;public key&lt;/a&gt;&lt;sup id=&quot;fnref-4&quot;&gt;&lt;a href=&quot;#fn-4&quot; class=&quot;footnote-ref&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;What happens if you then edit your photo in Photoshop? Photoshop will leave the camera’s metadata untouched, but will layer a “also, Photoshop was used” piece of metadata over the top, signed with Adobe’s private key (well, with the private key associated with your official copy of Photoshop, which is signed by Adobe’s official private key).&lt;/p&gt;
&lt;p&gt;Likewise, if you ask ChatGPT to generate an image for you, ChatGPT will sign its “made by ChatGPT” metadata with OpenAI’s private key. In theory, every single image could contain unforgeable C2PA metadata, allowing software like Twitter to trivially distinguish real photos from fake ones.&lt;/p&gt;
&lt;h3&gt;C2PA needs more regulation to boost adoption&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Right now, C2PA does not have anything like the adoption it’d need to work.&lt;/strong&gt; It’s hard to find hard data on how many images in the wild use C2PA, but FotoForensics &lt;a href=&quot;https://www.hackerfactor.com/blog/index.php?%2Farchives%2F1010-C2PAs-Butterfly-Effect.html&quot;&gt;reports&lt;/a&gt; around a dozen per week (so around 600 out of the &lt;a href=&quot;https://hackerfactor.com/blog/index.php?%2Farchives%2F1088-Fourteen-and-Video.html&quot;&gt;900,000&lt;/a&gt; images processed each year). This is even worse than it sounds, because basically all of the signed images are AI-generated. The adoption rate of C2PA for human-generated images is much, much lower: so far, Google’s Pixel 10 is the only phone camera to sign photos by default. The iPhone &lt;a href=&quot;https://c2pa.ai/smartphone-guide&quot;&gt;doesn’t sign&lt;/a&gt; photos. &lt;/p&gt;
&lt;p&gt;If almost all AI images are C2PA-signed, but almost no human-generated images are, consumers have no reliable way of identifying AI content, because anyone who wants to pretend their AI content is human can simply remove the signature. For C2PA to succeed, it needs to be on every camera and every phone, so that a photo with no signature is rare and suspicious.&lt;/p&gt;
&lt;p&gt;Is that realistic? Actually, I think it is. The appetite (at least in the EU) to regulate AI will increase over time, and while the current EU AI Act only mandates that AI-images are tagged (which by itself is useless), it’s plausible that some future regulation will enforce tagging of all images.&lt;/p&gt;
&lt;p&gt;Another adoption problem that must be solved for C2PA to work is &lt;strong&gt;preservation&lt;/strong&gt;. Right now, if you download a C2PA-tagged image, send it as a Facebook message, then re-download it, the C2PA manifest is stripped out. Most images we see on the internet have passed through some social media asset server at least once. All of these social media companies would need to update how they re-encode image content in order to preserve the C2PA data&lt;sup id=&quot;fnref-5&quot;&gt;&lt;a href=&quot;#fn-5&quot; class=&quot;footnote-ref&quot;&gt;5&lt;/a&gt;&lt;/sup&gt;. This would almost certainly require more regulation: C2PA adds tens or &lt;a href=&quot;https://www.tbray.org/ongoing/When/202x/2024/10/29/Lane-Provenance&quot;&gt;hundreds&lt;/a&gt; of kilobytes to each file, which at social media scale is big money&lt;sup id=&quot;fnref-6&quot;&gt;&lt;a href=&quot;#fn-6&quot; class=&quot;footnote-ref&quot;&gt;6&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;h3&gt;Forging C2PA signatures&lt;/h3&gt;
&lt;p&gt;Could a clever attacker forge a C2PA signature? Kind of. Neal Krawetz, who seems to have led the anti-C2PA charge, &lt;a href=&quot;https://www.hackerfactor.com/blog/index.php?/archives/919-Closed-Standards.html&quot;&gt;points out&lt;/a&gt; that with a camera development kit it’s straightforward to trick a digital camera into thinking that it’s taking an image when in fact it’s being fed one. This is very much not my area, so please write in if you know more about camera hardware and you think I got this wrong. I suppose you could also take a photo of an AI image on a screen, though I imagine you’d have to be careful to make it look real.&lt;/p&gt;
&lt;p&gt;If you exclude physical attacks on a digital camera, I think C2PA is more robust. You can &lt;a href=&quot;https://www.hackerfactor.com/blog/index.php?%2Farchives%2F1010-C2PAs-Butterfly-Effect.html&quot;&gt;sign&lt;/a&gt; a photo with a self-signed certificate, but the C2PA &lt;a href=&quot;https://spec.c2pa.org/specifications/specifications/2.4/specs/C2PA_Specification.html#_trust_lists&quot;&gt;spec&lt;/a&gt; and &lt;a href=&quot;https://opensource.contentauthenticity.org/docs/conformance/trust-lists&quot;&gt;docs&lt;/a&gt; say that validators must check that your certificate bubbles up to the official C2PA &lt;a href=&quot;https://spec.c2pa.org/conformance-explorer/&quot;&gt;trust list&lt;/a&gt;. This list currently contains only 26 certificates, and there’s a whole process for being added to it. That’ll slow down adoption, but at least it makes it hard to forge&lt;sup id=&quot;fnref-7&quot;&gt;&lt;a href=&quot;#fn-7&quot; class=&quot;footnote-ref&quot;&gt;7&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;h3&gt;Other file types and concerns&lt;/h3&gt;
&lt;p&gt;We’ve been talking exclusively about images, but it’s more or less the same story for any type of content. If the file doesn’t support JUMBF metadata (say, an Excel file or a PDF), then the C2PA metadata has to live in a “sidecar”: a separate &lt;code class=&quot;language-text&quot;&gt;.c2pa&lt;/code&gt; file, probably on some Microsoft or Adobe content server, which contains the signed checksum and the data about who created the file.&lt;/p&gt;
&lt;p&gt;However, the distinction between “real” and AI-generated content is fuzzier when you’re not talking about images. Here’s a trivial example: if I ask ChatGPT to create an Excel spreadsheet for me, the file will be tagged as AI-generated, but I can simply copy/paste the content into a new Excel doc and save it, which will tag it as human-generated&lt;sup id=&quot;fnref-8&quot;&gt;&lt;a href=&quot;#fn-8&quot; class=&quot;footnote-ref&quot;&gt;8&lt;/a&gt;&lt;/sup&gt;. There’s no software tool that can identify when I’m retyping some AI-generated text (except for perhaps &lt;a href=&quot;/text-ai-watermarks/&quot;&gt;text fingerprinting&lt;/a&gt;, which has its own raft of issues).&lt;/p&gt;
&lt;p&gt;There are also interesting questions around key management. ChatGPT and other AI tools have an easy problem — their users are all online, and so the files can be signed server-side — but how do you sign files created via Photoshop/Excel/Word? If the user doesn’t have internet, do you use some kind of local key? If so, how do you prevent that key being extracted and used to sign AI-generated content? &lt;/p&gt;
&lt;p&gt;Finally, is it a civil liberties problem to automatically fingerprint every photo? Does it make it impossible to be a whistleblower if every photograph can be traced back to your camera? I think this is a complicated question, but in short: I’d expect whistleblowers to already strip EXIF metadata from their images, C2PA metadata is similarly trivial to strip out, and overall I think image attribution is &lt;em&gt;positive&lt;/em&gt; for whistleblowers because it heads off “this was AI-generated” responses.&lt;/p&gt;
&lt;h3&gt;Conclusion&lt;/h3&gt;
&lt;p&gt;C2PA is probably here to stay. But it isn’t useful now, and won’t be useful until two huge programs of technical work are completed:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Every camera manufacturer (including phones) must C2PA-sign all images by default&lt;/li&gt;
&lt;li&gt;Every social media company must retain the C2PA metadata on uploaded images&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This will be a long organizational process, since each manufacturer must go through the approvals process (or decide to start their own competing system), evaluate the legal ramifications of storing attribution data in images, and so on. It will be a long technical process, because C2PA metadata is a substantial fraction of image sizes: storing it will add many petabytes of content.&lt;/p&gt;
&lt;p&gt;Of course, just because C2PA isn’t useful doesn’t mean we’re not all going to do it. Lots of companies are under pressure to signal that they care about AI safety and to head off regulatory attack. “We’re cryptographically signing AI-generated content” is a compelling “we’re doing &lt;em&gt;something&lt;/em&gt;” pitch, particularly for people who aren’t technically savvy enough to understand the limitations. In the near term, I expect large AI-involved companies to invest a substantial amount of engineering effort in C2PA-related activity.&lt;/p&gt;
&lt;p&gt;In the long run, once everyone gets on board, I think C2PA could end up working well. It’s awkward in some ways, but “attest content via a PKI certificate chain” is a good idea.&lt;/p&gt;
&lt;p&gt;Is it possible to defeat? Yes, of course. By design, private keys will be in the user’s hands — in their cameras, in their local versions of Photoshop or Microsoft Word, in their phones — so sufficiently technical users will be able to crack them out or use them to sign whatever content they want. I still think C2PA will end up stemming the tide of AI content, because most users are not going to be sophisticated enough to perform attacks like this. However, we should still retain some skepticism of unlikely-looking content, even if it has “created by a human” in its C2PA metadata.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;See sub-measure 1.1.1 of the Act’s associated &lt;a href=&quot;https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content&quot;&gt;Code of Practice&lt;/a&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;A previous version of this post criticized C2PA for &lt;a href=&quot;https://explorer.artificialintelligenceact.eu/en/&quot;&gt;incorrectly&lt;/a&gt; claiming to be the semi-official technology of the AI Act, but in fact this claim comes from &lt;a href=&quot;https://c2paviewer.com/articles/eu-ai-act-content-credentials&quot;&gt;C2PA Viewer&lt;/a&gt;, which is not affiliated with the official C2PA coalition. Thanks to &lt;a href=&quot;https://paul-friedl.github.io/&quot;&gt;Paul Friedel&lt;/a&gt;, who has recently written &lt;a href=&quot;https://verfassungsblog.de/the-problems-with-general-purpose-ai-detectability/&quot;&gt;his own post&lt;/a&gt; about C2PA, for emailing me with the correction.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;Otherwise if you cracked the key out of one Sony camera, you could spoof content from any Sony camera.&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-4&quot;&gt;
&lt;p&gt;In practice there are usually more “links in the chain”: a device will be signed by some intermediate certificate, which in turn will be signed by another intermediate certificate, which will be signed by the root certificate. That’s because the root key is so valuable. If an intermediate private key leaks, it can be revoked and replaced (via the root key), but if the root key leaks, it would take &lt;em&gt;years&lt;/em&gt; to rebuild the network of trust. So almost all signing is done by intermediates, and the root key stays on a USB drive locked in a safe somewhere.&lt;/p&gt;
&lt;a href=&quot;#fnref-4&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-5&quot;&gt;
&lt;p&gt;Not to mention that the whole &lt;em&gt;point&lt;/em&gt; of C2PA is that these social media companies will be displaying a “human or AI” sticker in their UI, which will require retaining the metadata.&lt;/p&gt;
&lt;a href=&quot;#fnref-5&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-6&quot;&gt;
&lt;p&gt;C2PA allows for storing the manifest &lt;em&gt;content&lt;/em&gt; as a separate &lt;code class=&quot;language-text&quot;&gt;.c2pa&lt;/code&gt; file, and just including a manifest &lt;em&gt;url&lt;/em&gt; in the image metadata itself, but that doesn’t solve the cloud provider problem: they still have to store all the &lt;code class=&quot;language-text&quot;&gt;.c2pa&lt;/code&gt; files on-disk somewhere.&lt;/p&gt;
&lt;a href=&quot;#fnref-6&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-7&quot;&gt;
&lt;p&gt;I think this defuses Neal Krawetz’s &lt;a href=&quot;https://www.hackerfactor.com/blog/index.php?/archives/1013-C2PAs-Worst-Case-Scenario.html&quot;&gt;“worst-case scenario”&lt;/a&gt;. I downloaded his forged image, and (as expected) it gets flagged as “signed, but we don’t trust the root”. I think Krawetz was right at the time, though, since the official “trust list” was only launched in mid-2025.&lt;/p&gt;
&lt;a href=&quot;#fnref-7&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-8&quot;&gt;
&lt;p&gt;You &lt;em&gt;could&lt;/em&gt; do the same thing with images by copying into Photoshop or Paint, but while that’d obscure the AI source, it would still be clear that the photo wasn’t taken by a camera.&lt;/p&gt;
&lt;a href=&quot;#fnref-8&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[Text AI watermarks will always be trivial to remove]]></title><link>https://seangoedecke.com/text-ai-watermarks/</link><guid isPermaLink="false">https://seangoedecke.com/text-ai-watermarks/</guid><pubDate>Thu, 02 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;The European Union &lt;a href=&quot;https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai&quot;&gt;AI Act&lt;/a&gt; will begin to be enforceable in August 2026, one month from now&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;. One of the biggest new requirements is &lt;a href=&quot;https://artificialintelligenceact.eu/article/50/&quot;&gt;Article 50&lt;/a&gt;, which requires all AI outputs to be “detectable as artificially generated”. In other words, if LLM providers want to do business in the EU, they will have to apply a watermark to their outputs&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;: some hidden signature that can be used to identify AI content.&lt;/p&gt;
&lt;p&gt;LLM text watermarking is a fascinating problem. Like the best engineering problems, it is theoretically hard to solve perfectly, but has multiple partial solutions: for instance, Google’s &lt;a href=&quot;https://deepmind.google/models/synthid/&quot;&gt;SynthID&lt;/a&gt;, and (as I’ll argue) some quiet Unicode trickery from OpenAI and Anthropic. It will be interesting to see how the AI labs navigate these tradeoffs before the end of the year.&lt;/p&gt;
&lt;h3&gt;Why text watermarking is hard&lt;/h3&gt;
&lt;p&gt;I wrote about AI watermarking at the end of last year in &lt;a href=&quot;/ai-detection/&quot;&gt;&lt;em&gt;AI detection tools cannot prove that text is AI-generated&lt;/em&gt;&lt;/a&gt;. It’s easy to watermark an image, because digital images contain lots of noise that the human eye can’t really see. For instance, you could apply a watermark like “these twenty pixels in these exact spots will always share a color”. Text is much, much harder. Unlike images, text is a very compressed medium: you cannot make any change to a sentence that a human wouldn’t notice (with one exception, which we’ll get to later). So how are you supposed to watermark it?&lt;/p&gt;
&lt;p&gt;It’s basically a &lt;a href=&quot;https://arxiv.org/pdf/1302.2718&quot;&gt;text steganography&lt;/a&gt; problem (concealing a secret code), made more difficult because the plaintext cannot be arbitrarily manipulated. Any changes you make to apply the watermark will compromise the quality of the output. For instance, “every fifth letter is an ‘e’” would be a good watermark, but applied naively would make the AI output full of typos. Could you just let the model figure out how to fit the watermark? Strong AI models are smart enough to juggle this kind of constraint&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;, but it’d still consume reasoning time that would be better spent on the user’s problem, and make the model sound much less capable than it is&lt;sup id=&quot;fnref-4&quot;&gt;&lt;a href=&quot;#fn-4&quot; class=&quot;footnote-ref&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;h3&gt;Do we need watermarks to detect AI content?&lt;/h3&gt;
&lt;p&gt;Do you really need a watermark? If you’re Anthropic, and you’re required to be able to verify whether your models produced a particular block of text, can’t you simply run the text through each model, measuring as you go how closely the model’s predicted tokens match each token from the text? &lt;/p&gt;
&lt;p&gt;Not really. The space of “all possible Claude Sonnet answers to a question” is way larger than the space of “all possible &lt;em&gt;watermarked&lt;/em&gt; answers to a question”. In other words, you’d get too many false positives for human text that reads like it was AI-written. It’s way more likely for a human to accidentally write like Claude than it is for a human to accidentally reproduce a watermark.&lt;/p&gt;
&lt;p&gt;It would also be prohibitively expensive to run every Anthropic model against a piece of text in order to watermark it. The EU AI Act will eventually require labs like Anthropic to offer free watermarking services to every EU citizen (see Commitment 2). You couldn’t do that with the “run the model” approach.&lt;/p&gt;
&lt;h3&gt;How SynthID works&lt;/h3&gt;
&lt;p&gt;As far as I know, the only AI provider to say they watermark text output is Google, who use a tool called &lt;a href=&quot;https://www.nature.com/articles/s41586-024-08025-4&quot;&gt;SynthID&lt;/a&gt;. Here’s how it works.&lt;/p&gt;
&lt;p&gt;When a LLM generates text, it’s generating a series of tokens (words or chunks of words). At each step, the model itself doesn’t output a single token, but instead outputs a full list of all (say) 100,000 tokens in its vocabulary, each annotated with the probability that that token will be the next one. Tools like ChatGPT or Claude Code will pick semi-randomly from the most likely options in order to get their outputs. &lt;strong&gt;This semi-random sampling process can be influenced in a detectable way.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;For instance, we could choose a sampling strategy like “we pick the second most likely token, then the first, then the second, then the first, and so on”. That would still produce high-quality output, but you’d be able to re-run the model against the generated text to verify that the pattern holds. However, that’d make verification really expensive, and any slight tweaks to the output would break the pattern and thus break the fingerprint. Is there a better way?&lt;/p&gt;
&lt;p&gt;Yes. SynthID is a process for assigning each token a “score” based on its previous tokens (for instance, sum the token’s ID with the IDs of its previous three tokens then take mod 5)&lt;sup id=&quot;fnref-5&quot;&gt;&lt;a href=&quot;#fn-5&quot; class=&quot;footnote-ref&quot;&gt;5&lt;/a&gt;&lt;/sup&gt;. To apply the watermark, the model adopts a sampling strategy like “out of the top five most likely tokens, pick the one with the top SynthID score”&lt;sup id=&quot;fnref-6&quot;&gt;&lt;a href=&quot;#fn-6&quot; class=&quot;footnote-ref&quot;&gt;6&lt;/a&gt;&lt;/sup&gt;. The watermark can then be detected by calculating the aggregate SynthID score of a block of text. If it’s suspiciously high, it’s very likely to have been AI-generated.&lt;/p&gt;
&lt;p&gt;This is basically a version of the common advice that you can identify LLMs by use of the &lt;a href=&quot;/em-dashes/&quot;&gt;em-dash&lt;/a&gt;, except that instead of a list of keywords, it relies on subtle mathematical relationships between words that humans can’t identify. Because the process for assigning the score is trivial, it’s very cheap to run watermark detection.&lt;/p&gt;
&lt;h3&gt;Unicode watermarks via homoglyphs&lt;/h3&gt;
&lt;p&gt;Google have a complicated mathematical rationale for why SynthID doesn’t make the model dumber: supposedly the SynthID scoring is random enough to act like a normal pseudo-random token sampler, just one that leaves a detectable fingerprint on the outputs. But of course this is suspicious. For instance, it’s common to do inference setting temperature to zero, which always picks the model’s most likely next token. In that case, you can’t leave a fingerprint at all (or you have to ignore the user’s preference and pick the second or third choice anyway).&lt;/p&gt;
&lt;p&gt;If you can’t alter the model outputs, can you still fingerprint the content? Well, kind of. I’m pretty sure OpenAI and Anthropic are sometimes applying fancy Unicode tricks. For instance, you might go through and replace your normal ” ” spaces (unicode &lt;code class=&quot;language-text&quot;&gt;U+0020&lt;/code&gt;) with a three-per-em ” ” space (unicode &lt;code class=&quot;language-text&quot;&gt;U+2004&lt;/code&gt;), or a CJK ideographic ”　” space (unicode &lt;code class=&quot;language-text&quot;&gt;U+3000&lt;/code&gt;). These are called “homoglyphs”, and you can find more of them &lt;a href=&quot;https://www.irongeek.com/homoglyph-attack-generator.php&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Of course, lots of human-generated text uses homoglyphs. But it’s trivial to encode a &lt;em&gt;pattern&lt;/em&gt; of homoglyphs (say, “every third space becomes a three-per-em”) that is much less likely to occur in the wild. Like the SynthID watermark, a homoglyph-based watermark can be detected very cheaply. A homoglyph-based watermark is cheaper to apply than SynthID: you could even do it entirely on the client.&lt;/p&gt;
&lt;p&gt;I don’t think this is a conspiracy theory. Claude Code was &lt;a href=&quot;https://thereallo.dev/blog/claude-code-prompt-steganography&quot;&gt;definitely doing this&lt;/a&gt; to tag suspicious requests from Chinese users (exploiting homoglyphs for the ’ character in “Today’s date”, though they’ve since walked that back). In the last few years, I’ve noticed that when I copy blocks of text from ChatGPT and paste them into VSCode, sometimes VSCode marks some or all of the spaces as unusual Unicode characters&lt;sup id=&quot;fnref-7&quot;&gt;&lt;a href=&quot;#fn-7&quot; class=&quot;footnote-ref&quot;&gt;7&lt;/a&gt;&lt;/sup&gt;. Are OpenAI and Anthropic using homoglyphs as an AI-generated watermark? I’m not sure. But they’re definitely using homoglyphs.&lt;/p&gt;
&lt;h3&gt;Text watermarks can be trivially removed&lt;/h3&gt;
&lt;p&gt;The AI Act (specifically, its associated &lt;a href=&quot;https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content&quot;&gt;Code of Practice&lt;/a&gt;) requires watermarking to be “embedded within the content in a manner that is difficult for it to be separated from the content”. However, text watermarks can be trivially removed.&lt;/p&gt;
&lt;p&gt;To remove unicode homoglyph watermarking, you simply have to replace all the homoglyphs with their “real” character equivalents. If you have access to even a relatively weak un-watermarked LLM&lt;sup id=&quot;fnref-8&quot;&gt;&lt;a href=&quot;#fn-8&quot; class=&quot;footnote-ref&quot;&gt;8&lt;/a&gt;&lt;/sup&gt;, you can strip out SynthID watermarking by asking that LLM to paraphrase the text content. Because the watermark is inherent to subtle vocabulary choices, re-wording the content will remove the watermark. You could even do it by hand, although at that point it’s not really AI-generated content anymore. Since there will be some kind of free public watermark testing tool, you can just keep tweaking until it comes back negative.&lt;/p&gt;
&lt;p&gt;Moreover, the AI Act requires watermarking techniques to be “interoperable… as far as this is technically feasible”. That means AI providers would have to publish their watermarking process, and potentially even attempt to standardize on applying the same kind of watermarks. I just don’t see how this is compatible with the kind of security-by-obscurity that LLM text watermarking depends on. Unlike image and video watermarks, text watermarks will always be trivial to remove.&lt;/p&gt;
&lt;h3&gt;What about C2PA?&lt;/h3&gt;
&lt;p&gt;The AI Act and Code of Practice talk a lot about “digitally signed metadata”. The idea here is that you can include an AI disclosure in the file’s metadata itself, ideally in a way that cannot be tampered with (for instance, by signing a hash of the file’s contents). This signed-metadata process is basically &lt;a href=&quot;https://c2paviewer.com/articles/eu-ai-act-content-credentials&quot;&gt;C2PA Content Credentials&lt;/a&gt;. While you can remove C2PA metadata, you (theoretically) can’t &lt;em&gt;fake&lt;/em&gt; it, so a file with “created by a human” metadata can be trusted, and files with no metadata at all can be held in suspicion.&lt;/p&gt;
&lt;p&gt;This post is already too long to get into what I think about C2PA, but I do want to say that &lt;strong&gt;C2PA is not a substitute for text watermarking&lt;/strong&gt;. It only really applies to &lt;em&gt;files&lt;/em&gt;. In the words of the Code of Practice, that’s “a data format that supports attaching metadata (e.g., an audio, image, video, or containerised text)“. The output of chat tools (and most of the output of AI agents) is not containerized text, but plain old regular text, and so can’t be signed. What would it even look like to sign ChatGPT outputs? There’s no artifact to pass around.&lt;/p&gt;
&lt;p&gt;I think it’s a fascinating question whether Claude Code has to C2PA-sign any HTML files or PDFs it generates for you. That seems kind of tricky to get right. But in any case, the AI Act also mandates some kind of actual watermarking as well.&lt;/p&gt;
&lt;h3&gt;Conclusion&lt;/h3&gt;
&lt;p&gt;So what’s going to happen this year? If I had to guess, I’d say that each AI provider (not just labs like OpenAI or Anthropic, but third-party providers like Fireworks or Groq) will stick a SynthID token sampler in front of their inference stacks. This might be limited to users in the EU, but it might not be, since SynthID is at least as good as a normal top-k token sampling approach.&lt;/p&gt;
&lt;p&gt;AI providers will then offer a “check for watermark” page that re-tokenizes user-provided text, runs the scoring, and checks whether it’s above a certain threshold. Depending on how seriously the interoperability clause is taken, providers might even standardize on the same SynthID setup, in which case there could be a single EU-hosted “watermark this text” page.&lt;/p&gt;
&lt;p&gt;I don’t think unicode-based watermarking is going to be considered compliant with the AI Act, but some providers which don’t want to set up SynthID might try it. Either way, technical users will be able to strip out the watermark at will, and there will be a plethora of tools that non-technical users will use for this purpose.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;Well, for new systems; existing ones get until December.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;I don’t think the plain text of Article 50 requires this, but &lt;a href=&quot;https://artificialintelligenceact.eu/recital/133/&quot;&gt;Recital 133&lt;/a&gt; and the Code of Practice makes it pretty clear that they’re looking for watermarks.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;Even with extra high thinking, GPT-5.5 could not explain SynthID to me with every fifth letter being an “e”, but GPT-5.5-Pro produced this puzzling koan: “These hidden codes label model-made image, voice, movie, prose. Probe trace: maybe a model-made piece. Maybe erase trace; maybe leave trace. Hence trace alone? No.”&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-4&quot;&gt;
&lt;p&gt;I leave the analogy with AI safety guardrails as an exercise for the reader.&lt;/p&gt;
&lt;a href=&quot;#fnref-4&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-5&quot;&gt;
&lt;p&gt;That’s a toy example. In practice there are multiple different (but still mathematically simple) scoring methods that get combined together, including a random seed. Why include the seed? Otherwise the watermark would bias towards the same set of tokens.&lt;/p&gt;
&lt;a href=&quot;#fnref-5&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-6&quot;&gt;
&lt;p&gt;The tokens are scored in a multi-round knockout against each other, but I think that’s more of an implementation detail and not required to get the core intuition behind why SynthID works.&lt;/p&gt;
&lt;a href=&quot;#fnref-6&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-7&quot;&gt;
&lt;p&gt;When this became &lt;a href=&quot;https://www.rumidocs.com/newsroom/new-chatgpt-models-seem-to-leave-watermarks-on-text&quot;&gt;public knowledge&lt;/a&gt;, OpenAI claimed it was just a model quirk, which is certainly possible.&lt;/p&gt;
&lt;a href=&quot;#fnref-7&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-8&quot;&gt;
&lt;p&gt;All AI providers might be legally required to watermark, but even tiny local models are good enough to paraphrase text.&lt;/p&gt;
&lt;a href=&quot;#fnref-8&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[Saying the obvious thing]]></title><link>https://seangoedecke.com/saying-the-obvious-thing/</link><guid isPermaLink="false">https://seangoedecke.com/saying-the-obvious-thing/</guid><pubDate>Sat, 27 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Stating the obvious is &lt;a href=&quot;https://blog.jim-nielsen.com/2026/blogging-stating-the-obvious/&quot;&gt;surprisingly useful&lt;/a&gt;. Most of your knowledge lives below the threshold of conscious awareness, so it’s possible for a piece of writing to remind you of what you already know. It’s common to know you don’t like something without being quite sure why, and reading an obvious statement (such as “accuracy &lt;a href=&quot;https://www.astralcodexten.com/p/if-its-worth-your-time-to-lie-its&quot;&gt;matters&lt;/a&gt;, even when you agree with the broad strokes”) can help clarify why you find certain things distasteful.&lt;/p&gt;
&lt;p&gt;Sometimes you can see some obvious truth that nobody seems to be talking about, and reading it in someone else’s words can prompt an “oh god, I’m not crazy” moment of catharsis. For many junior engineers, it’s almost a rite of passage to notice that some percentage of software engineers &lt;a href=&quot;https://x.com/yegordb/status/1859290734257635439&quot;&gt;do virtually no work&lt;/a&gt;. Since nobody talks about it (how would you even bring it up in the workplace?), they often feel like they’re losing their minds: &lt;em&gt;surely&lt;/em&gt; this state of affairs wouldn’t be allowed to continue, so they must be completely misreading the situation. But in fact it’s true.&lt;/p&gt;
&lt;p&gt;Stating the obvious is hard. It can even be dangerous: sometimes there’s a good reason nobody says the obvious thing. But I think the bigger reason it’s hard is for the same reason that it’s hard to &lt;a href=&quot;https://drawingacademy.com/drawing-what-you-see-vs-drawing-what-you-know&quot;&gt;draw what you actually see&lt;/a&gt;. When I look at a person and try to draw them, I’m not drawing the lines and shades my eye sees (like a printer or camera might). I’m drawing &lt;em&gt;what I know the person looks like&lt;/em&gt;, which is a kind of stick-figure approximation. It takes time and effort to drop the layer of interpretation and draw what’s actually there&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;Many of the posts I’m most proud of are times when I’ve managed to articulate something I think is obviously true: &lt;a href=&quot;/ratchet-effects/&quot;&gt;engineer reputation is determined by ratchet effects&lt;/a&gt;, &lt;a href=&quot;/being-right-a-lot/&quot;&gt;good engineers are right most of the time&lt;/a&gt;, &lt;a href=&quot;/party-tricks/&quot;&gt;you shouldn’t just do JIRA tickets&lt;/a&gt; (or &lt;a href=&quot;/glue-work-considered-harmful/&quot;&gt;glue work&lt;/a&gt;), and so on. These are all things I’ve believed for a while, but have only (relatively) recently been able to &lt;em&gt;notice&lt;/em&gt; that I believe them. Sometimes I’m helped along by reading something I vehemently disagree with (like “nobody gets promoted for doing &lt;a href=&quot;/simple-work-gets-rewarded/&quot;&gt;simple work&lt;/a&gt;”, or ”&lt;a href=&quot;/big-tech-needs-big-egos/&quot;&gt;big egos&lt;/a&gt; have no place in tech”).&lt;/p&gt;
&lt;p&gt;Stating the obvious doesn’t mean avoiding nuance. Every obvious claim carries with it a host of subtle, non-obvious claims. For example, I believe that having a big ego can be very useful as a software engineer. But why exactly is that, and what do I mean by ego? Obviously it’s not good to be constantly flexing your status on other people, or to be unable to tolerate the possibility of being wrong. However, I do think you need to be able to take &lt;a href=&quot;/taking-a-position/&quot;&gt;firm technical positions&lt;/a&gt; even when the situation is uncertain, which means you have to be confident in your technical instincts. Teasing out that distinction (and its implications) is very interesting, but in order to do it you need to be able to first articulate the obvious part.&lt;/p&gt;
&lt;p&gt;I’ve been talking about stating the obvious in technical blogging. But this principle applies just as well to other kinds of communication. When I write a technical design document at work, it’s very important to state the obvious. In fact, technical communication is &lt;a href=&quot;/technical-communication/&quot;&gt;so hard&lt;/a&gt; and general understanding is &lt;a href=&quot;/nobody-knows-how-software-products-work/&quot;&gt;so poor&lt;/a&gt; that just getting people aligned on the obvious things is often &lt;em&gt;enormously&lt;/em&gt; valuable. Much great literature and poetry aims to bring out some obvious but hard-to-articulate part of human experience.&lt;/p&gt;
&lt;p&gt;Don’t avoid writing something down just because you think it’s obvious. The thing you think is obvious now might recede into your subconscious in an hour; get it written down while you can! And don’t avoid writing something down because you think it’s dangerous to say and everyone already knows it. For people new to the area, reading your words can help them feel like they’re not losing their minds. Finally, once you write down the obvious thing, it allows you to go on and draw out the parts that are less obvious, in a way that you couldn’t do if you try to just skip straight to the subtleties.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;Incidentally, this is why most people &lt;a href=&quot;https://www.reddit.com/r/pics/comments/4ew345/as_it_turns_out_most_people_cannot_draw_a_bike/&quot;&gt;cannot draw a bicycle&lt;/a&gt; on their first attempt. Unless you’re a mechanical engineer, you probably do not have a stick-figure-level approximation of what a bicycle looks like in your head, so you begin confidently (after all, you’ve seen a thousand bicycles) and get stuck after the first few lines.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[AI inference is obviously profitable]]></title><link>https://seangoedecke.com/ai-inference-is-obviously-profitable/</link><guid isPermaLink="false">https://seangoedecke.com/ai-inference-is-obviously-profitable/</guid><pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Many people &lt;a href=&quot;https://www.wheresyoured.at/why-everybody-is-losing-money-on-ai/&quot;&gt;claim&lt;/a&gt; that AI inference is unprofitable to serve, and thus must be subsidized by an ocean of dumb money from investors who believe that some future AI model will come to dominate the world economy. When that dumb money goes away, so will AI products. According to this view, LLMs are just inherently too expensive (in terms of money, power, and water) to be used in consumer products. In fact, they can only be used today by externalizing the costs: money onto VC funds and now retail ETF &lt;a href=&quot;https://www.investopedia.com/spacex-stock-joins-major-index-funds-what-regular-investors-need-to-know-spcx-ipo-vanguard-blackrock-vti-itot-12004764&quot;&gt;investors&lt;/a&gt;, power onto electric utility &lt;a href=&quot;https://salatainstitute.harvard.edu/how-you-subsidize-big-tech-with-your-electricity-bill/&quot;&gt;consumers&lt;/a&gt;, and water onto the &lt;a href=&quot;https://theconversation.com/5-ways-data-centers-endanger-their-local-communities-and-the-country-as-a-whole-282348&quot;&gt;communities&lt;/a&gt; where datacenters are built.&lt;/p&gt;
&lt;p&gt;There are &lt;a href=&quot;/is-ai-wrong/&quot;&gt;good reasons&lt;/a&gt; to dislike AI, but this really isn’t one of them. In fact, &lt;strong&gt;AI inference is obviously profitable&lt;/strong&gt;.&lt;/p&gt;
&lt;h3&gt;Doing the math demonstrates that inference is profitable&lt;/h3&gt;
&lt;p&gt;Frontier AI providers are reporting 70%-80% &lt;a href=&quot;https://www.morningstar.com/stocks/anthropics-gross-margin-is-most-important-number-tech&quot;&gt;gross&lt;/a&gt; &lt;a href=&quot;https://www.saastr.com/have-ai-gross-margins-really-turned-the-corner-the-real-math-behind-openais-70-compute-margin-and-why-b2b-startups-are-still-running-on-a-treadmill/&quot;&gt;margins&lt;/a&gt; on inference, but maybe we can’t trust them. Let’s do some very rough estimates on the actual cost.&lt;/p&gt;
&lt;p&gt;A Nvidia A100 consumes 400W of power under full load. In practice, even a carefully-tuned inference server will not be at full load all the time, but it’s at least an upper bound. Suppose you’re running a dense 70B model&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;, which will &lt;a href=&quot;https://dlewis.io/evaluating-llama-33-70b-inference-h100-a100/&quot;&gt;fit&lt;/a&gt; comfortably (unquantized) on four A100s at around 2M tokens per hour. At industrial power prices, that’s about 13c/hr in the &lt;a href=&quot;https://www.eia.gov/electricity/monthly/update/end-use.php&quot;&gt;USA&lt;/a&gt;. Suppose (pessimistically) cooling is the same cost. That’s about 13 cents per million output tokens&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;Let’s amortize the cost of the GPUs, since that’s going to be the most expensive part. An A100 costs about $20k. If each A100 lasts around five years&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;, you’ll have to make 16k/yr in profit to recoup your capital investment (or $1.80 per hour). At lower utilization, it’ll take longer to recoup, but your GPUs will also last longer. Either way, your overall inference costs are at about one dollar per million tokens.&lt;/p&gt;
&lt;p&gt;GPT-5.4-mini &lt;a href=&quot;https://openai.com/business/pricing/#api&quot;&gt;charges&lt;/a&gt; $4.50 per million tokens, and stronger OpenAI or &lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic&lt;/a&gt; models are three to six times as expensive. It’s hard to make a direct comparison because we don’t know the size of OpenAI or Anthropic models, but the claimed 70% or 80% profit margin is extremely plausible.&lt;/p&gt;
&lt;h3&gt;Open LLMs demonstrate that inference is profitable&lt;/h3&gt;
&lt;p&gt;What if you don’t trust my estimates either? Let’s look at the pricing of open-weights Chinese LLMs. DeepSeek have &lt;a href=&quot;https://github.com/deepseek-ai/open-infra-index/blob/main/202502OpenSourceWeek/day_6_one_more_thing_deepseekV3R1_inference_system_overview.md&quot;&gt;claimed&lt;/a&gt; a bit over 80% profit margin on inference for DeepSeek-R1. Since their API pricing for R1 is less than half that of OpenAI or Anthropic&lt;sup id=&quot;fnref-4&quot;&gt;&lt;a href=&quot;#fn-4&quot; class=&quot;footnote-ref&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;, that suggests that my estimates above for inference cost might be too expensive. Cooling at scale is probably &lt;a href=&quot;https://massedcompute.com/faq-answers/?question=What%20are%20the%20estimated%20annual%20power%20consumption%20costs%20of%20NVIDIA%20A100%20and%20H100%20GPUs%20in%20a%20typical%20data%20center?&quot;&gt;cheaper&lt;/a&gt; than power, R1 only has half the active parameters of a dense 70B model, modern GPUs are more efficient than the A100, and there are significant &lt;a href=&quot;/inference-batching-and-deepseek/&quot;&gt;economies of scale&lt;/a&gt; in inference.&lt;/p&gt;
&lt;p&gt;Since DeepSeek’s models are available for anyone to download, they can’t get away with extracting a large profit margin. One of the other inference providers would undercut them with the same model. Inference costs for DeepSeek-V4-Pro on the market are around 87 cents per million output tokens, which is probably pretty close to the actual cost of serving the model.&lt;/p&gt;
&lt;h3&gt;For AI labs, inference must subsidize training&lt;/h3&gt;
&lt;p&gt;All of this doesn’t mean that &lt;em&gt;OpenAI&lt;/em&gt; or &lt;em&gt;Anthropic&lt;/em&gt; are profitable. Those companies are making huge capital &lt;a href=&quot;https://openai.com/index/building-the-compute-infrastructure-for-the-intelligence-age/&quot;&gt;investments&lt;/a&gt; that may or may not pan out, and are spending enormous amounts of money on talent and compute to train brand-new models and retain users.&lt;/p&gt;
&lt;p&gt;They’re doing crazy things like offering per-month subscription models for nearly unlimited inference, which is almost certainly not profitable. If you used an API token instead of your Anthropic subscription in Claude Code, you’d pay ten times the cost. But that doesn’t mean API-based Claude Code couldn’t be a good deal. Some people are &lt;a href=&quot;https://www.reddit.com/r/opencodeCLI/comments/1tril88/test_of_prices_of_deepseek_in_opencode_go_and_api/&quot;&gt;already using&lt;/a&gt; DeepSeek’s inference API for agentic coding, because once you take away the huge profit margin it’s cheaper than the relative per-month subscription.&lt;/p&gt;
&lt;p&gt;Why won’t OpenAI or Anthropic lower their prices? Supposedly OpenAI has &lt;a href=&quot;https://www.wsj.com/tech/ai/openai-considers-drastic-price-cuts-anticipating-war-for-users-with-anthropic-9b8c178e&quot;&gt;thought about it&lt;/a&gt;, but for an AI lab, &lt;strong&gt;inference has to subsidize training costs&lt;/strong&gt;. A company like OpenAI has to fund the production of new models from the inference margins on existing models (at least partially). That’s why the margins on inference are so high: the AI labs are trying to squeeze out every dollar so they can stay alive in the training arms race.&lt;/p&gt;
&lt;p&gt;However, inference only has to subsidize training costs &lt;strong&gt;for an AI lab&lt;/strong&gt;. If you’re merely an inference provider, you don’t have to do any training at all. Therefore, even if OpenAI and Anthropic go out of business, whoever snaps up the rights to their frontier models will be able to continue selling Opus and GPT inference at a profit&lt;sup id=&quot;fnref-5&quot;&gt;&lt;a href=&quot;#fn-5&quot; class=&quot;footnote-ref&quot;&gt;5&lt;/a&gt;&lt;/sup&gt;. The AI bubble popping will not mean the end of the inference business, because &lt;strong&gt;AI inference is obviously profitable&lt;/strong&gt;.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;Expensive frontier models are probably mixture-of-experts, not dense, which is tougher to estimate. However, I think a 70B dense model and a MoE with 70B active params will come out to basically the same numbers at scale (though the MoE will require more GPU memory and thus a greater upfront cost). Are frontier models around 70B params? Nobody outside the AI labs really knows, but my guess is that 70B is probably larger than a Haiku/mini class model.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;I think it’s reasonable to estimate the cost of output tokens only, since they’re by far the most expensive part of serving inference. Input tokens are cheaper for two reasons: transformers let you prefill them in parallel, and for most real-world use cases they can be aggressively cached in the KV cache.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;It’s common (and wrong) to estimate GPU lifespan at three years. I wrote a lot about this in &lt;a href=&quot;/ai-gpus-live-longer-than-three-years/&quot;&gt;&lt;em&gt;AI GPUs probably live longer than three years&lt;/em&gt;&lt;/a&gt;. &lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-4&quot;&gt;
&lt;p&gt;Again, this is just an guess, since we don’t know what OpenAI or Anthropic model is equivalent in size to R1.&lt;/p&gt;
&lt;a href=&quot;#fnref-4&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-5&quot;&gt;
&lt;p&gt;I do wonder if Anthropic would be able to prevent other people from being able to access the model if the company goes out of business. Anthropic is currently in &lt;a href=&quot;https://www.bloomberg.com/news/articles/2026-06-02/broadcom-backing-lowers-debt-costs-on-36-billion-anthropic-deal&quot;&gt;debt&lt;/a&gt; to Broadcom, Google, and a bunch of private equity firms. Would they get the Mythos and Opus weights, over Dario’s protestations? &lt;/p&gt;
&lt;a href=&quot;#fnref-5&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[AI GPUs probably live longer than three years]]></title><link>https://seangoedecke.com/ai-gpus-live-longer-than-three-years/</link><guid isPermaLink="false">https://seangoedecke.com/ai-gpus-live-longer-than-three-years/</guid><pubDate>Mon, 15 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;a href=&quot;https://www.wheresyoured.at/ai-is-slowing-down&quot;&gt;People&lt;/a&gt; who think current AI use is unsustainable often rely on the &lt;a href=&quot;https://www.tomshardware.com/pc-components/gpus/datacenter-gpu-service-life-can-be-surprisingly-short-only-one-to-three-years-is-expected-according-to-unnamed-google-architect&quot;&gt;claim&lt;/a&gt; that inference GPUs only last “three years at the most” under load&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;. The idea here is that once the AI bubble money drains away, current infrastructure will rapidly become obsolete, and there won’t be enough money floating around to buy a whole slate of brand-new GPUs. Inference costs would thus rapidly become way too expensive for current AI products to make any financial sense.&lt;/p&gt;
&lt;p&gt;Where does this “three years at the most” claim come from? Is it plausible? &lt;/p&gt;
&lt;h3&gt;Sourcing the quote&lt;/h3&gt;
&lt;p&gt;The original Tom’s Hardware article quotes this &lt;a href=&quot;https://x.com/techfund1/status/1849031571421983140&quot;&gt;tweet&lt;/a&gt; from Tech Fund, an anonymous former PM and tech investor, who quotes an anonymous “GenAI principal architect” at Google as saying “if you have a high utilization rate, then constant high utilization rate for a year or two, I think the lifespan will be three years at most”.&lt;/p&gt;
&lt;p&gt;&lt;span
      class=&quot;gatsby-resp-image-wrapper&quot;
      style=&quot;position: relative; display: block; margin-left: auto; margin-right: auto; max-width: 590px; &quot;
    &gt;
      &lt;a
    class=&quot;gatsby-resp-image-link&quot;
    href=&quot;/static/9cf90a11e418ecd73f81e5178a9328c2/1b853/tweet.png&quot;
    style=&quot;display: block&quot;
    target=&quot;_blank&quot;
    rel=&quot;noopener&quot;
  &gt;
    &lt;span
    class=&quot;gatsby-resp-image-background-image&quot;
    style=&quot;padding-bottom: 85.8108108108108%; position: relative; bottom: 0; left: 0; background-image: url(&apos;data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAARCAYAAADdRIy+AAAACXBIWXMAABJ0AAASdAHeZh94AAACiUlEQVQ4y42U6XKjQAyEef8X3Gxim3M4DT4xxnai7U+E1P7ZqqVqmBkhaVrqHqK+H6xuWsuKYHXdWBZKX2/jxOeqri0vCiur2pI8+FyUle0Vx/P19eVjXUe8Pj8/7fV6ybDM7B+Phw/WT7ctgXx/Pp8+//2siaO+7+33ZmO7NLW3943FaWa5UIIwyTILQg3KWN8/tluLs9y6fW/zPFu732v0djidfhJHt2myw/Fo58vFhsPBLterr0/ni6/H281ut0n784/fON7sIZTE4ne/3x011URHORWhsn44CGVmTbe3UJbqa+P9ytQ/bHkIjjhVX5u2tVx9rIQ+TjL3zfLC6razaDgMcswd3a/3D09WBEhJbZekFiAhLGTsFMy6kE+qtmBrlIQ2bXY7ryg6n86evem6ZZYDSUsxX9WtK2CxVTr06IhT+dFHkK6Ml+rzqNZEnRoL9FBVLpNW5dWeTCUzZD/q0G0cezCEUf7qs8aCzkmh2an0lRfBSwdBpn0iVmG0UiAJOIwqCk9YedkkWxQR20YtGsRH1A+DvX1s3CkPi6ApKVWyQidzGIGgAwl2mIXV+Vurd0lomu6+j6B/HEeXgH+UBOaZefYgHJkJ8qHviPxfT4QDOmOgORJcdcCkAy7X0Q9jxo5twl8zelz1eVb/rq7ZaSk5TnPvDRKpXHuLNLDTH+4wJGCDbcpHKqiCNsAwvefWOCkkatVwCKgQNER8J+Xase+lU9YkhDT8OARVrIqgCr8pSbb8TWrpDjGDKJUNllfkMEwgN8X/PEKI+DkgBa38/OpxQ4DvCdvOryE/AlAtaIofcYOQaxa+f2FJnru///6kEHodnSUBnJHK0rtFHon2XDOCKeV/nz8newEInqX7qwAAAABJRU5ErkJggg==&apos;); background-size: cover; display: block;&quot;
  &gt;&lt;/span&gt;
  &lt;img
        class=&quot;gatsby-resp-image-image&quot;
        alt=&quot;tweet&quot;
        title=&quot;tweet&quot;
        src=&quot;/static/9cf90a11e418ecd73f81e5178a9328c2/fcda8/tweet.png&quot;
        srcset=&quot;/static/9cf90a11e418ecd73f81e5178a9328c2/12f09/tweet.png 148w,
/static/9cf90a11e418ecd73f81e5178a9328c2/e4a3f/tweet.png 295w,
/static/9cf90a11e418ecd73f81e5178a9328c2/fcda8/tweet.png 590w,
/static/9cf90a11e418ecd73f81e5178a9328c2/1b853/tweet.png 592w&quot;
        sizes=&quot;(max-width: 590px) 100vw, 590px&quot;
        style=&quot;width:100%;height:100%;margin:0;vertical-align:middle;position:absolute;top:0;left:0;&quot;
        loading=&quot;lazy&quot;
      /&gt;
  &lt;/a&gt;
    &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;This screenshot looks like it was from an interview. What interview? I scrolled back to October 2024 on Tech Fund’s Twitter feed and saw a bunch of &lt;a href=&quot;https://x.com/techfund1/status/1828858794480140391?s=20&quot;&gt;similarly-formatted&lt;/a&gt; &lt;a href=&quot;https://x.com/techfund1/status/1826875528751534448?s=20&quot;&gt;screenshots&lt;/a&gt;, some of which were cited as coming from &lt;a href=&quot;https://tegus.com/&quot;&gt;Tegus&lt;/a&gt;. Tegus is apparently a company with a &lt;a href=&quot;https://www.reddit.com/r/expertnetworks/comments/1ghe2ls/tegus_analyst_reached_out_to_me/&quot;&gt;business model&lt;/a&gt; of reaching out to insiders (in this case, AI company employees) and paying them hundreds of dollars an hour in order to answer specific technical questions. It’s essentially gig work for &lt;em&gt;almost-but-not-quite&lt;/em&gt; insider trading: the more informed and confident you sound, the more likely Tegus analysts will pick you for future interviews.&lt;/p&gt;
&lt;p&gt;I’m sure the source for this tweet is in fact a GenAI principal architect, since Tegus would have presumably asked for some proof of that before they paid them out. But it’s pretty clear that the incentives here are to sound confident and authoritative, even on questions that you’re not sure about. With that in mind, the quote itself also reads a bit suspiciously. I’ve worked with enough principal engineers and architects to take their casual back-of-envelope estimates with a grain of salt. If they knew the actual rate at which GPUs fail and get retired in Google datacenters, wouldn’t they have just said that?&lt;/p&gt;
&lt;h3&gt;Evidence for a longer lifespan&lt;/h3&gt;
&lt;p&gt;We have some anecdotal evidence that points the other way. Google has &lt;a href=&quot;https://www.datacenterdynamics.com/en/news/google-says-tpu-demand-is-outstripping-supply-claims-8yr-old-hardware-iterations-have-100-utilization&quot;&gt;publicly claimed&lt;/a&gt; to have eight year old TPUs (their version of GPUs) running in production at “100% utilization”. Nvidia only made A100 GPUs from &lt;a href=&quot;https://www.amax.com/nvidia-h100-vs-nvidia-a100/&quot;&gt;2020-2024&lt;/a&gt;, but in February 2026 the AWS CEO &lt;a href=&quot;https://www.datacenterdynamics.com/en/news/aws-has-never-retired-an-nvidia-a100-server-ceo-matt-garman-claims/&quot;&gt;claimed&lt;/a&gt; that AWS had never retired an A100 server (and you can still easily rent A100s for AI work)&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;. AI GPU usage isn’t exactly like crypto mining GPU usage, but it certainly seems like years-old ex-crypto GPUs are &lt;a href=&quot;https://www.youtube.com/watch?v=UFytB3bb1P8&quot;&gt;functional&lt;/a&gt;. There’s also &lt;a href=&quot;https://news.ycombinator.com/item?id=48456717&quot;&gt;this comment&lt;/a&gt; from Hacker News I noticed where someone claims that their GPU cluster in academia has lasted six years with less than 20% failure rate.&lt;/p&gt;
&lt;p&gt;What about hard data? It’s hard to get concrete data on the lifespan of AI GPUs, because modern AI datacenters have only existed for a handful of years. But an interesting case study would be recent supercomputer clusters like Oak Ridge’s &lt;a href=&quot;https://www.datacenterdynamics.com/en/news/oak-ridge-national-laboratory-to-retire-summit-supercomputer-in-november-2024/&quot;&gt;Summit&lt;/a&gt;, which had over 27 thousand Nvidia V100s running from 2018 to 2024, or its predecessor, the Cray &lt;a href=&quot;https://christian-engelmann.de/publications/ostrouchov20gpu.pdf&quot;&gt;Titan&lt;/a&gt; supercomputer that ran from 2012 to 2019. I couldn’t find any evidence that Summit had to buy an additional 27,000 GPUs to replace their old ones, and GPU failures in Titan have been &lt;a href=&quot;https://christian-engelmann.de/publications/ostrouchov20gpu.pdf&quot;&gt;carefully studied&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;span
      class=&quot;gatsby-resp-image-wrapper&quot;
      style=&quot;position: relative; display: block; margin-left: auto; margin-right: auto; max-width: 590px; &quot;
    &gt;
      &lt;a
    class=&quot;gatsby-resp-image-link&quot;
    href=&quot;/static/3922b6bf4aaf6d3202cacd2f472034aa/019a6/figure10.png&quot;
    style=&quot;display: block&quot;
    target=&quot;_blank&quot;
    rel=&quot;noopener&quot;
  &gt;
    &lt;span
    class=&quot;gatsby-resp-image-background-image&quot;
    style=&quot;padding-bottom: 60.810810810810814%; position: relative; bottom: 0; left: 0; background-image: url(&apos;data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAAMCAYAAABiDJ37AAAACXBIWXMAABYlAAAWJQFJUiTwAAAB5ElEQVQoz3VT7ZKbMBDL+79UH6E/e53m0gTCVwjGgG2MDaha53JNOy0zmxjZq5V3xQF8+l4ju2S4nM84Ho8oigLX6zVFnl9xPP3E9+MPlEX5ieV5jvfTCSdGWRbp3VqLw7ZtGMcRIUT4ecbiPWKMWJYFxliEJWKwDtpZbOsKv3jiJv075+D9gm1bMU0T6rrGYeUhQ2Y7KlgzYd8BwUIImFlACu4f8Te+fWA7k4RQ1CeFlko6Eupp+CNRFMhaFEv8DxdCuaVc+yCLvlfcAK8Zk0I5EOOrkv2fCvcXhdK/pmlw6PsebdtiZtVpMuhVy9AYRodhGKie/dMKShuemZloiI+Ps8wVtVJA1olQGqyUShUf15mg7jl0d4ObPRVscHZE1eRodOT+kopvxNu2Rj84+mRPc0iEIldIfzcY2HZOzdRMfuC8MeaxxP2W425WKnIJi8HizMlGCxi2Lsuyx1BeCSVkHeP62auEkaCrvmLqLrQMz9Eq/EF9e8O3akDXWlRV8erDkLwnEcKSyKTQE5cCzhvooWIf6Vf6MEQ6pM3wdv6CsqvR1LyyfCmvZM/wNLjE0yYyIDG/OMF7En7kCOkyWwow6T31UNSIMbXWSa0kS8i7hAxNpih7SnUJE/WPYB7Pylclzy+EUp6sFg3U1gAAAABJRU5ErkJggg==&apos;); background-size: cover; display: block;&quot;
  &gt;&lt;/span&gt;
  &lt;img
        class=&quot;gatsby-resp-image-image&quot;
        alt=&quot;fig10&quot;
        title=&quot;fig10&quot;
        src=&quot;/static/3922b6bf4aaf6d3202cacd2f472034aa/fcda8/figure10.png&quot;
        srcset=&quot;/static/3922b6bf4aaf6d3202cacd2f472034aa/12f09/figure10.png 148w,
/static/3922b6bf4aaf6d3202cacd2f472034aa/e4a3f/figure10.png 295w,
/static/3922b6bf4aaf6d3202cacd2f472034aa/fcda8/figure10.png 590w,
/static/3922b6bf4aaf6d3202cacd2f472034aa/efc66/figure10.png 885w,
/static/3922b6bf4aaf6d3202cacd2f472034aa/c83ae/figure10.png 1180w,
/static/3922b6bf4aaf6d3202cacd2f472034aa/019a6/figure10.png 1818w&quot;
        sizes=&quot;(max-width: 590px) 100vw, 590px&quot;
        style=&quot;width:100%;height:100%;margin:0;vertical-align:middle;position:absolute;top:0;left:0;&quot;
        loading=&quot;lazy&quot;
      /&gt;
  &lt;/a&gt;
    &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;These cages of GPUs are stacked vertically, and cold air is pumped in from the bottom, which explains why cage 0 (at the bottom) has better survival rates than cage 2 (at the top). Let’s consider cage 0, so we’re just looking at the GPU lifespan instead of at the lifespan of improperly-cooled GPUs. At three years, over 95% of GPUs survived&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;. At six years, nodes 2 and 3 (the GPUs closest to the bottom of the cage) were still at above 90% survival rate, and the highest nodes were over 60%.&lt;/p&gt;
&lt;p&gt;It’s possible that newer Nvidia GPUs are less reliable than older ones (they certainly draw more power), or that AI datacenters are under-cooled, or that something about LLM utilization is more stressful than the workloads that ran on traditional GPU datacenters. But this is at least circumstantial evidence that GPUs can survive under load for far longer than three years.&lt;/p&gt;
&lt;h3&gt;Economic lifespans&lt;/h3&gt;
&lt;p&gt;This discussion is complicated by the fact that GPUs may have a short &lt;em&gt;economic&lt;/em&gt; lifespan. Supposedly a B100 GPU &lt;a href=&quot;https://bizon-tech.com/blog/nvidia-b200-b100-h200-h100-a100-comparison?srsltid=AfmBOoqugX-R8Y9AoVlyxRMheglf4gJ2Xc5hefXVxL6Cv3Htl0P_rHx1&quot;&gt;draws&lt;/a&gt; twice as much power as an A100, but can do five times as much work. For some AI providers, that might mean that A100s are only worth running until they can be replaced with B100s (if you’re bottlenecked on electricity, you should spend it all on B100s and throw out your obsolete A100s). This is why the Titan supercomputer was decommissioned in favor of Summit: it could have continued to operate, but it was more profitable to spend the money and maintenance effort on newer hardware.&lt;/p&gt;
&lt;p&gt;It should be obvious that this doesn’t support the “inference will become more expensive when the bubble pops” argument. So long as A100s are profitable &lt;em&gt;right now&lt;/em&gt;, cash-poor AI providers can continue profitably serving inference from them, even if there are more efficient options available for those with the capital to upgrade.&lt;/p&gt;
&lt;p&gt;On top of that, GPUs only represent one part of AI datacenter infrastructure spending. If your GPUs wear out, you don’t have to go and build an entirely new datacenter. About 30-50% of &lt;a href=&quot;https://epoch.ai/assets/images/data-insights/ai-datacenter-cost-breakdown/ai-datacenter-cost-breakdown-upfront.png&quot;&gt;datacenter&lt;/a&gt; &lt;a href=&quot;https://www.reuters.com/commentary/breakingviews/how-big-techs-630-bln-ai-splurge-will-fall-short-2026-03-26/&quot;&gt;spend&lt;/a&gt; goes to land, power, cooling, and so on. The remaining 50-70% is the cost of the entire server rack, which includes a bunch of things that aren’t GPUs.&lt;/p&gt;
&lt;h3&gt;Conclusion&lt;/h3&gt;
&lt;p&gt;Like the idea that AI inference &lt;a href=&quot;/water-impact-of-ai/&quot;&gt;requires using huge amounts of water&lt;/a&gt;, the idea that AI GPUs only live a year or two is popular because it’s a useful idea for AI skeptics, not because it’s true. It comes from a pseudonymous tweet quoting an anonymous source who’s being paid hundreds of dollars to sound like a credible expert on AI. Other public communications from AI inference providers cite much higher lifespan numbers, and the statistics from supercomputers (the traditional examples of large GPU clusters) don’t bear out the claim that the maximum lifespan is three years.&lt;/p&gt;
&lt;p&gt;It might be true that the &lt;em&gt;economic&lt;/em&gt; lifespan is three years, in a world where new GPUs come out every eighteen months and GPU providers are flush with cash to upgrade, but that doesn’t tell us much about the economics of inference in an AI winter. If money becomes a lot more scarce, it’s likely that AI datacenters will continue profitably&lt;sup id=&quot;fnref-4&quot;&gt;&lt;a href=&quot;#fn-4&quot; class=&quot;footnote-ref&quot;&gt;4&lt;/a&gt;&lt;/sup&gt; running their B300s (or their H100s or even A100s) for six years or longer.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;Of course, like previous claims about AI and water usage, “three years at the most” is often cited as &lt;a href=&quot;https://ithy.com/article/data-center-gpu-lifespan-explained-7mpjwwyp&quot;&gt;“1-2 years, with some lasting up to 3 years under optimal conditions”&lt;/a&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;Of course, pronouncements from CEOs/CTOs should be taken with a grain of salt as well (for instance, maybe they have a big backlog of unused A100s they keep swapping out), but (a) executives don’t often straight-up lie about concrete technical facts, and (b) they’re going up against an unsourced quote from a tweet, so the bar isn’t that high.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;What about proactive GPU replacement? In the “Survival Analysis” section, the study attempts to account for this. I haven’t dug into exactly how.&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-4&quot;&gt;
&lt;p&gt;Assuming inference is profitable, which I believe (when you’re not attempting to amortize the cost of training).&lt;/p&gt;
&lt;a href=&quot;#fnref-4&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[Doing nothing at work]]></title><link>https://seangoedecke.com/doing-nothing-at-work/</link><guid isPermaLink="false">https://seangoedecke.com/doing-nothing-at-work/</guid><pubDate>Mon, 08 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Many engineers should be doing less work. I don’t necessarily mean producing less code or fewer changes, but literally working fewer hours in the day. When they do work, they should be working at a slower pace. I like to aim to be running at 80% utilization by default: unless I have a high-pressure project going on, I spend 20% of my workday away from the computer.&lt;/p&gt;
&lt;h3&gt;High-impact opportunities&lt;/h3&gt;
&lt;p&gt;Why? &lt;strong&gt;Performance at tech companies is dominated by outlier events&lt;/strong&gt;. When I think about the most impactful changes I’ve made, many of them involved a surprisingly trivial amount of work. There are no points for effort in software development. What matters is solving the right problem at the right time.&lt;/p&gt;
&lt;p&gt;In large engineering organizations, there are usually trivial pieces of engineering work you could do that would make tens or hundreds of millions of dollars for the company. Here are three common examples:&lt;/p&gt;
&lt;p&gt;First, when the company is trying to sign a big enterprise deal, stepping in with a feature or bugfix can make the deal happen. It doesn’t even have to be a &lt;em&gt;good&lt;/em&gt; feature: sometimes just showing that you’re willing and able to make a concrete change will be enough.&lt;/p&gt;
&lt;p&gt;Second, preventing or mitigating an incident early (even by just knowing the right feature flag to turn off) can save huge amounts of money: both immediate lost revenue during the incident and future lost revenue from customers who would have pulled their business or refused to sign pending contracts.&lt;/p&gt;
&lt;p&gt;Third, when the company is trying to ship a high-profile feature, success or failure often hinges on trivial but obscure changes (e.g. the ability to rapidly add a new field in user settings, or to update the crufty enterprise-data-export functionality nobody has touched in years). Familiarity with the system can be the difference between one of these changes taking a few hours or a whole week.&lt;/p&gt;
&lt;p&gt;What do these examples have in common? They’re all &lt;em&gt;time-dependent&lt;/em&gt;. You can’t just log on in the morning and decide to unblock a big deal, or mitigate an incident, or speed up a high-profile feature. Is it just a matter of being in the right place at the right time? Not quite. &lt;strong&gt;You also have to not already be busy.&lt;/strong&gt;&lt;/p&gt;
&lt;h3&gt;Staying loose&lt;/h3&gt;
&lt;p&gt;I wrote about this a couple of years ago in &lt;a href=&quot;https://www.seangoedecke.com/party-tricks/&quot;&gt;&lt;em&gt;Crushing JIRA tickets is a party trick, not a path to impact&lt;/em&gt;&lt;/a&gt;. If you’re always 100% utilized on a steady stream of low-priority work (for instance, if you’re just picking up tickets from the backlog, crushing them, then picking up the next one), you’ll miss your chance to do high-impact work in two ways.&lt;/p&gt;
&lt;p&gt;First, you’ll be too busy to even &lt;em&gt;notice&lt;/em&gt; the opportunities. You won’t be chatting with people who are working on other things, or reading team updates, or keeping an eye on ongoing incidents. So you’ll miss out on the best way to get involved in high-impact work, which is to volunteer your expertise.&lt;/p&gt;
&lt;p&gt;Second, if you perpetually look busy, your manager won’t want to volunteer for you. This is the second-best way to get involved in high-impact work: to have your manager or product manager say “oh, Sean has capacity to help out here, let me tag him in”. Why is this better? Because managers and product managers usually have a much better read on what high-impact work is going on. They’re in meetings that you aren’t in.&lt;/p&gt;
&lt;h3&gt;Doing nothing&lt;/h3&gt;
&lt;p&gt;If you’re supposed to keep your time free for high-impact work, and you’re not supposed to just grind tickets, what should you be doing on a minute-by-minute basis? Should you just be doing nothing? Yep!&lt;/p&gt;
&lt;p&gt;Doing nothing is good, actually. Software engineering can be a stressful job, but it’s typically not &lt;em&gt;consistently&lt;/em&gt; stressful: the stress comes from the occasional incident, or high-pressure urgent piece of work, or (these days) layoff. If you approach the comparatively low-pressure parts of your work with urgent intensity, you’ll already be exhausted and frazzled when you have to handle the high-pressure parts.&lt;/p&gt;
&lt;p&gt;Even in high-pressure parts of the job, doing nothing can still be good. One thing I recommend for engineers new to on-call is to avoid rushing: take a few breaths before joining the call or before speaking, and in general try to &lt;a href=&quot;/thinking-clearly/&quot;&gt;“think in slow motion”&lt;/a&gt;. Most incidents resolve on their own. Most frantic “maybe this will help” changes during incidents make things worse, not better. As a general rule, if you can simply avoid panicking, you will be doing better than most engineers at incident response.&lt;/p&gt;
&lt;p&gt;Nothing is a space things can happen in&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;. If you give your brain a chance to rest, you will find you’re more likely to have new ideas. If someone hands you an important task, you can tackle it with your full attention (instead of juggling it with the three other things you’re working on in the background). When you’re not busy, you have time to just &lt;em&gt;look at things&lt;/em&gt; and take in new data.&lt;/p&gt;
&lt;h3&gt;Deliberately not doing specific things&lt;/h3&gt;
&lt;p&gt;A lot of engineers are uncomfortable seeing a task that needs doing and not doing it. I’m like this as well. I wrote about it in &lt;a href=&quot;/addicted-to-being-useful/&quot;&gt;&lt;em&gt;I’m addicted to being useful&lt;/em&gt;&lt;/a&gt;: it’s a psychological quirk that many software engineers share, because having that quirk (to a point) makes you a good fit for the job. In order to spend time doing nothing, sometimes you need to force yourself to not step in.&lt;/p&gt;
&lt;p&gt;For instance, I believe that &lt;strong&gt;engineers should generally avoid glue work&lt;/strong&gt;&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;. Most glue work — making sure people talk to each other, updating docs for work you’re not leading, volunteering to address technical debt — reflects the fact that the organization is not explicitly prioritizing this work. If they were, you wouldn’t need to volunteer for it. Either that’s fine, or it’s a big mistake. If it’s fine, then you shouldn’t step up and do it: you’ll be wasting your time and annoying your manager. If it’s a big mistake, &lt;em&gt;you still shouldn’t do it&lt;/em&gt;, because you’ll be insulating the company from the consequences of its own mistakes at the cost of your own career and mental well-being.&lt;/p&gt;
&lt;p&gt;That’s a bad deal for you, and a bad example for your junior colleagues, and sets a bad precedent for someone else to jump into the same position when you inevitably burn out&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;. If the consequences truly are severe, let them happen, so the organization can feel the pain and change its policies.&lt;/p&gt;
&lt;p&gt;I also believe that &lt;strong&gt;being too helpful leaves you vulnerable to predators&lt;/strong&gt;. Tech companies are full of people who want to extract uncompensated work from software engineers&lt;sup id=&quot;fnref-4&quot;&gt;&lt;a href=&quot;#fn-4&quot; class=&quot;footnote-ref&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;. This is different from work that arrives via normal channels, and for which you’re compensated by promotions, bonuses (and just your normal salary). I’m talking about work that arrives via backchannels, from people who don’t have the ability or willingness to ensure that work is formally recorded under your name. For instance, a product manager from another organization messaging you to say “you’re so good at querying data, would you mind pulling some statistics for me about X?”, or an engineer from another team asking you to “pair” on a piece of work that will ultimately involve you writing all the code and them quietly submitting the change under their own name.&lt;/p&gt;
&lt;p&gt;Doing some amount of this kind of work is fine. You may as well help people out when you can. But you need to be able to apply backpressure, either by saying no or simply delaying your response by a few hours or days.&lt;/p&gt;
&lt;p&gt;It’s also a good idea to &lt;strong&gt;avoid investing too much in work that is likely going to disappear&lt;/strong&gt;. For instance, suppose you’re working with a product designer who is figuring out what they want in real time. At 9am they message you saying they want the page header to look one way, then at 10am they have tweaks, and more changes at 11am, and so on. You should not throw yourself into fully rewriting the page every hour. Instead, you should do nothing (say, go for a walk) and rewrite the page once in the afternoon, based on the most recent design. Another common instance of this is “big idea from a manager without the political clout to follow through on it”. Often you can just run out the clock until the project gets inevitably cancelled&lt;sup id=&quot;fnref-5&quot;&gt;&lt;a href=&quot;#fn-5&quot; class=&quot;footnote-ref&quot;&gt;5&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;h3&gt;Conclusion&lt;/h3&gt;
&lt;p&gt;A lot of software engineering advice and tooling is designed around the ability to scale up your ability to exert technical effort: to do more things at the same time, to take on projects of larger scope, or to just write more code. But software engineering success is not determined by any of these. It is determined by the ability to do the &lt;em&gt;right&lt;/em&gt; things at the &lt;em&gt;right&lt;/em&gt; time, which requires that you deliberately hold back some of your effort during ordinary work.&lt;/p&gt;
&lt;p&gt;In my experience, it’s still possible to be a “high performing engineer” at 80% effort. In fact, it’s &lt;em&gt;easier&lt;/em&gt;, because you’ll be less likely to make silly mistakes from stress, and you’ll be in a position to jump on the kind of high-impact tasks that deliver outsized returns.&lt;/p&gt;
&lt;p&gt;This doesn’t mean you should never grind at 100% effort. I think there are probably two or three times a year where I work as hard as I possibly can: long hours, intense focus, thinking about the problem from when I wake up to when I go to bed. But I reserve this mode of work for &lt;a href=&quot;/the-spotlight&quot;&gt;when the rewards are really high&lt;/a&gt;. For the rest of the year, I take it relatively easy.&lt;/p&gt;
&lt;p&gt;edit: this post got some comments on &lt;a href=&quot;https://news.ycombinator.com/item?id=48442880&quot;&gt;Hacker News&lt;/a&gt;. Commenters discuss &lt;a href=&quot;https://news.ycombinator.com/item?id=48446245&quot;&gt;how to not get in trouble&lt;/a&gt; with your manager when you’re taking slack time (in my experience, if you’re generally productive it’s fine, but managers vary a lot) and &lt;a href=&quot;https://news.ycombinator.com/item?id=48443273&quot;&gt;whether engineers really do have control&lt;/a&gt; over their workload.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;One of my big influences is Rich Hickey’s talk &lt;a href=&quot;https://github.com/matthiasn/talk-transcripts/blob/master/Hickey_Rich/HammockDrivenDev.md&quot;&gt;&lt;em&gt;Hammock Driven Development&lt;/em&gt;&lt;/a&gt;. This is &lt;em&gt;kind of&lt;/em&gt; like what he’s talking about, except (a) Hickey is more talking about what it takes to design solutions to really hard problems, rather than what it takes to be a strong engineer in an ordinary tech company, and so (b) Hickey recommends using your time-away-from-the-computer to focus on a hard problem, instead of to simply decompress and let solutions congeal in your head. It’s also like Zvi Mowshowitz’s post on &lt;a href=&quot;https://thezvi.substack.com/p/slack&quot;&gt;“slack”&lt;/a&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;I wrote about this a lot more in &lt;a href=&quot;/glue-work-considered-harmful/&quot;&gt;&lt;em&gt;Glue work considered harmful&lt;/em&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;Why inevitably? Because in my view, burnout is &lt;em&gt;hard work unrewarded&lt;/em&gt;, and taking on a personal crusade that your job doesn’t care about is a great way to do a lot of unrewarded work.&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-4&quot;&gt;
&lt;p&gt;I wrote about this in &lt;a href=&quot;/predators&quot;&gt;&lt;em&gt;Protecting your time from predators in large tech companies&lt;/em&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-4&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-5&quot;&gt;
&lt;p&gt;Of course, you have to be careful with this. If you try this strategy and you’re wrong about the level of political support for the project, you will come off like a slacker and then have to deliver in a rush.&lt;/p&gt;
&lt;a href=&quot;#fnref-5&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[Working with product managers]]></title><link>https://seangoedecke.com/working-with-product-managers/</link><guid isPermaLink="false">https://seangoedecke.com/working-with-product-managers/</guid><pubDate>Mon, 08 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;The relationship engineers have with product management is more dysfunctional than with any other part of the company. There’s no shared culture or language like there is with other engineers, and the rules of “who gets to tell who what to do” aren’t as clear-cut as they are with managers. Engineers don’t have a lot in common with legal, or design, or sales, but they also don’t need to interact much with those roles. In my experience, engineers are communicating with product managers almost every single day.&lt;/p&gt;
&lt;h3&gt;Against the “product mommy”&lt;/h3&gt;
&lt;p&gt;The worst version of the product/engineering relationship goes something like this:&lt;/p&gt;
&lt;p&gt;Engineers are technically competent but are too autistic to be fully trusted. They need a kind-but-stern parental figure who knows how to communicate to other stakeholders in the organization (for instance, by being comfortable using the word “stakeholders”), and how to keep engineers from going off in the wrong direction.&lt;/p&gt;
&lt;p&gt;This entire gross dynamic is neatly captured by the popular term &lt;a href=&quot;https://x.com/search?q=%22product%20mommy%22&amp;#x26;src=typed_query&quot;&gt;“product mommy”&lt;/a&gt;&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;.
I really, really don’t like that term, or this entire dynamic in general. Almost none of my relationships with my product managers have been anything like this, though I have seen it at a distance.&lt;/p&gt;
&lt;p&gt;Working well with product managers can be the difference between succeeding and failing at a company. Why is it so hard to maintain good relationships between engineering and product? What does a good relationship look like?&lt;/p&gt;
&lt;h3&gt;Why it’s so hard to build trust&lt;/h3&gt;
&lt;p&gt;Product managers and engineers have largely non-overlapping skillsets. Product managers don’t understand the technical work engineers do and aren’t equipped to talk about it: if an engineer gives a technical reason for something, product managers generally have to shrug and say “sure, I guess”. Likewise, engineers don’t have anything like the visibility into the organization that product managers do. Particularly in large organizations, it is the product manager who is the source of truth about who wants what and which features are important. When a product manager says that something is critical, engineers generally have to shrug and say “sure, I guess”.&lt;/p&gt;
&lt;p&gt;This obviously requires a lot of trust. What’s a little less obvious is that &lt;strong&gt;this trust is continually broken by both sides&lt;/strong&gt;. Every single product manager has been told &lt;em&gt;thousands&lt;/em&gt; of times that technical task X is technically impossible or would be disastrous, only for that task to end up being done fairly smoothly and successfully. Every single engineer has been told &lt;em&gt;thousands&lt;/em&gt; of times that requirement X is absolutely critical and worth going to enormous effort for, only for that requirement to be silently dropped or changed with no apology.&lt;/p&gt;
&lt;p&gt;Of course this isn’t malicious. Engineers often give wrong estimates because &lt;a href=&quot;/how-i-estimate-work/&quot;&gt;estimation is impossible&lt;/a&gt;, and sometimes the dire consequences they warn about really do happen (they’re just handled behind the scenes, like engineers handle many other kinds of technical dysfunction). Product managers “change their minds” because what’s important in a large tech company does genuinely change hour-by-hour&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;, and even the best attempts to only filter the most reliable priorities through to the engineering team will sometimes go wrong.&lt;/p&gt;
&lt;h3&gt;Manipulation and lies&lt;/h3&gt;
&lt;p&gt;The consequence of this broken trust is that the relationship becomes very difficult to maintain. When you’re an engineer, and you explain something to your product manager, and you &lt;em&gt;know&lt;/em&gt; they don’t believe you (despite having no ability themselves to judge the question), it can be incredibly frustrating. Likewise, when you’re a product manager, and you’re desperately trying to explain what we need to do to an engineer, and you know they’re internally shrugging their shoulders, it must be unbearable. Don’t they know this is critical to the company? You were just in a meeting with the leaders of the organization!&lt;/p&gt;
&lt;p&gt;The natural tool for a mistrustful product manager is &lt;em&gt;manipulation&lt;/em&gt;. I still remember a product manager who tried to extract a commitment from my team by asking us to go around and all say “I commit to getting this work done in two weeks”, after a conversation where we’d explained the risks that cause it to take longer. I suppose the idea was that we’d all work much harder, having taken a sacred oath? More subtle variants of this approach involve suggesting that you would be really disappointed if this work was delayed (in true “product mommy” style), or vaguely suggesting the possibility of some abstract reward (that the product manager is not empowered to deliver) if work gets done ahead of schedule.&lt;/p&gt;
&lt;p&gt;The natural tool for a mistrustful engineer is &lt;em&gt;lies&lt;/em&gt;. The most benign version of this is exaggerating estimates: for instance, the classic advice to &lt;a href=&quot;https://news.ycombinator.com/item?id=19671824&quot;&gt;double your estimate and add 20%&lt;/a&gt;. I’ve seen engineers claim that they’ve had to follow up on all sorts of largely-fake tasks (one common example is “reaching out to a neighbor team to confirm X”) in order to gain more time. In the worst case, engineers might even straight-out lie that work has been completed, and then track the “it doesn’t work in production” feedback as a bug.&lt;/p&gt;
&lt;p&gt;Once this starts happening, it’s nearly impossible to repair the relationship. I can’t bring myself to trust a product manager who’s clearly trying to pull my strings, and I’m sure a product manager can’t trust an engineer who’s lied to their face in the past. That’s why it’s so important to avoid getting into a bad relationship in the first place.&lt;/p&gt;
&lt;h3&gt;Don’t fight with the product manager&lt;/h3&gt;
&lt;p&gt;Why bother? If it’s so hard to hammer out a good working relationship with product managers, why not just settle for a bad one? Product managers can absolutely &lt;em&gt;bury&lt;/em&gt; you if you’re not careful.&lt;/p&gt;
&lt;p&gt;Product managers are almost always more politically sophisticated than engineers. This is partly structural: product managers are simply in more conversations with the company’s movers and shakers, and so naturally have a better relationship with them (and are thus better attuned to which way the wind is blowing). It’s also partly selection bias: engineers can be hired even with relatively poor social skills, because they’re primarily being assessed on technical ability, but social skills are a core part of the product role&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;If you are feuding with a product manager, you will probably lose&lt;/strong&gt;. Unless you’re unusually influential, they will simply have far more opportunities to quietly talk you down in influential circles than you will. All it takes is a few comments like “oh, I probably wouldn’t pick Sean for that project” to wreck your reputation. In the case where you are &lt;em&gt;openly&lt;/em&gt; feuding with a product manager, the company’s leaders will by default take the product manager’s side over yours. They’re likely to know them better, have more shared cultural context with them, and in general be willing to interpret the situation as “another engineer who doesn’t understand how the organization works”.&lt;/p&gt;
&lt;p&gt;There are huge benefits to being trusted by a product manager. Product managers &lt;em&gt;want to ship things&lt;/em&gt;, and typically understand a fair amount about all of the non-technical barriers to shipping. If you also want to ship things, you can become a fearsome team.&lt;/p&gt;
&lt;p&gt;On top of that, because trust between engineers and product managers is so difficult, once you’re in you’re in all the way. Product managers often pick one or two engineers as their go-to for getting the “real story” on technical questions. If that’s you, you have an outsized position of influence in the organization, which you can use to &lt;a href=&quot;/how-to-influence-politics/&quot;&gt;get the things you want done&lt;/a&gt;.&lt;/p&gt;
&lt;h3&gt;How can you build trust with product managers?&lt;/h3&gt;
&lt;p&gt;As an engineer, how can you build trust with your product manager?&lt;/p&gt;
&lt;p&gt;The first step is to &lt;strong&gt;understand where they’re coming from&lt;/strong&gt;. When they tell you something is important or that a requirement has come in, be aware that this is rarely their decision. It’s not them who’s jerking you around, it’s someone higher up in the food chain jerking you both around. If you can adopt a conspiratorial mindset &lt;em&gt;with&lt;/em&gt; them, instead of &lt;em&gt;against&lt;/em&gt; them, that’s a good start. Try just asking “oh man, alright, what can we do about this?” instead of complaining.&lt;/p&gt;
&lt;p&gt;The second step is to &lt;strong&gt;be right, a lot&lt;/strong&gt;. This is a silly-sounding Amazon leadership principle that turns out to be entirely accurate. I wrote more about it &lt;a href=&quot;https://www.seangoedecke.com/being-right-a-lot/&quot;&gt;here&lt;/a&gt;, but (as unfair as it sounds) you really do have to be mostly accurate if you want to build trust with a product manager. When you say something will ship, it has to ship; when you say something is impossible, it can’t happen days or weeks later. It’s okay to be wrong &lt;em&gt;sometimes&lt;/em&gt;, but you have to establish a pattern of you providing them useful, correct technical information.&lt;/p&gt;
&lt;p&gt;The third step is to &lt;strong&gt;let them make the political calls most of the time&lt;/strong&gt;. If you expect them to trust your technical calls, you have to extend them the same trust when it comes to navigating the organization. Don’t publicly undermine them in meetings, bring up your concerns in private. If they say something is important and you’re not so sure, at least act like it is. Accept that sometimes they’re going to be wrong, just like you’re sometimes wrong about technical questions.&lt;/p&gt;
&lt;p&gt;The fourth step is to &lt;strong&gt;get lucky&lt;/strong&gt;. Sometimes your product manager will just be a dud. You can’t build trust with someone incompetent: there’s nothing for you to trust them with, and they aren’t in a position where they can usefully extend trust to you. Working in large organizations requires getting comfortable with the fact that some of your colleagues will be stronger than others, and figuring out ways to work with (or bypass) people who make the work harder, not easier.&lt;/p&gt;
&lt;h3&gt;“Technical” product managers&lt;/h3&gt;
&lt;p&gt;Many product managers were once engineers. If your product manager is technical, does that make you immune from these problems? Absolutely not!&lt;/p&gt;
&lt;p&gt;You likely won’t have much choice in which product managers you work with, but be aware that having once been an engineer is a &lt;em&gt;negative&lt;/em&gt;, not a positive. No product manager can ever be technical enough to matter, because &lt;a href=&quot;/you-cant-design-software-you-dont-work-on/&quot;&gt;they don’t work on the codebase&lt;/a&gt;: even if they were a full-time engineer, they wouldn’t have the time to build the specific context on the system they’d need to be a real participant in technical discussions. It’s thus better to have a product manager who knows they’re not technical than to have one who mistakenly thinks they might be.&lt;/p&gt;
&lt;p&gt;The worst-case scenario is an ex-engineering product manager who believes they’re technical enough to detect when engineers are lying to them. This kind of paranoia is an easy trap for “technical” product managers to fall into, particularly when they don’t have a trusted engineer on the team they can lean on. If you’re dealing with one of these, prepare to spend a lot of time explaining why you can’t “just” do things (and prepare to have those explanations not be believed).&lt;/p&gt;
&lt;h3&gt;Conclusion&lt;/h3&gt;
&lt;p&gt;At its worst, a product manager relationship is like an unhealthy family: driven by condescension, emotional manipulation, lies, and mistrust. This isn’t because product managers are bad people! It’s because the structure of the relationship creates conflict. Both sides must make commitments (about the technical system or goals of the organization) that are (a) often wrong, and that (b) the other side is unable to independently verify. To avoid the trap, both sides have to be generous, willing to trust each other in their areas of expertise, and most importantly &lt;em&gt;competent&lt;/em&gt;.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;Unlike most roles in tech, product management (particularly the lower-level roles that are more engineer-facing) has close to an &lt;a href=&quot;https://www.productplan.com/blog/gender-diversity-better-products&quot;&gt;even&lt;/a&gt; gender split.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;For instance, based on the whims (or snap decisions, more charitably) of the CEO.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;I have worked with product managers with poor social skills, but it’s rare: about as rare as working with engineers with genuinely poor (i.e. by general-population standards) technical skills.&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[Anti-AI nostalgia and the cult of the past]]></title><link>https://seangoedecke.com/anti-ai-nostalgia/</link><guid isPermaLink="false">https://seangoedecke.com/anti-ai-nostalgia/</guid><pubDate>Thu, 04 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Programmers were better back in the day, weren’t they? Back when we had real programmers. Not just people who got paid to write code, but people who &lt;em&gt;lived&lt;/em&gt; it, who were obsessed with their craft, and whose code was a lively expression of themselves. Hackers were hackers in those days before money took over the industry.&lt;/p&gt;
&lt;p&gt;Don’t even get me started on LLMs. Could there be a better example of today’s degenerate spirit? A machine to mass-produce software (not good software, just barely good enough), so that the weak minds that dominate the industry can indulge their obsession with &lt;em&gt;quantity&lt;/em&gt;: of slop code, of features, and ultimately of money, which is the only way they can understand value. If they weren’t destroying our way of life, they would be pitiable. All of them together don’t have a fraction of the spiritual integrity of someone like &lt;a href=&quot;https://users.cs.utah.edu/~elb/folklore/mel.html&quot;&gt;Mel&lt;/a&gt;. But as it is, we must band together to crush them and drive them from our industry like the parasites they are.&lt;/p&gt;
&lt;h3&gt;Returning to the past&lt;/h3&gt;
&lt;p&gt;Okay, that’s not actually what I believe. But there sure are a lot of posts&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt; and comments on the internet that sound a bit like the paragraph above. Here are some older quotes that might sound similar:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;…the third collapse, in which power tends to pass into the hands of the lowest of the traditional castes, the caste of the beasts of burden and the standardized individuals. The result of this transfer of power was a reduction of horizon and value to the plane of matter, the machine, and the reign of quantity.&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;Usura rusteth the chisel \ It rusteth the craft and the craftsman \ It gnaweth the thread in the loom&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;The actual accomplishments of the past will nevertheless remain accomplishments, while the artistic stammerings of the painting, music, sculpture, and architecture produced by these types of charlatans will one day be nothing but proof of the magnitude of a nation’s downfall.&lt;sup id=&quot;fnref-4&quot;&gt;&lt;a href=&quot;#fn-4&quot; class=&quot;footnote-ref&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;These are all from the writings (or speeches) of famous fascists: Julius Evola, Ezra Pound, and Hitler himself. Mussolini’s &lt;a href=&quot;https://sjsu.edu/faculty/wooda/2B-HUM/Readings/The-Doctrine-of-Fascism.pdf&quot;&gt;&lt;em&gt;Doctrine of Fascism&lt;/em&gt;&lt;/a&gt; begins by defining fascism as a “spiritual attitude”, which the fascist man adopts in order to regain the mysterious qualities that were lost by the transition to modern life. In his classic &lt;a href=&quot;https://theanarchistlibrary.org/library/umberto-eco-ur-fascism&quot;&gt;&lt;em&gt;Ur-Fascism&lt;/em&gt;&lt;/a&gt;, Umberto Eco’s first two defining features of fascism are the “cult of tradition” and the “rejection of modernism”. So when someone tells me that the industry has lost its way and we must deny the corrupting influence of modern technology in order to &lt;a href=&quot;https://www.urbandictionary.com/define.php?term=retvrn&quot;&gt;retvrn&lt;/a&gt; to the time of virile &lt;a href=&quot;https://users.cs.utah.edu/~elb/folklore/mel.html&quot;&gt;real programmers&lt;/a&gt; (who understood and appreciated the spiritual dimension of programming), I get suspicious.&lt;/p&gt;
&lt;h3&gt;Fascism and crypto-fascism&lt;/h3&gt;
&lt;p&gt;It’s strange to describe anti-AI sentiment as potentially fascist, since a very &lt;a href=&quot;https://tante.cc/2026/04/21/ai-as-a-fascist-artifact/&quot;&gt;popular argument&lt;/a&gt; is that LLMs themselves are an inherently fascist tool. Surely both sides of the debate can’t be fascist? I do think that the structure of fascist arguments is &lt;a href=&quot;/many-anti-ai-arguments-are-conservative&quot;&gt;generally persuasive&lt;/a&gt;, and that many avowedly anti-fascist groups do sometimes fall into this trap: describing the world as a &lt;a href=&quot;https://en.wikipedia.org/wiki/300_(film)&quot;&gt;struggle&lt;/a&gt; between the spiritual power of the macho, traditional man and the corrupting influence of degenerate (often foreign) capital.&lt;/p&gt;
&lt;p&gt;For instance, I am a big fan of Lord of the Rings. I’ve read the series and watched the films multiple times, and even made a failed attempt to learn Elvish as a kid. But it’s hard to deny that fascists absolutely &lt;em&gt;love&lt;/em&gt; Lord of the Rings. “Marble statue of a Roman emperor” might be the most popular avatar for fascists on the internet, but Aragorn is the second most popular. Neo-fascist movements &lt;a href=&quot;https://www.cbc.ca/radio/tapestry/lord-of-the-rings-italy-1.6756668&quot;&gt;in Italy&lt;/a&gt; explicitly take up Lord of the Rings as a foundational text. Why? Because the core conflict in the text is between the traditional, nostalgic heroism of the Shire and Gondor, and the corrupting modern industrial (partly &lt;a href=&quot;https://tolkiengateway.net/wiki/Haradrim&quot;&gt;foreign&lt;/a&gt;) influence of Saruman and Sauron&lt;sup id=&quot;fnref-5&quot;&gt;&lt;a href=&quot;#fn-5&quot; class=&quot;footnote-ref&quot;&gt;5&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;I don’t think Lord of the Rings (or anti-AI rhetoric) is intrinsically fascist. In fact, the surface-level reading of the text is anti-fascist: the plucky people of the West banding together to fight Sauron’s command-and-control totalitarian society. But I can see why fascists love it.&lt;/p&gt;
&lt;h3&gt;The Luddites&lt;/h3&gt;
&lt;p&gt;One common historical touch-point for anti-AI folks is the Luddites, who were a violent conservative labor movement in early 1800s England. Anti-AI blogs adopt Luddite language like “smashing frames”, and positively &lt;a href=&quot;https://tante.cc/2026/04/21/ai-as-a-fascist-artifact/&quot;&gt;cite&lt;/a&gt; the Luddites as “the go-to enemies of fascism since its inception”. I’ve written at length about what we can learn from the Luddites in &lt;a href=&quot;/luddites-and-ai-datacenters/&quot;&gt;&lt;em&gt;Luddites and burning down AI datacenters&lt;/em&gt;&lt;/a&gt;, but one point I think is under-emphasized by the (generally pro-Luddite) books is that &lt;strong&gt;the Luddites were a little bit fascist themselves&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Brian Merchant’s &lt;a href=&quot;https://www.amazon.com.au/Blood-Machine-Origins-Rebellion-Against/dp/0316487740&quot;&gt;&lt;em&gt;Blood in the Machine&lt;/em&gt;&lt;/a&gt; is the most popular recent book on the Luddites. I enjoyed it, but Merchant’s attempts to paint the Luddites as a friendly, left-wing, proto-feminist movement&lt;sup id=&quot;fnref-6&quot;&gt;&lt;a href=&quot;#fn-6&quot; class=&quot;footnote-ref&quot;&gt;6&lt;/a&gt;&lt;/sup&gt; seemed really unconvincing to me. From the writings of the Luddites, it’s clear that they were interested in protecting the rights of their all-male elite guild fraternity. Here’s one Luddite threat to a workshop that explicitly includes a threat against the female workers&lt;sup id=&quot;fnref-7&quot;&gt;&lt;a href=&quot;#fn-7&quot; class=&quot;footnote-ref&quot;&gt;7&lt;/a&gt;&lt;/sup&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;We think it quite inconsistent with our duty as men, as husbands and as fathers to suffer ourselves to be ruined any longer by a set of vagabond strumpets and those gibbet-deserving rascals that are looking over them. We will lead them to their satisfaction. We sincerely hope, gentlemen, that you will discharge the bitches and take men into your employ again, or they must take what they get.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;These were fundamentally conservative people who felt (correctly) that modernity had deprived them of their elite status, handing it instead to lower-paid inferiors: women, vagabonds, and foreigners.&lt;/p&gt;
&lt;p&gt;The Luddites were obviously not fascists&lt;sup id=&quot;fnref-8&quot;&gt;&lt;a href=&quot;#fn-8&quot; class=&quot;footnote-ref&quot;&gt;8&lt;/a&gt;&lt;/sup&gt;. However, the basic ingredients were there: wounded pride, a masculine elite identity, hatred of modern economics, and violence aimed at restoring their previous position in society. The currents that produced Luddism are the same currents that guided so many unhappy people towards fascism. When things are looking grim for an elite group, they often turn towards any movement that promises a return to an idealized past.&lt;/p&gt;
&lt;h3&gt;Everything is permitted&lt;/h3&gt;
&lt;p&gt;If my blog has themes, one of them is surely that many software engineers labor under a delusion that their job is to be excellent at their craft. Of course, &lt;em&gt;wanting&lt;/em&gt; to be an excellent programmer is not a delusion; it is a completely legitimate value to hold, and a legitimate purpose to pursue. It’s just not what you’re paid to do at work. Your &lt;em&gt;job&lt;/em&gt;, unfortunately, is producing &lt;a href=&quot;/shareholder-value&quot;&gt;shareholder value&lt;/a&gt;. This delusion has been punctured by the &lt;a href=&quot;/good-times-are-over&quot;&gt;end of ZIRP&lt;/a&gt;, and again more recently by the rise of AI coding.&lt;/p&gt;
&lt;p&gt;In this environment, I worry that some software engineers will form exactly the kind of disillusioned elite that was the audience for Ezra Pound’s poems about “usury” or the Luddites’ campaign against unapprenticed (often female) textile workers. I worry that AI, and the companies that build AI, are becoming an enemy against which anything is permitted: an enemy which in Umberto Eco’s &lt;a href=&quot;https://theanarchistlibrary.org/library/umberto-eco-ur-fascism&quot;&gt;words&lt;/a&gt; is “at the same time too strong and too weak”, &lt;a href=&quot;/illusion-of-thinking/&quot;&gt;unable to reason&lt;/a&gt; and yet powerful enough to drastically reshape the global labor market for the worse.&lt;/p&gt;
&lt;h3&gt;Nuance&lt;/h3&gt;
&lt;p&gt;The enemy of fascism is nuance. Fascism presents a good, clean, rousing story about a spiritual conflict between right and wrong. It is anathema to fascism to stop and muddy the waters a bit: in this case, to explore the ways in which LLMs, like any transformative technology, can both support and endanger traditional values.&lt;/p&gt;
&lt;p&gt;In &lt;a href=&quot;/the-left-wing-case-for-AI&quot;&gt;&lt;em&gt;The left-wing case for AI&lt;/em&gt;&lt;/a&gt; I wrote about how AI is being used &lt;em&gt;right now&lt;/em&gt; as a disability aid, and many disabled readers wrote in to share their positive experiences with LLMs, and often how alienated they feel by the anti-AI mainstream on the left. I recently got an email describing how there’s a sudden flood of accessibility software for blind people&lt;sup id=&quot;fnref-9&quot;&gt;&lt;a href=&quot;#fn-9&quot; class=&quot;footnote-ref&quot;&gt;9&lt;/a&gt;&lt;/sup&gt; that’s &lt;em&gt;actually built by blind people&lt;/em&gt;, who can now iterate with a LLM to get a product that meets their needs. Framing AI as an ontological evil erases experiences like these.&lt;/p&gt;
&lt;p&gt;Being anti-AI is not inherently fascist. Many of the anti-AI posts I’ve quoted are thoughtful, sensitive pieces exploring how the author thinks about one of the biggest changes to our industry. I still think the world needs more articles like that, not less, but the more of them I read, the more I recognize the tropes: spiritually pure lovers of the craft, degenerate peddlers of corrupt modernism, a need to return to the traditional ways of the hacker, and a lament for the (potentially) waning power of an elite fraternity of programmers.&lt;/p&gt;
&lt;h3&gt;Conclusion&lt;/h3&gt;
&lt;p&gt;I know I’m tiptoeing around &lt;a href=&quot;https://www.lesswrong.com/posts/yCWPkLi8wJvewPbEp/the-noncentral-fallacy-the-worst-argument-in-the-world&quot;&gt;the worst argument in the world&lt;/a&gt;. It isn’t a refutation of anti-LLM arguments to say that they are structurally similar in some ways to fascist arguments, any more than it’s a devastating critique to say the same thing about Lord of the Rings. Sometimes it is good to try and halt the march of progress! Some of our past traditions really were purer and more spiritually robust! It just bothers me, that’s all.&lt;/p&gt;
&lt;p&gt;I used to read &lt;a href=&quot;https://users.cs.utah.edu/~elb/folklore/mel.html&quot;&gt;The Story of Mel&lt;/a&gt; with unalloyed pleasure. Now it makes me nervous. If you believe you’re fighting &lt;a href=&quot;https://tante.cc/2026/04/21/ai-as-a-fascist-artifact/&quot;&gt;the embodiment of fascism&lt;/a&gt;, or for &lt;a href=&quot;https://sinclairtarget.com/blog/2026/06/01/quality-in-the-age-of-slop/&quot;&gt;the idea of value itself&lt;/a&gt;, what &lt;a href=&quot;https://www.theguardian.com/technology/2026/apr/18/sam-altman-house-attack-ai&quot;&gt;tactics&lt;/a&gt; are off-limits? What positions might you eventually come to accept?&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;It feels wrong to directly associate my caricature with any actual posts, but it also feels wrong to make a blanket assertion without examples. Just so you know what I’m talking about, &lt;a href=&quot;https://huronbikes.mataroa.blog/blog/i-am-not-a-software-engineer/&quot;&gt;here&lt;/a&gt; &lt;a href=&quot;https://lpcvoid.com/blog/0018_why_i_am_against_genai/index.html&quot;&gt;are&lt;/a&gt; &lt;a href=&quot;https://alextardif.com/AI.html&quot;&gt;some&lt;/a&gt; &lt;a href=&quot;https://sinclairtarget.com/blog/2026/06/01/quality-in-the-age-of-slop/&quot;&gt;posts&lt;/a&gt; that have elements of this attitude. I like some of these posts and dislike others.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;Page 329 of my copy of Julius Evola’s &lt;em&gt;Revolt Against the Modern World&lt;/em&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;Ezra Pound, Canto XLV. “Usura” should be read as “usury”, or today we could gloss it as “capitalism”: all Pound’s examples of great art were from the pre-capitalist patronage era of art.&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-4&quot;&gt;
&lt;p&gt;Adolf Hitler, from his speech at the 1933 Party Congress in Nuremberg.&lt;/p&gt;
&lt;a href=&quot;#fnref-4&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-5&quot;&gt;
&lt;p&gt;Of course, there’s also historically been a strong &lt;em&gt;pro&lt;/em&gt;-technology current in fascist thinking (even specificially &lt;em&gt;Italian&lt;/em&gt; fascist &lt;a href=&quot;https://artmejo.com/how-italian-futurism-influenced-the-rise-of-fascism/&quot;&gt;thinking&lt;/a&gt;).&lt;/p&gt;
&lt;a href=&quot;#fnref-5&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-6&quot;&gt;
&lt;p&gt;Page 134 of &lt;em&gt;Blood in the Machine&lt;/em&gt; has a brief argument that Luddism was feminist because the (exclusively male) artisans’ wives would provide food for their meetings. No, really.&lt;/p&gt;
&lt;a href=&quot;#fnref-6&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-7&quot;&gt;
&lt;p&gt;From Kevin Binfield’s &lt;em&gt;Writings of the Luddites&lt;/em&gt;, page 40. I’ve taken the liberty of re-rendering it in modern spelling and grammar.&lt;/p&gt;
&lt;a href=&quot;#fnref-7&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-8&quot;&gt;
&lt;p&gt;Aside from being too early, they didn’t have any connection to the state apparatus of power (in fact, they were ultimately crushed by it) and they famously lacked a singular leader.&lt;/p&gt;
&lt;a href=&quot;#fnref-8&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-9&quot;&gt;
&lt;p&gt;The example cited was &lt;a href=&quot;https://github.com/serrebidev/BlindRSS&quot;&gt;BlindRSS&lt;/a&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-9&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[Weird projects I shipped with AI]]></title><link>https://seangoedecke.com/weird-projects-i-shipped-with-ai/</link><guid isPermaLink="false">https://seangoedecke.com/weird-projects-i-shipped-with-ai/</guid><pubDate>Mon, 01 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Where are all the AI-generated projects? This is a &lt;a href=&quot;https://news.ycombinator.com/item?id=46262545&quot;&gt;common question&lt;/a&gt; from AI skeptics: if LLMs are so good at writing code, where is the tsunami of new AI-generated apps, services and games?&lt;/p&gt;
&lt;p&gt;I personally don’t find this to be much of a paradox. Writing code is only one of the bottlenecks involved in actually &lt;a href=&quot;/how-to-ship&quot;&gt;shipping&lt;/a&gt; a new product, after all. It’s also impossible to talk about the paid work I’ve done with AI (you’ll simply have to take my word that it’s increased my productivity). But one thing I can do is share a list of personal projects I’ve built with AI in the last twelve months.&lt;/p&gt;
&lt;p&gt;I definitely would not have done &lt;em&gt;all&lt;/em&gt; of these by hand. I might have found the time to do one or two of them, but based on my pre-AI track record they would probably have stayed in the “GitHub repo with a few commits” stage. This list is a kind of &lt;a href=&quot;https://sites.millersville.edu/bikenaga/math-proof/existence-proofs/existence-proofs.html&quot;&gt;existence proof&lt;/a&gt;: a bunch of weird projects, useful to at least some people, that would not have existed without AI assistance&lt;sup id=&quot;fnref-0&quot;&gt;&lt;a href=&quot;#fn-0&quot; class=&quot;footnote-ref&quot;&gt;0&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;h3&gt;Skifreedle&lt;/h3&gt;
&lt;p&gt;Most recently I’ve built &lt;a href=&quot;https://skifreedle.com/&quot;&gt;skifreedle.com&lt;/a&gt;, a daily-game version of the classic Windows SkiFree &lt;a href=&quot;http://ski.ihoc.net/&quot;&gt;game&lt;/a&gt; (i.e. “like Wordle, but for SkiFree”). The code for that is &lt;a href=&quot;https://github.com/sgoedecke/skifreedle&quot;&gt;here&lt;/a&gt;&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;. &lt;/p&gt;
&lt;p&gt;&lt;span
      class=&quot;gatsby-resp-image-wrapper&quot;
      style=&quot;position: relative; display: block; margin-left: auto; margin-right: auto; max-width: 590px; &quot;
    &gt;
      &lt;a
    class=&quot;gatsby-resp-image-link&quot;
    href=&quot;/static/6cdb3210b3e29b8b9f8ba68fee4d2502/105d8/skifreedle1.png&quot;
    style=&quot;display: block&quot;
    target=&quot;_blank&quot;
    rel=&quot;noopener&quot;
  &gt;
    &lt;span
    class=&quot;gatsby-resp-image-background-image&quot;
    style=&quot;padding-bottom: 114.86486486486487%; position: relative; bottom: 0; left: 0; background-image: url(&apos;data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAAXCAYAAAALHW+jAAAACXBIWXMAABYlAAAWJQFJUiTwAAACpElEQVQ4y5VUy2sTQRjPzT+kN8G/wJsn8aIHb54FQRAUBO+CPfpAD70p+CDUB20pptDXQVCqNWlsmuwr+2qSptluEvLa7M7sz9nZbJpNS9gOfMw388385je/75tJGT+zEN8tQv2+jmqlggoz2z5hfQ1LGwpWtmRougHTNGBZFk6YHQqrqApp1EwBJ3YLrVYTsqJA0StIbd99jN25q1i7dhsry8tYYlY82Mf65g/MXf+AKzc/4lP6Kz4vppHL5VHYz2PlzWVkXl3CdmYBxVIZpVIR2exfLK9mkNKNOta+rSK/swNd16FpGhRZhiSKyOVLyO4V2ViBqqqQR/OK+A+qlGe+AEmSoLHYxuYWCpKB1BCAUj2GSzzMar4/M8xkOYTV8ZDqMpyipKLTbvMAJWRsZNS7QxeOM4zFKCPA13iE7xPlMo6afaR6bFySdXQ7nZAJpTELqLmuhyEDjcaTMUoo3ycpKuqtwfmAp5tY71MO5jhxwKhPCDg6nYbX8Zks/YFzChQIGhzIGZJkDCOwYvcARl+H26cg/lmGUZsJGBihLh8/EO7jtfmc+31WxBGzCKxj12CbBYiSzACd2UkJ2p3sDcyLT7j/J7MLq2qF6wjl+tr1NmoFG4JYns0w0DGQqt0x0e03eOzIaqDX68XXTV65OTPLcY3O+H5YAQmTQieAybn1F81TeqE6PCerU9lNxHD6pcQKfQR2PKihPWwyRCRjGBqN1WVwTULDT+Sp8BBvtRfJ63BSq6jIwzjz2Vktp4Gum5BhcLVprcyeBmtwHOoWHeBfUMNgg9Er8/l54RHeGy/HQNGaRFmOwAj78+7t3UK+9Zv5Qz63UV/CQvnZiKmXnGGYCPa0WDYdr8vBgrbX/IXN+pcxy8S/zTjbE1ek1DvzWqb/w/+ANdSk/ZHjHwAAAABJRU5ErkJggg==&apos;); background-size: cover; display: block;&quot;
  &gt;&lt;/span&gt;
  &lt;img
        class=&quot;gatsby-resp-image-image&quot;
        alt=&quot;skifreedle&quot;
        title=&quot;skifreedle&quot;
        src=&quot;/static/6cdb3210b3e29b8b9f8ba68fee4d2502/fcda8/skifreedle1.png&quot;
        srcset=&quot;/static/6cdb3210b3e29b8b9f8ba68fee4d2502/12f09/skifreedle1.png 148w,
/static/6cdb3210b3e29b8b9f8ba68fee4d2502/e4a3f/skifreedle1.png 295w,
/static/6cdb3210b3e29b8b9f8ba68fee4d2502/fcda8/skifreedle1.png 590w,
/static/6cdb3210b3e29b8b9f8ba68fee4d2502/efc66/skifreedle1.png 885w,
/static/6cdb3210b3e29b8b9f8ba68fee4d2502/105d8/skifreedle1.png 1170w&quot;
        sizes=&quot;(max-width: 590px) 100vw, 590px&quot;
        style=&quot;width:100%;height:100%;margin:0;vertical-align:middle;position:absolute;top:0;left:0;&quot;
        loading=&quot;lazy&quot;
      /&gt;
  &lt;/a&gt;
    &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span
      class=&quot;gatsby-resp-image-wrapper&quot;
      style=&quot;position: relative; display: block; margin-left: auto; margin-right: auto; max-width: 590px; &quot;
    &gt;
      &lt;a
    class=&quot;gatsby-resp-image-link&quot;
    href=&quot;/static/6ee543da4626f592eb2537653f7d2b18/a13c9/skifreedle2.png&quot;
    style=&quot;display: block&quot;
    target=&quot;_blank&quot;
    rel=&quot;noopener&quot;
  &gt;
    &lt;span
    class=&quot;gatsby-resp-image-background-image&quot;
    style=&quot;padding-bottom: 127.02702702702702%; position: relative; bottom: 0; left: 0; background-image: url(&apos;data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAAZCAYAAAAxFw7TAAAACXBIWXMAABYlAAAWJQFJUiTwAAAFL0lEQVQ4y5WVSWwbZRTHfUJCYimtOEFBaYVQD+wHkDhx4IJAlKpQLlxKqZIKqUVlq7gUKoQQqDQt6kIV1DRN48ZZmtVOxksS23HcLE4cL2OPZ7zvSewk3rc/b8ZOF0oOjPT3+8bv+37zvuW9TxYYGESs7Rp4si4PC7vNinBIwILVhdYOC7pHluB22eBwLCORTCDos4E1tcIz/zfcbhv1X8ZqIgFB4BAMByFLBsMwTQ/DbFUhEY8hGgkiHo3Azi9Cu6QDLwQQi4Xo/xAi5IuE/YiGWESCLJLRGKz+OQw7FGBdLFLracjylTLCK1Fki2lUygWUSzlSntp51KqiitJ7qZhHsZCr9ymXSEWgUoFv3YXJYD+i8SjyxSJk2XyOIoiiRLZEMHFgqZiTVCwQpJi9C8vnMw1/tt5H/AB9rJzLIBKloIghy9KgWDyM7MYqcusbyGXXkNtI1yW219eRy6SQ2VzDxlrqYf/GOvlSiCfjBMxDVihVYXcYcfVyO+RX5OjS/4zr54bRrfgTnePn0fnbKG5pfsH1mzfQeaEbcsMZdFwYRPfNy+hkzqLrDxXO3ziN8UklKtUaZEVaQ8v8HJb6lBDUOnCMBpxyAh6GATc0Bm6YgUelAjeoBqciMeqGXw33mAZ+jRG69nZ0tF9DDaAI6ceqYuB9bB8Su15DXNRTLyHx3NtIH/kG6eZTWHv/c6RbTiH12QmkDrYg8eybiO94GbFdryL19OuYfWQvui5eQbUOrGFpcgr2Dw7Dc/ArsAeOw/HhcbAfn4Tn+3PwfHsWXMsZ8D9cgOfEr+COnobrk5NwfXRc6us79DUM736Km1fb6sA87ZrDxYEPWGE3HIXf0oKk8xiClmbwhsMQjCTTkXqbrDDzBYTZZghzLfDNN8NpOoapGQ16+gdQqVVFYI6AXsR4HRzyR2HpeAIzbY9jhdkBOHcBNtLyTrINLTdk3Sn58/onoVXeQN/ACAEr9QidbgE+txGDrbvR9uNutH73DG6fex7LPXuwIG+CrfcFLJMWu/fAqtiLxVtNsJCsiiborzVBM6Yg4PAWsCCljc/npRTUwO2cg9djgc1qxsK8GcvWOShHh6DVjGHGNIXJCQaWBTPpDpYWZ2nMFPR6PXr7+7emXICDdSIQ8EuAkZFBeDgnLW8F1UoO4jM7a4LRMEH+GQi8E5sbSSktgRLlfwgMM46+27fvAZ0UYSgUhMEwCa2WkazROIVpkpHaY6oRMONK6LTjFKEGBr0O83MzJDNMJgNFOU0R9jWApSJlih3ZTIZyvUr5mkcqlQbHeag8cfB6vfAF4uC9EfiDCXiEMAJkU+tFZHMlaQZ+vx+3FN10bGr3gBkCik+UkpzneZo2h5W1VYrgDiJzX8Knew+BqQMQdPvhYvaDG38HC1MXpTFeL/8fwM1NycmyLHp7e9HT04NR2oxLl/6CVn4I8t9fpLx+BYrWN6A4/xYUZ/dhoOsnaYwg8NtHWCqVkM1mqd6VqaBGYDaboZmYwYR+HsNKOqusX+qXoT0pV6QmnRDfw8BsA3j/UygUaIMM0E9qaeH1GBoaoCUJN7w11Gq17YGbjSlXq3RcqlVpg2q0a4VCEbFkBsnVjFSpRUilstWn0lhD74NAGwHFKW738JEaQivbuqWNlHc3gAVaMxGoVI5CRXVPpVLelVJZt2pGRedQdff9fr9arUY71UMRKNXDHF0BkViEjgkL1mmDi7U/IFaUs65/+7bkpCtW8AkQk0Qm3gMxurHE2w1SRav8T9GYWgnxRP1O+Qe4Rrb0JcRZEQAAAABJRU5ErkJggg==&apos;); background-size: cover; display: block;&quot;
  &gt;&lt;/span&gt;
  &lt;img
        class=&quot;gatsby-resp-image-image&quot;
        alt=&quot;skifreedle2&quot;
        title=&quot;skifreedle2&quot;
        src=&quot;/static/6ee543da4626f592eb2537653f7d2b18/fcda8/skifreedle2.png&quot;
        srcset=&quot;/static/6ee543da4626f592eb2537653f7d2b18/12f09/skifreedle2.png 148w,
/static/6ee543da4626f592eb2537653f7d2b18/e4a3f/skifreedle2.png 295w,
/static/6ee543da4626f592eb2537653f7d2b18/fcda8/skifreedle2.png 590w,
/static/6ee543da4626f592eb2537653f7d2b18/efc66/skifreedle2.png 885w,
/static/6ee543da4626f592eb2537653f7d2b18/a13c9/skifreedle2.png 1178w&quot;
        sizes=&quot;(max-width: 590px) 100vw, 590px&quot;
        style=&quot;width:100%;height:100%;margin:0;vertical-align:middle;position:absolute;top:0;left:0;&quot;
        loading=&quot;lazy&quot;
      /&gt;
  &lt;/a&gt;
    &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;I enjoy coding small web games by hand, but &lt;em&gt;definitely&lt;/em&gt; would not have had the time to wire up all the different SkiFree objects or build neat features like a ghost of your fastest run. I also tried out a lot of different visual themes for the game UI before landing on something I liked. If I’d done this by hand, I would have only had time to try out two or three different looks, instead of fifteen or twenty.&lt;/p&gt;
&lt;p&gt;I’m very happy with how this turned out. I’ve been enjoying competing against my brother to get better times, since both of us have a lot of nostalgia for the original SkiFree game.&lt;/p&gt;
&lt;h3&gt;Autodeck&lt;/h3&gt;
&lt;p&gt;Last year I built &lt;a href=&quot;https://www.autodeck.pro/&quot;&gt;Autodeck&lt;/a&gt;! I wrote a &lt;a href=&quot;/autodeck/&quot;&gt;blog post&lt;/a&gt; about this before, but this came from my partner wishing there was some way to automatically generate Anki cards about random topics she wanted to learn about. It ended up being relatively straightforward to set up an endless feed of auto-generated spaced repetition cards:&lt;/p&gt;
&lt;p&gt;&lt;span
      class=&quot;gatsby-resp-image-wrapper&quot;
      style=&quot;position: relative; display: block; margin-left: auto; margin-right: auto; max-width: 590px; &quot;
    &gt;
      &lt;a
    class=&quot;gatsby-resp-image-link&quot;
    href=&quot;/static/ab189b44f0c389b77ca4df74dcf8259b/71c1d/autodeck.png&quot;
    style=&quot;display: block&quot;
    target=&quot;_blank&quot;
    rel=&quot;noopener&quot;
  &gt;
    &lt;span
    class=&quot;gatsby-resp-image-background-image&quot;
    style=&quot;padding-bottom: 70.27027027027026%; position: relative; bottom: 0; left: 0; background-image: url(&apos;data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAAOCAYAAAAvxDzwAAAACXBIWXMAABYlAAAWJQFJUiTwAAABdklEQVQ4y41T227CMAzl/39r2tM07WkTu7FxhwlG0zZpkib1fFygBBWGpVOniX1y6tqDqtLknBV43/l07RIc4k9RVYaMKWlgraGmIYr8qEOgECL7KOsYea8OnOCZqG7B732Gi4QQzLCi1KTyglHSZpvRdqfY72iX5VRZl6Dhy08Bg8o9oRZ1UBZiC6iCh6LI/pLhDABnQgiCvDRU6orywrBaw8oK8ZkqZS9TBXst+5tfJfvW1QJfx5QQ6h6fnunu/oFehp+SMHwb0fdkQcP3L5ovf+iV/WyxpjHvYf0xmvDne7I+SF1ThUyIuo2nCypYJQKdxw9quBzE553nUFmjEojpVYgbrAs0mS1pOl/J7av1VoJFhfj6+ImV69a9hKghEpB8xD4YKq7hisK6N/gWSPx5Df9TcgAu75q/BQYgITxv1EuAoTxodmkpHgK0E9oNk6J10U3KrQZe63wHrnc7nq5TaAzGTpFSGUNdQXue53kCzfktT0l/h7tC4XKA5C0AAAAASUVORK5CYII=&apos;); background-size: cover; display: block;&quot;
  &gt;&lt;/span&gt;
  &lt;img
        class=&quot;gatsby-resp-image-image&quot;
        alt=&quot;autodeck&quot;
        title=&quot;autodeck&quot;
        src=&quot;/static/ab189b44f0c389b77ca4df74dcf8259b/fcda8/autodeck.png&quot;
        srcset=&quot;/static/ab189b44f0c389b77ca4df74dcf8259b/12f09/autodeck.png 148w,
/static/ab189b44f0c389b77ca4df74dcf8259b/e4a3f/autodeck.png 295w,
/static/ab189b44f0c389b77ca4df74dcf8259b/fcda8/autodeck.png 590w,
/static/ab189b44f0c389b77ca4df74dcf8259b/efc66/autodeck.png 885w,
/static/ab189b44f0c389b77ca4df74dcf8259b/c83ae/autodeck.png 1180w,
/static/ab189b44f0c389b77ca4df74dcf8259b/71c1d/autodeck.png 1536w&quot;
        sizes=&quot;(max-width: 590px) 100vw, 590px&quot;
        style=&quot;width:100%;height:100%;margin:0;vertical-align:middle;position:absolute;top:0;left:0;&quot;
        loading=&quot;lazy&quot;
      /&gt;
  &lt;/a&gt;
    &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;I set up Stripe payments for this one, more because I was worried about someone running away with my Groq balance than because I wanted to make money, but I was pleasantly surprised to see a bunch of people actually use this. Over five hundred people have tried it out, with enough paid subscribers to cover inference and hosting.&lt;/p&gt;
&lt;p&gt;I &lt;em&gt;might&lt;/em&gt; have built this without LLM assistance, but I almost certainly would not have &lt;em&gt;deployed&lt;/em&gt; it as a website. The hassle of setting up a database and Stripe would have just been too much work.&lt;/p&gt;
&lt;h3&gt;Endless Wiki&lt;/h3&gt;
&lt;p&gt;I also built an AI-generated &lt;a href=&quot;https://www.endlesswiki.com/&quot;&gt;endless wiki&lt;/a&gt;. I wrote a &lt;a href=&quot;https://www.seangoedecke.com/endless-wiki/&quot;&gt;blog post&lt;/a&gt; about this one as well. Like Autodeck, I was fascinated with the idea of non-chat interfaces for LLMs, and I thought a wiki-based approach where you interact with the model by clicking links was pretty cool.&lt;/p&gt;
&lt;p&gt;&lt;span
      class=&quot;gatsby-resp-image-wrapper&quot;
      style=&quot;position: relative; display: block; margin-left: auto; margin-right: auto; max-width: 590px; &quot;
    &gt;
      &lt;a
    class=&quot;gatsby-resp-image-link&quot;
    href=&quot;/static/50f19e416dd78a54917fe8f645449839/6578c/ewiki.png&quot;
    style=&quot;display: block&quot;
    target=&quot;_blank&quot;
    rel=&quot;noopener&quot;
  &gt;
    &lt;span
    class=&quot;gatsby-resp-image-background-image&quot;
    style=&quot;padding-bottom: 39.189189189189186%; position: relative; bottom: 0; left: 0; background-image: url(&apos;data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAAICAYAAAD5nd/tAAAACXBIWXMAABYlAAAWJQFJUiTwAAABUElEQVQoz01RCZKDMAzj/6/YX+xjdtpyE5JAUsJVLq3tFnY1I2wC2JKIwvjCQw2omgnTNCP0A/phxDBOUpnrtoGx74fwOIDXssD5DvM04CcL+PouMMwbIt8FjNOEE89nh7wooLVGWZZomgZK1bBUGceH/8FL1nWjRQeiXPcozXA97EKPQrUo9ROV9qitQ13XQmutLHDOoW3biyGEa01k/IxMj1jWXQ68d0iShNRVpJRZIs9zUZumqfRZXiKhnt/js1obESIK+bJt72wYjQt4pAppYZFlOX0QU80+PQ8skJYWLeV3KjWNo6xnURn9JXFcliuloWqipvyMJ/tO7HOunKdSlVif5xnjOFKdpO77/lZ4kjH0PX1YC41poG0LbUiFsTRIycAzUx4qJJUL/XXJ8NL3Gcg2brcb7vc7ZRSTzQRxHIvd06L3XtjSPZ9xz+oYv3aLYoFdwXwIAAAAAElFTkSuQmCC&apos;); background-size: cover; display: block;&quot;
  &gt;&lt;/span&gt;
  &lt;img
        class=&quot;gatsby-resp-image-image&quot;
        alt=&quot;endlesswiki&quot;
        title=&quot;endlesswiki&quot;
        src=&quot;/static/50f19e416dd78a54917fe8f645449839/fcda8/ewiki.png&quot;
        srcset=&quot;/static/50f19e416dd78a54917fe8f645449839/12f09/ewiki.png 148w,
/static/50f19e416dd78a54917fe8f645449839/e4a3f/ewiki.png 295w,
/static/50f19e416dd78a54917fe8f645449839/fcda8/ewiki.png 590w,
/static/50f19e416dd78a54917fe8f645449839/efc66/ewiki.png 885w,
/static/50f19e416dd78a54917fe8f645449839/c83ae/ewiki.png 1180w,
/static/50f19e416dd78a54917fe8f645449839/6578c/ewiki.png 2242w&quot;
        sizes=&quot;(max-width: 590px) 100vw, 590px&quot;
        style=&quot;width:100%;height:100%;margin:0;vertical-align:middle;position:absolute;top:0;left:0;&quot;
        loading=&quot;lazy&quot;
      /&gt;
  &lt;/a&gt;
    &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;I learned the hard way that putting a LLM generation call on the end of a regular link was a bad idea: scrapers would exhaust my inference budget quickly. I ended up faking the no-article-exists-yet links with JavaScript, which at least so far has defeated scrapers. People still email me about Endless Wiki, and there are over 280 thousand pages generated.&lt;/p&gt;
&lt;p&gt;My original goal was to see if you could eventually generate a page for Neon Genesis Evangelion, starting at the root page and only following links (kind of like &lt;a href=&quot;https://dev.to/zmbailey/wikigolf-an-automated-traversal-of-wikipedia-9o0&quot;&gt;wiki golf&lt;/a&gt;). I was successful! You can read the “Evangelion Anime” page &lt;a href=&quot;https://www.endlesswiki.com/wiki/evangelion_anime&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Almost exactly a month after I launched Endless Wiki, xAI launched &lt;a href=&quot;https://en.wikipedia.org/wiki/Grokipedia&quot;&gt;Grokipedia&lt;/a&gt;. Obviously they didn’t plagiarize me. This is a very easy idea to have, and my site was not the first infinite wiki (though I think it was the first one where you had to discover new pages by clicking on links). But it did take some of the shine off.&lt;/p&gt;
&lt;h3&gt;VicFlora Offline&lt;/h3&gt;
&lt;p&gt;I built a &lt;a href=&quot;https://vicfloraoffline.netlify.app/&quot;&gt;PWA&lt;/a&gt; that caches the VicFlora plant identification database so it could be used with low or no internet. This was more of a utility project for my partner, who likes plants and occasionally goes on field trips where internet is spotty.&lt;/p&gt;
&lt;p&gt;I would definitely not have done this without LLMs. It was reasonably difficult to scrape the basic dichotomous key from the VicFlora website: their API documentation was out of date, there were multiple possible pathways for fetching data (most of which were not functional), and the format of the data I did manage to fetch was hard to parse. I think I &lt;em&gt;could&lt;/em&gt; have done it, with enough effort, but it would have been a substantial amount of work.&lt;/p&gt;
&lt;p&gt;I’m very happy with how this turned out. It’s not perfect, but it’s functional, and I’ve even had the occasional Victorian botanist email me with bug reports or feature requests, so it’s clearly seeing a little bit of usage.&lt;/p&gt;
&lt;h3&gt;Other projects&lt;/h3&gt;
&lt;p&gt;I did a bunch of other stuff that doesn’t necessarily rise to the level of a “deployed project”: my &lt;a href=&quot;https://github.com/sgoedecke/gh-standup&quot;&gt;gh-standup&lt;/a&gt; GitHub CLI extension to automatically generate a standup report, which has just over a hundred stars, my (low quality) image geolocation &lt;a href=&quot;https://github.com/sgoedecke/ai_geolocation&quot;&gt;benchmark&lt;/a&gt;, which I blogged about &lt;a href=&quot;/the-o3-geoguessr-prompt-did-not-work/&quot;&gt;here&lt;/a&gt;, or my &lt;a href=&quot;https://github.com/sgoedecke/skills/blob/main/skills/extract-features-clamp-inference/SKILL.md&quot;&gt;skill&lt;/a&gt; for extracting features from open-source models.&lt;/p&gt;
&lt;p&gt;There may not be a flood of AI-generated companies (yet), but at least for me there’s been a flood of small, weird projects that would not have existed without significant LLM assistance.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-0&quot;&gt;
&lt;p&gt;I also want to shout out Simon Willison’s &lt;a href=&quot;https://simonwillison.net/2025/Sep/4/highlighted-tools/&quot;&gt;version of this&lt;/a&gt;, which is another great example of “weird useful tools that only exist because the cost of creating them was so low”.&lt;/p&gt;
&lt;a href=&quot;#fnref-0&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;I did lift the spritesheet from DanielHough’s &lt;a href=&quot;https://github.com/basicallydan/skifree.js&quot;&gt;SkiFree.js&lt;/a&gt;, which attributes it to &lt;a href=&quot;http://spriters-resource.com/submitter/Wing%20Wang%20Wao&quot;&gt;Wing Wang Wao&lt;/a&gt;. Of course, the original sprites and art belong to Chris Pirih’s SkiFree and Microsoft.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[Build agents, not pipelines]]></title><link>https://seangoedecke.com/build-agents-not-pipelines/</link><guid isPermaLink="false">https://seangoedecke.com/build-agents-not-pipelines/</guid><pubDate>Sun, 31 May 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;There are only two ways to use LLMs in a computer program: as part of a pipeline, or as an agent. In other words, either you express the control flow of the program in code, or you give a LLM tools and allow it to manage the control flow itself&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;Here’s how you might structure a trivial “summarize a bunch of information and email it to me” program as a pipeline:&lt;/p&gt;
&lt;div class=&quot;gatsby-highlight&quot; data-language=&quot;ruby&quot;&gt;&lt;pre class=&quot;language-ruby&quot;&gt;&lt;code class=&quot;language-ruby&quot;&gt;context &lt;span class=&quot;token operator&quot;&gt;=&lt;/span&gt; gather_context&lt;span class=&quot;token punctuation&quot;&gt;(&lt;/span&gt;various&lt;span class=&quot;token punctuation&quot;&gt;,&lt;/span&gt; data&lt;span class=&quot;token punctuation&quot;&gt;,&lt;/span&gt; sources&lt;span class=&quot;token punctuation&quot;&gt;)&lt;/span&gt;
llm_response &lt;span class=&quot;token operator&quot;&gt;=&lt;/span&gt; llm_summarize&lt;span class=&quot;token punctuation&quot;&gt;(&lt;/span&gt;context&lt;span class=&quot;token punctuation&quot;&gt;)&lt;/span&gt;
summary &lt;span class=&quot;token operator&quot;&gt;=&lt;/span&gt; parse&lt;span class=&quot;token punctuation&quot;&gt;(&lt;/span&gt;llm_response&lt;span class=&quot;token punctuation&quot;&gt;)&lt;/span&gt;
email_me&lt;span class=&quot;token punctuation&quot;&gt;(&lt;/span&gt;summary&lt;span class=&quot;token punctuation&quot;&gt;,&lt;/span&gt; my_email&lt;span class=&quot;token punctuation&quot;&gt;)&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;And here’s how you’d do it as an agent:&lt;/p&gt;
&lt;div class=&quot;gatsby-highlight&quot; data-language=&quot;ruby&quot;&gt;&lt;pre class=&quot;language-ruby&quot;&gt;&lt;code class=&quot;language-ruby&quot;&gt;read_data_tool &lt;span class=&quot;token operator&quot;&gt;=&lt;/span&gt; build_read_data_tool&lt;span class=&quot;token punctuation&quot;&gt;(&lt;/span&gt;various&lt;span class=&quot;token punctuation&quot;&gt;,&lt;/span&gt; data&lt;span class=&quot;token punctuation&quot;&gt;,&lt;/span&gt; sources&lt;span class=&quot;token punctuation&quot;&gt;)&lt;/span&gt;
email_tool &lt;span class=&quot;token operator&quot;&gt;=&lt;/span&gt; build_email_tool&lt;span class=&quot;token punctuation&quot;&gt;(&lt;/span&gt;my_email&lt;span class=&quot;token punctuation&quot;&gt;)&lt;/span&gt;
run_agent&lt;span class=&quot;token punctuation&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;token symbol&quot;&gt;tools&lt;/span&gt;&lt;span class=&quot;token operator&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;token punctuation&quot;&gt;[&lt;/span&gt;read_data_tool&lt;span class=&quot;token punctuation&quot;&gt;,&lt;/span&gt; email_tool&lt;span class=&quot;token punctuation&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;)&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;It’s like the difference &lt;a href=&quot;https://news.ycombinator.com/item?id=46375199&quot;&gt;between&lt;/a&gt; a library and a framework. When you use a library, you define the structure of the program yourself, and call out to various library helpers along the way. When you use a framework, the main structure of the program lives in the framework, and it calls your code at various points. There are tradeoffs involved in both approaches. Frameworks let you get started more quickly and typically give you features “for free”, but can be difficult when you want to do something that isn’t part of the framework’s design. Libraries give you a lot more control, but require you to write (and maintain) more boilerplate code.&lt;/p&gt;
&lt;p&gt;In the trivial case, the distinction between a pipeline and an agent melts away. If you only have a few paragraphs of possible context for the problem, an agent with a &lt;code class=&quot;language-text&quot;&gt;gather_context&lt;/code&gt; and an &lt;code class=&quot;language-text&quot;&gt;email_me&lt;/code&gt; tool will perform exactly the same steps as a pipeline that calls a reasoning model with the context injected into the prompt (i.e. the agent will reproduce the trivial control flow of your pipeline). But when you have more context than will fit into a single prompt, or you want to take an action and then react to the result, the choice between pipelines and agents becomes very significant.&lt;/p&gt;
&lt;h3&gt;Predictability, flexibility and intelligence&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Pipelines are more predictable, but agents are more flexible&lt;/strong&gt;. When you give a problem to an agent, work stops when the LLM thinks it’s done. Depending on the perceived difficulty of the problem, this can take anywhere from a few LLM turns to hundreds (and thus cost anywhere from a few cents to many dollars). If you’re building something intended to run at scale, this unpredictability can be a nightmare. Any subtle change to the user data could cause the LLM to take twice as long on each task, which would double your latency&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt; and cost.&lt;/p&gt;
&lt;p&gt;Pipelines are only immune to this problem if they don’t use reasoning models, or don’t allow the model to “think out loud” in its output tokens (for instance, by using &lt;a href=&quot;https://developers.openai.com/api/docs/guides/structured-outputs&quot;&gt;structured output&lt;/a&gt;). However, individual LLMs offer much tighter control over model reasoning than over how long an agentic loop will take. In all frontier model APIs, you can explicitly set the level of reasoning you want. That doesn’t give you total control, but it does cap “take longer” at maybe ten or twenty percent (instead of with agents, where it can be 2x or more).&lt;/p&gt;
&lt;p&gt;Why use agents, then? &lt;strong&gt;Agents are smarter&lt;/strong&gt;. If you’re happy to accept the unpredictability, an agentic system can handle &lt;em&gt;much&lt;/em&gt; more difficult tasks, by virtue of being able to loop for longer, and to gather more information after thinking about the problem. There’s a reason that the most successful AI products (coding agents like Claude Code, Codex, Cursor, and Copilot&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;) are agents: coding is a hard enough task that you simply cannot build a functional coding agent with pipelines.&lt;/p&gt;
&lt;h3&gt;Context-gathering&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The context-gathering stage is far more delicate for pipelines than for agents&lt;/strong&gt;. If an agent is trying to solve a problem and realizes it needs more data, it can simply go and get it. But for a pipeline, all the required data has to be present in the context already, because the LLM only gets to run once.&lt;/p&gt;
&lt;p&gt;Much of the work involved in building pipelines is in getting context-gathering right. &lt;strong&gt;Agents are much easier.&lt;/strong&gt; For instance, with a coding agent, you can basically just provide a “grep” and “read file” tool and let the agent figure out what chunks of code are relevant to the current file. In a pipeline, you have to figure that out yourself: good luck, it’s an unsolved technical problem! Typically you’ll end up doing some set of clever tricks, like walking the AST to identify which parts of code “contribute” to the current file, or indexing the whole codebase with semantic embeddings and doing some kind of nearest-neighbor search to build the context (called RAG, or “retrieval-augmented generation”). Neither of these will work as well as using an agent.&lt;/p&gt;
&lt;p&gt;In 2023 and 2024, many people believed that RAG would solve context-gathering. Every LLM would have a fully-indexed context base that would magically surface the precise information the LLM needed at any given moment. This did not happen. Instead, we went &lt;em&gt;backwards&lt;/em&gt;, getting our agents to do plain-text search and figure it out like a human would. Why didn’t RAG work? This is a topic for a whole other post, but the short answer is this: “find what information is relevant to this problem” is often as hard a task as &lt;em&gt;actually solving the problem&lt;/em&gt;. Semantic embeddings and cosine similarity are simply not powerful enough tools for the job.&lt;/p&gt;
&lt;h3&gt;Multi-model pipelines&lt;/h3&gt;
&lt;p&gt;Pipelines that make multiple LLM invocations do have an extra dimension of flexibility: they can use different LLMs for different tasks. For instance, if one LLM benchmarks better at task A, or is cheaper for an easier task B, you can use the right model for the job. Agents (at least right now) have to stay the same model the whole time, so you’re always pinned to the highest level of intelligence you need.&lt;/p&gt;
&lt;p&gt;Is this a big deal? I’m suspicious. One pattern I see a lot is tasking a cheaper model with collating or summarizing data for a smarter model to do something with. But often the signal is in the raw data itself! I think designs like this are really shooting themselves in the foot, for the same reasons that RAG didn’t work: context-gathering was a harder problem than people anticipated.&lt;/p&gt;
&lt;p&gt;In any case, if you do want to farm out tasks to different models, you can also do it via careful agentic tool design. For instance, you could build your &lt;code class=&quot;language-text&quot;&gt;web_search&lt;/code&gt; tool so that it uses a cheap model to summarize web pages.&lt;/p&gt;
&lt;h3&gt;Small contexts and future-proofing&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Pipelines allow working with smaller contexts, and thus with local models&lt;/strong&gt;. An agent’s ability to fetch its own context means that it almost always ingests more data than it needs. On top of that, agents run in loops, so each agent turn increases the size of the context. This isn’t a big problem for systems built on top of frontier model APIs, because:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;frontier models all expose large context windows,&lt;/li&gt;
&lt;li&gt;frontier models tend to hold up pretty well for the first 200k tokens, and&lt;/li&gt;
&lt;li&gt;KV caching means that passing around the same large context block is surprisingly cheap.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;However, it is a big problem for local models. The context window consumes &lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1j6xpvt/how_large_is_your_local_llm_context/&quot;&gt;a lot of VRAM&lt;/a&gt;, so most people running local models stay below 32k (or even 6k) tokens. If you’re writing a program to run in this environment, you likely will not be able to give an agent the space it needs, and you will be instead forced to use a pipeline.&lt;/p&gt;
&lt;p&gt;In my opinion, &lt;strong&gt;agents are more future-proof&lt;/strong&gt;. This is partly because models are now being explicitly built to be better agents, and partly because agents delegate more to the LLM and thus benefit more from LLM improvements. If you have a pipeline-based system, new models will probably do a bit better than old ones. If you have an agentic system, new models might do &lt;em&gt;much&lt;/em&gt; better than old ones (to the point that it’s worth building an agentic systems for tasks that are currently too hard, on the assumption that by the time you’ve finished the models may be good enough). I have been banging this drum &lt;a href=&quot;/llm-driven-agents/&quot;&gt;since 2023&lt;/a&gt;, before tool-calling was even a part of model APIs.&lt;/p&gt;
&lt;h3&gt;Safety and legibility&lt;/h3&gt;
&lt;p&gt;In general, I disagree with the &lt;a href=&quot;https://www.decodingai.com/p/stop-building-ai-agents&quot;&gt;popular advice&lt;/a&gt; that workflows are safer than agents. Workflows offer more control &lt;em&gt;over budget&lt;/em&gt;, but when it comes to taking action based on LLM output, you have exactly the same problem whether you’re checking at the tool-call level or at the next stage in the pipeline: either you make some heuristic assessment via code, which might be wrong, or you queue the action up for a human to approve, which will be slow.&lt;/p&gt;
&lt;p&gt;Don’t agents open you up to prompt injection? Yes, but pipelines do too. In both cases, you’re feeding some block of human-generated data (e.g. the files in a codebase, or the results of a web search) into the LLM. Any prompt injections in that data will be consumed by the LLM just the same whether they’re the result of a tool-call or directly injected into the prompt by the pipeline. You have to sanitize user content and double-check LLM-triggered actions, no matter what design you choose&lt;sup id=&quot;fnref-4&quot;&gt;&lt;a href=&quot;#fn-4&quot; class=&quot;footnote-ref&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;I do want to acknowledge that &lt;strong&gt;pipelines are slightly more &lt;em&gt;legible&lt;/em&gt;&lt;/strong&gt;. You can trace most of what a pipeline is doing because you’re in control over more of it. It’s harder to figure out why an agent queried for a particular piece of information or took some action. But even in a pipeline, you’ll never know for sure why the LLM responded in the way it did. That’s just what it means to program with LLMs.&lt;/p&gt;
&lt;h3&gt;LLM-driven mass surveillance&lt;/h3&gt;
&lt;p&gt;Let’s apply some of these principles to a real-world, non-trivial example. Suppose you are the NSA, and you are attempting to use LLMs to get a grip on the wild firehose&lt;sup id=&quot;fnref-5&quot;&gt;&lt;a href=&quot;#fn-5&quot; class=&quot;footnote-ref&quot;&gt;5&lt;/a&gt;&lt;/sup&gt; of covert email surveillance data&lt;sup id=&quot;fnref-6&quot;&gt;&lt;a href=&quot;#fn-6&quot; class=&quot;footnote-ref&quot;&gt;6&lt;/a&gt;&lt;/sup&gt;. Should you use pipelines or agents? Well, if you’re building something that’s supposed to run on every single piece of email in America, you probably shouldn’t use agents: keeping performance and cost strictly bounded requires a pipeline. However, you’re definitely well-resourced enough to use agents &lt;em&gt;in general&lt;/em&gt;, and the problem is definitely hard enough to benefit from the extra intelligence. I’d probably recommend using both: a low-context, cheap pipeline that can run once against each email and flag it, and a fleet of agents that can dig into those flags, make ordinary queries, and act more like human analysts would.&lt;/p&gt;
&lt;p&gt;The pipeline would have to scale with the total volume of data, which should be &lt;em&gt;mostly&lt;/em&gt; fine, since pipelines scale in a predictable-ish manner. The fleet of unpredictable agents can be scaled entirely independently, though in practice it would get bottlenecked on GPU availability and the necessity for human review. The majority of the engineering work&lt;sup id=&quot;fnref-7&quot;&gt;&lt;a href=&quot;#fn-7&quot; class=&quot;footnote-ref&quot;&gt;7&lt;/a&gt;&lt;/sup&gt; would likely go into context-assembly for the pipeline: feeding in enough data about who’s involved in the email conversation so that the LLM can make a sensible decision on whether or not to flag it.&lt;/p&gt;
&lt;h3&gt;Summary&lt;/h3&gt;
&lt;p&gt;Overall, I’d suggest following these guidelines:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Use pipelines when you have strict requirements around context size&lt;/li&gt;
&lt;li&gt;Use pipelines when you need to be able to accurately predict (or limit) GPU cost&lt;/li&gt;
&lt;li&gt;Use pipelines when you have to use local models&lt;/li&gt;
&lt;li&gt;Use agents when you’re not confident you’ll be able to assemble all of the relevant context in one shot&lt;/li&gt;
&lt;li&gt;Use agents when the problem is hard enough that you’re not sure a pipeline will be able to solve it&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;When in doubt, use agents.&lt;/strong&gt; I am aware of several AI projects that have migrated from pipelines to agents in the last year, but none that have gone the other way around. As a general point about software design, if you’re not sure what to do, pick the solution that’s easier to build and more likely to be able to solve your actual problem. If you want to change to a cheaper, pipeline-based system later on, at least you’ll be able to compare it to a working agentic design and make an informed decision.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;This distinction was popularized by Anthropic’s &lt;a href=&quot;https://www.anthropic.com/engineering/building-effective-agents&quot;&gt;&lt;em&gt;Building effective agents&lt;/em&gt;&lt;/a&gt;, written in December 2024, and now (I believe) made at least partially obsolete by advances in agents since then. They say “workflow”, but I slightly prefer the term “pipeline”.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;Yes, I know this is technically not what “latency” means, but there’s no other single-word shorthand for “the duration of a standard unit of work”.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;If you’re building your own coding agent, I suggest you begin with the letter “C”.&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-4&quot;&gt;
&lt;p&gt;For instance, in my trivial example at the top of the post, doesn’t the agent have a failure mode where it might send a ton of emails, or email a bunch of different people? No, because you ought to constrain the email tool so that it can only send to the right address, and (if this is important) that it can only be called once.&lt;/p&gt;
&lt;a href=&quot;#fnref-4&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-5&quot;&gt;
&lt;p&gt;In a &lt;a href=&quot;https://github.com/sgoedecke/gatsby-blog/blob/5b6205fbe191a591bbcf61d094a6edbcfbd6475d/content/drafts/_icebox/ai-mass-surveillance/index.md&quot;&gt;draft post&lt;/a&gt; I never published, I ballpark-estimated all non-spam American email data at around seven trillion tokens per day (around a third of OpenAI’s total daily token usage).&lt;/p&gt;
&lt;a href=&quot;#fnref-5&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-6&quot;&gt;
&lt;p&gt;Should you do this? Probably not, but it’s a fascinating engineering problem, and I imagine the NSA has been thinking about these questions for several years by now. If the example bothers you, substitute some other more-ethical firehose of English language.&lt;/p&gt;
&lt;a href=&quot;#fnref-6&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-7&quot;&gt;
&lt;p&gt;Not counting evals, operations, standing up a trusted GPU cluster somewhere, scaling the physical hardware, and all the other thousand things you have to do in order to ship anything.&lt;/p&gt;
&lt;a href=&quot;#fnref-7&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[The famous o3 "GeoGuessr" prompt did not work]]></title><link>https://seangoedecke.com/the-o3-geoguessr-prompt-did-not-work/</link><guid isPermaLink="false">https://seangoedecke.com/the-o3-geoguessr-prompt-did-not-work/</guid><pubDate>Thu, 21 May 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;In April last year, Kelsey Piper &lt;a href=&quot;https://x.com/KelseyTuoc/status/1917340813715202540&quot;&gt;discovered&lt;/a&gt; that OpenAI’s o3 model was surprisingly good at figuring out where a photo was taken from. Like human “geoguessr” &lt;a href=&quot;https://www.youtube.com/@georainbolt&quot;&gt;pros&lt;/a&gt;, o3 could sometimes take a nondescript photo of a beach and tell you exactly where it is. Here’s the example Kelsey gave:&lt;/p&gt;
&lt;p&gt;&lt;span
      class=&quot;gatsby-resp-image-wrapper&quot;
      style=&quot;position: relative; display: block; margin-left: auto; margin-right: auto; max-width: 590px; &quot;
    &gt;
      &lt;a
    class=&quot;gatsby-resp-image-link&quot;
    href=&quot;/static/4113a246112c8b6db424a58af58a9a90/4d836/kelsey-geoguessr.jpg&quot;
    style=&quot;display: block&quot;
    target=&quot;_blank&quot;
    rel=&quot;noopener&quot;
  &gt;
    &lt;span
    class=&quot;gatsby-resp-image-background-image&quot;
    style=&quot;padding-bottom: 130.40540540540542%; position: relative; bottom: 0; left: 0; background-image: url(&apos;data:image/jpeg;base64,/9j/2wBDABALDA4MChAODQ4SERATGCgaGBYWGDEjJR0oOjM9PDkzODdASFxOQERXRTc4UG1RV19iZ2hnPk1xeXBkeFxlZ2P/2wBDARESEhgVGC8aGi9jQjhCY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2NjY2P/wgARCAAaABQDASIAAhEBAxEB/8QAGQAAAwADAAAAAAAAAAAAAAAAAAIEAQMF/8QAFgEBAQEAAAAAAAAAAAAAAAAAAAID/9oADAMBAAIQAxAAAAHdTzKdIsIQiExlbCB//8QAGRAAAwEBAQAAAAAAAAAAAAAAAQIRABAg/9oACAEBAAEFAleZTWJGDrQ95dfP/8QAFBEBAAAAAAAAAAAAAAAAAAAAIP/aAAgBAwEBPwEf/8QAFBEBAAAAAAAAAAAAAAAAAAAAIP/aAAgBAgEBPwEf/8QAGBAAAgMAAAAAAAAAAAAAAAAAABARMFH/2gAIAQEABj8CJW0//8QAHBABAAIDAAMAAAAAAAAAAAAAAQARECFRMWFx/9oACAEBAAE/IWRyE1Ja8k2FsKNWHqL9wv1luDH/2gAMAwEAAgADAAAAECAhz//EABQRAQAAAAAAAAAAAAAAAAAAACD/2gAIAQMBAT8QH//EABURAQEAAAAAAAAAAAAAAAAAABEg/9oACAECAQE/EGP/xAAbEAEAAwEBAQEAAAAAAAAAAAABABExIWGRUf/aAAgBAQABPxBwXjRnYAMagGIP2U1l6UQmIDkWufUER60Xgxu1i9mWK/s//9k=&apos;); background-size: cover; display: block;&quot;
  &gt;&lt;/span&gt;
  &lt;img
        class=&quot;gatsby-resp-image-image&quot;
        alt=&quot;geo&quot;
        title=&quot;geo&quot;
        src=&quot;/static/4113a246112c8b6db424a58af58a9a90/1c72d/kelsey-geoguessr.jpg&quot;
        srcset=&quot;/static/4113a246112c8b6db424a58af58a9a90/a80bd/kelsey-geoguessr.jpg 148w,
/static/4113a246112c8b6db424a58af58a9a90/1c91a/kelsey-geoguessr.jpg 295w,
/static/4113a246112c8b6db424a58af58a9a90/1c72d/kelsey-geoguessr.jpg 590w,
/static/4113a246112c8b6db424a58af58a9a90/a8a14/kelsey-geoguessr.jpg 885w,
/static/4113a246112c8b6db424a58af58a9a90/4d836/kelsey-geoguessr.jpg 920w&quot;
        sizes=&quot;(max-width: 590px) 100vw, 590px&quot;
        style=&quot;width:100%;height:100%;margin:0;vertical-align:middle;position:absolute;top:0;left:0;&quot;
        loading=&quot;lazy&quot;
      /&gt;
  &lt;/a&gt;
    &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Several people &lt;a href=&quot;https://www.astralcodexten.com/p/testing-ais-geoguessr-genius&quot;&gt;reproduced this&lt;/a&gt; with good results: not a 100% success rate, but clearly &lt;em&gt;far&lt;/em&gt; better than you’d do with a random human guess. The lesson here is that &lt;strong&gt;model capabilities can surprise us&lt;/strong&gt;. The o3 model had been released for two weeks before Kelsey’s tweet without anyone noticing how good it was at geolocation. What obscure capabilities did we never find? What capabilities of current models are we missing today?&lt;/p&gt;
&lt;p&gt;Some people drew &lt;a href=&quot;https://newsletter.angularventures.com/p/ai-s-geoguessr-genius-and-the-art-of-prompting-well&quot;&gt;another&lt;/a&gt; &lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1kep2bp/comment/mqlvv1a/&quot;&gt;lesson&lt;/a&gt; from this: that “prompt engineering” can unlock brand-new capabilities. This is because Kelsey had a &lt;a href=&quot;https://raw.githubusercontent.com/sgoedecke/ai_geolocation/refs/heads/main/prompts/geoguessr_protocol.txt&quot;&gt;magic prompt&lt;/a&gt; that she built over time. When o3 got something wrong, she would ask it how it could have avoided the mistake, and then included that in the prompt. Here’s the first 10% of that prompt, so you get the idea:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;You are playing a one-round game of GeoGuessr. Your task: from a single still image, infer the most likely real-world location. Note that unlike in the GeoGuessr game, there is no guarantee that these images are taken somewhere Google’s Streetview car can reach: they are user submissions to test your image-finding savvy. Private land, someone’s backyard, or an offroad adventure are all real possibilities (though many images are findable on streetview). Be aware of your own strengths and weaknesses: following this protocol, you usually nail the continent and country…&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This prompt impressed a lot of people, who &lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1kep2bp/comment/mqo3yzz/&quot;&gt;tried&lt;/a&gt; &lt;a href=&quot;https://www.thealgorithmicbridge.com/p/upload-a-picture-to-chatgpt-itll&quot;&gt;it&lt;/a&gt; &lt;a href=&quot;https://www.astralcodexten.com/p/testing-ais-geoguessr-genius&quot;&gt;out&lt;/a&gt; and reported that it correctly identified a lot of images. But of course, o3 correctly identified a lot of images with just a basic “think carefully about where this picture was taken?” prompt. Did the prompt actually help? It’d be tough to figure that out just from playing around in ChatGPT. You’d need to build an evaluation set of images and run o3 against them twice: once with the fancy prompt and once without it.&lt;/p&gt;
&lt;p&gt;So &lt;a href=&quot;https://github.com/sgoedecke/ai_geolocation/tree/main&quot;&gt;that’s what I did&lt;/a&gt;. I pulled 200 images from Wikimedia Commons, Geograph Britain and Ireland, and iNaturalist for the benchmark. You can read the AI-generated summary &lt;a href=&quot;https://github.com/sgoedecke/ai_geolocation/blob/main/results/dataset_mixed_200_o3_high_report.md&quot;&gt;here&lt;/a&gt;, but here’s the key table:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Prompt&lt;/th&gt;
&lt;th align=&quot;right&quot;&gt;n&lt;/th&gt;
&lt;th align=&quot;right&quot;&gt;Median km&lt;/th&gt;
&lt;th align=&quot;right&quot;&gt;Mean km&lt;/th&gt;
&lt;th align=&quot;right&quot;&gt;P25 km&lt;/th&gt;
&lt;th align=&quot;right&quot;&gt;P75 km&lt;/th&gt;
&lt;th align=&quot;right&quot;&gt;&amp;#x3C;=25 km&lt;/th&gt;
&lt;th align=&quot;right&quot;&gt;&amp;#x3C;=100 km&lt;/th&gt;
&lt;th align=&quot;right&quot;&gt;&amp;#x3C;=500 km&lt;/th&gt;
&lt;th align=&quot;right&quot;&gt;&amp;#x3C;=1000 km&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Default&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;200&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;&lt;strong&gt;83.2&lt;/strong&gt;&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;&lt;strong&gt;440.7&lt;/strong&gt;&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;&lt;strong&gt;16.4&lt;/strong&gt;&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;&lt;strong&gt;221.9&lt;/strong&gt;&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;58&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;&lt;strong&gt;109&lt;/strong&gt;&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;&lt;strong&gt;176&lt;/strong&gt;&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;&lt;strong&gt;182&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GeoGuessr prompt&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;200&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;102.3&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;481.9&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;18.5&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;277.8&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;&lt;strong&gt;59&lt;/strong&gt;&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;99&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;172&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;180&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;In general, the basic prompt did better on average. It consistently guessed closer to the actual location. Both prompts did pretty well, actually. Despite the fancy prompt being 10x larger, it only caused o3 to think for slightly longer (about one second on average, though the max was about double, at 10 minutes instead of 5 minutes). The images in my benchmark were fairly generic geoguessr-style outdoor images, with twelve indoor images thrown in for an extra challenge (the fancy prompt also did slightly worse on these).&lt;/p&gt;
&lt;p&gt;What’s going on? I think this shows &lt;strong&gt;how easy it is to fool yourself about the quality of prompting&lt;/strong&gt;. When the model is already pretty good at a task, you can give it a very elaborate prompt without impacting performance. It’ll still be pretty good, except this time it’s good &lt;em&gt;because of what you did&lt;/em&gt;. This is particularly true if you’re iterating with the model and asking it “what should I add to the prompt” for each mistake. Models will happily make up stories for you about their own reasoning processes, and will almost always say “yes, that helped a lot!” when you ask them if a particular prompt tweak made things better. The only way to actually know is by constructing some kind of benchmark&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;It’s also interesting to me that nobody checked this at the time. It took me about six hours of fairly-distracted work and about $15 to construct and run this benchmark. Why didn’t anyone do this when they were writing articles about how good the o3 prompt was?&lt;/p&gt;
&lt;p&gt;One charitable reason might be that the story was more about o3’s real geolocation ability than about the magic prompt. The pricing for o3 also used to be about five times more expensive (though a benchmark of 40 images instead of 200 would still have thrown doubt on how much water the prompt was carrying). Also, AI just moves so &lt;em&gt;fast&lt;/em&gt;. Geolocation was only the story for about a week: after that, GPT-4o’s &lt;a href=&quot;/ai-sycophancy&quot;&gt;sycophancy&lt;/a&gt; was what people were talking about. Another reason is that AI tooling wasn’t as good then. The benchmark was so easy for me to run because GPT-5.5 did most of the heavy lifting. Prior to strong agents, you would have had to write the (simple) benchmark yourself. I can’t point the finger too hard: I didn’t bother at the time either.&lt;/p&gt;
&lt;p&gt;Maybe my benchmark isn’t very good? The photos look reasonable enough: a wide variety of geoguessr-like shots of roads and landscapes, mostly. I could have tried to gather a few thousand photos instead of a few hundred, but if the magic prompt really was a big improvement you’d still expect to see that manifest on a benchmark this size. If someone wants to go and build a hundred-dollar geolocation benchmark instead of my fifteen-dollar one, I think that’d be an interesting project.&lt;/p&gt;
&lt;p&gt;Finally, let’s use the benchmark to answer a question I’ve had for a while: do gpt-5.4 and gpt-5.5 have o3’s geolocation abilities? The answer, apparently, is no.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Run&lt;/th&gt;
&lt;th align=&quot;right&quot;&gt;Median km&lt;/th&gt;
&lt;th align=&quot;right&quot;&gt;Mean km&lt;/th&gt;
&lt;th align=&quot;right&quot;&gt;&amp;#x3C;=25 km&lt;/th&gt;
&lt;th align=&quot;right&quot;&gt;&amp;#x3C;=100 km&lt;/th&gt;
&lt;th align=&quot;right&quot;&gt;&amp;#x3C;=500 km&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;o3 default&lt;/strong&gt;&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;&lt;strong&gt;83.2&lt;/strong&gt;&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;&lt;strong&gt;440.7&lt;/strong&gt;&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;58&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;&lt;strong&gt;109&lt;/strong&gt;&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;&lt;strong&gt;176&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;o3 GeoGuessr&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;102.3&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;481.9&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;&lt;strong&gt;59&lt;/strong&gt;&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;99&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;172&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gpt-5.4 default&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;163.3&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;638.9&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;26&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;74&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;148&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gpt-5.5 default&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;156.5&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;645.9&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;39&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;77&lt;/td&gt;
&lt;td align=&quot;right&quot;&gt;161&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Whatever o3 had that made it good at this task hasn’t transferred to newer models. &lt;/p&gt;
&lt;p&gt;edit: This post got some comments on &lt;a href=&quot;https://news.ycombinator.com/item?id=48219682&quot;&gt;Hacker News&lt;/a&gt;. The top &lt;a href=&quot;https://news.ycombinator.com/item?id=48220126&quot;&gt;comment&lt;/a&gt; worried that the models already knew the images, since they’re public domain. I thought about this but didn’t think it was worth sourcing brand new images: first, if the image/location pairs were in the training data, the models would have done better; second, even if they’re in the training data it still gives us useful comparison data from the prompt and for other models. I did confirm the images didn’t have EXIF metadata, so we’re not testing whether the prompt makes the model more or less likely to cheat.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;Benchmarks can mislead as well, but they’re better than just vibes.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[Prompts are technical debt too]]></title><link>https://seangoedecke.com/prompts-are-technical-debt-too/</link><guid isPermaLink="false">https://seangoedecke.com/prompts-are-technical-debt-too/</guid><pubDate>Wed, 20 May 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;It’s &lt;a href=&quot;https://www.tokyodev.com/articles/all-code-is-technical-debt&quot;&gt;common&lt;/a&gt; and correct to say that “all code is technical debt”. Adding code is a necessary evil for developing new features: you almost always have to do it, but each line of code adds to the complexity and maintenance burden of the system. All future changes to the system have to work with the existing code, or at least avoid breaking it. Once systems accumulate enough code, they become impossible for a single person to understand: instead of reading the code and understanding what it does, you must rely on guesses, theories and heuristics&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;. Sensible engineers write as little code as possible.&lt;/p&gt;
&lt;p&gt;They write a lot of prompts, though! Many large projects now have a set of codebase-specific prompt files: AGENTS.md, CLAUDE.md, those same files in sub-directories, and &lt;a href=&quot;https://github.com/anthropics/skills&quot;&gt;skills&lt;/a&gt;. If you’re building a program that uses AI&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;, you’ll have separate prompts for &lt;a href=&quot;https://github.com/anomalyco/opencode/tree/dev/packages/opencode/src/agent/prompt&quot;&gt;capabilities&lt;/a&gt; and for each &lt;a href=&quot;https://github.com/anomalyco/opencode/blob/dev/packages/opencode/src/tool/lsp.txt&quot;&gt;tool&lt;/a&gt;, as well as a whole set of &lt;a href=&quot;https://github.com/anomalyco/opencode/tree/dev/packages/opencode/src/session/prompt&quot;&gt;system prompts&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Prompts are important. Minor tweaks to a LLM’s prompt can unlock &lt;em&gt;significant&lt;/em&gt; performance improvements. If the same model feels different across Codex, Cursor, OpenCode, and Copilot, it’s almost certainly due to subtle differences in prompting. AI companies spend a lot of time testing and tweaking their prompts, so it makes sense why engineers would spend a lot of time tweaking their AGENTS.md files&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt; for their projects. I’d even call switching tools or workflows to be a form of prompting. If I start wrapping my agents in a &lt;a href=&quot;https://github.com/anomalyco/opencode/tree/dev/packages/opencode/src/session/prompt&quot;&gt;Ralph loop&lt;/a&gt;, pull in a new skill file, or install an &lt;a href=&quot;https://www.seangoedecke.com/model-context-protocol/&quot;&gt;MCP&lt;/a&gt; server, that’s still a change to my prompts even though I’m not the one who wrote it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;I think it is a bad idea to spend a ton of time tweaking a bespoke agentic coding setup.&lt;/strong&gt; Why is that, given that prompt adjustments can deliver a lot of value? Because prompt adjustments are &lt;em&gt;model-specific&lt;/em&gt;. Earlier I said that AI companies spend a lot of time tweaking their prompts. In fact, they spend that amount of time for each new model release. A prompt that worked great for GPT-5.4 won’t necessarily work as well for GPT-5.5. You have to “learn how to hold the model” each time. &lt;/p&gt;
&lt;p&gt;In other words, a set of prompts that you carefully crafted in January this year might be out of date or actively harmful by February. Worse still, you might not even notice. Model capabilities are already so hard to pin down (unless you’re running every problem through different models and tools), and even weak AI systems are surprisingly good at some problems. You might just think “huh, the new Anthropic model isn’t as impressive as the hype”, or “wow, Claude Code has gotten worse recently”.&lt;/p&gt;
&lt;p&gt;In this sense, &lt;strong&gt;prompts are a worse form of technical debt than code&lt;/strong&gt;. When technical debt blows up, it usually causes errors or a tangible slowdown as you try to understand the code. Prompts will decay silently. Also, even janky code tends to be relatively stable when untouched, but every single model upgrade could turn a functional prompt into a non-functional one.&lt;/p&gt;
&lt;p&gt;Could you simply decide not to upgrade models? Some people are trying this, but the pace of improvement is fast enough that that isn’t really practical. A delicately-prompted agentic harness built around GPT-4.1 is always going to underperform a bare-bones harness built around Opus 4.7. This might be a sensible strategy at some point in the future, when the rate of model improvement slows down (or when models are so capable that you don’t need the extra intelligence for normal engineering tasks), but I don’t believe it’s a good strategy today.&lt;/p&gt;
&lt;p&gt;In my view, most people should just be picking an AI coding tool maintained by a third-party company (Claude Code, Codex, Cursor, Copilot, etc) and leaving it as unconfigured as possible, so they can piggyback on the work of teams of engineers who are evaluating and tweaking prompts with each new model. Avoid MCP and skills unless absolutely necessary, and keep them off by default. At least this way if one of those teams gets it badly wrong, users will notice eventually and complain about it.&lt;/p&gt;
&lt;p&gt;When you write AGENTS.md files, try to avoid behavior steering (like the now-outdated “think step by step”, “you are a skilled engineer”, or “if you get a task right I will tip you $200”). Keep them limited to specific, concrete facts about the project. Don’t let models fill your AGENTS.md with pages of barely-reviewed text, for the same reason that you wouldn’t let them fill your codebase with pages of barely-reviewed code. Write your prompts yourself, and delete them whenever you get the chance.&lt;/p&gt;
&lt;p&gt;edit: this ended up being the topic of a Theo &lt;a href=&quot;https://www.youtube.com/watch?v=WnBx1Vi7M6w&quot;&gt;video&lt;/a&gt; on YouTube.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;Almost every system you might get paid to work on is in this category (if not in the code of the system itself, then in its dependencies and libraries).&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;Instead of just using AI to build a program. This distinction was a real pain when I was working on &lt;a href=&quot;https://github.blog/news-insights/product-news/introducing-github-models/&quot;&gt;GitHub Models&lt;/a&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[The just-say-no engineer was a ZIRP phenomenon]]></title><link>https://seangoedecke.com/the-just-say-no-engineer-was-a-zirp-phenomenon/</link><guid isPermaLink="false">https://seangoedecke.com/the-just-say-no-engineer-was-a-zirp-phenomenon/</guid><pubDate>Mon, 18 May 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;The engineer who &lt;a href=&quot;https://www.nair.sh/guides-and-opinions/communicating-your-expertise/why-senior-developers-fail-to-communicate-their-expertise#a-senior-developer-is-a-problem-avoider&quot;&gt;says no all the time&lt;/a&gt; is a real archetype among senior and staff engineers. Their role is to slow things down, to block the development of features that add complexity, and to ensure that as little code gets written as possible (since code is a liability).&lt;/p&gt;
&lt;p&gt;We can think of this as the just-say-no engineer&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;, as opposed to the just-say-yes engineer. The just-say-yes engineer is obsessed with moving fast, approves code changes by default, values &lt;a href=&quot;https://en.wikipedia.org/wiki/Mean_time_to_repair&quot;&gt;MTTR&lt;/a&gt; over &lt;a href=&quot;https://en.wikipedia.org/wiki/Mean_time_between_failures&quot;&gt;MTBF&lt;/a&gt;, and tends to ship a lot of code. The just-say-no engineer is obsessed with quality, is happy to move slowly, and blocks code changes by default. Most engineers are somewhere in the middle of the spectrum. By “just-say-no engineer”, I’m talking about the group of engineers who most strongly identify with that archetype.&lt;/p&gt;
&lt;p&gt;The just-say-no engineer is having a hard time in the era of AI. It used to be that they only had to say no to more junior engineers’ handwritten PRs, but now they have to say no to a barrage of AI-generated code, some of it generated by managers and VPs who are politically difficult to say no to. For the first time in their careers, they’re under a lot of pressure to lower their standards and start saying yes. However, &lt;strong&gt;this isn’t because of AI.&lt;/strong&gt; It’s because of the end of ZIRP.&lt;/p&gt;
&lt;h3&gt;ZIRP and the just-say-no engineer&lt;/h3&gt;
&lt;p&gt;ZIRP, or the “zero interest rate policy”, is a shorthand for the era of software development between 2008 and 2022 when banks were allowing companies to borrow money at near-zero interest rates. During this period, investors were throwing borrowed money at &lt;em&gt;anything&lt;/em&gt;, which meant that tech companies were incentivized to constantly hire engineers for low-risk high-reward projects&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;. Successful companies would routinely grow from tens of engineers to thousands, who would go and work on all kinds of things: tangential open-source projects, endless technology migrations, rewrites into other languages, and so on.&lt;/p&gt;
&lt;p&gt;It was a great time to be a software engineer. We had a lot of bargaining power, and could get paid top dollar to do almost anything. The bosses largely didn’t care, because (a) teams were growing so fast they couldn’t pay attention, and (b) just having more engineers around was beneficial to the stock price, which was the main thing they cared about. But tech companies did have one problem: with so many engineers running wild, how would they keep their systems from becoming completely unmanageable? &lt;strong&gt;Enter the just-say-no engineer.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In this environment, having a very senior engineer whose only job is to say no to things was actually quite valuable to the company. There are a few reasons for this:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Having half of the company’s engineers enmeshed in an endless loop of proposing changes and being told no was totally fine — they didn’t need to be productive anyway, and this way they weren’t impacting business-critical systems.&lt;/li&gt;
&lt;li&gt;It also solved the problem of the 5% of engineers who would get drunk on their technical freedom and make wild proposals like migrating to a hand-rolled database. &lt;/li&gt;
&lt;li&gt;Having a reputation for a very high technical bar is a positive for hiring (and remember, during ZIRP every tech company was always hiring)&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;The end of ZIRP&lt;/h3&gt;
&lt;p&gt;When banks hiked interest rates, almost every tech company immediately laid off 5-20% of their engineers. It was just no longer profitable to keep a bloated engineering staff around to boost the stock price. Instead, companies had to actually make money&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;. However, that wasn’t a good public explanation for the layoffs, since it sounds weak to admit that you were paying hundreds of engineers to do unprofitable work. Fortunately, the end of ZIRP coincided roughly with the rise of ChatGPT, so tech companies were able to to blame their layoffs on the power of AI. Saying “with this transformative new technology, we’re able to deliver 10x the value with half the engineers” is a much stronger message, even though it doesn’t make much sense (if this is true, why not keep your engineers and deliver 20x the value?)&lt;/p&gt;
&lt;p&gt;Something like this dynamic has been happening to the just-say-no engineer. Tech companies are now more focused than at any time in the past two decades. They are not doing a bunch of random crap anymore; instead they’re desperately chasing new capabilities and features that can make money (mostly built on AI, for obvious reasons). This new environment is &lt;em&gt;actively inimical&lt;/em&gt; to the just-say-no engineer. It’s as if a shark got pulled out of the deep ocean and dropped into a fast-flowing river: what was once a powerful apex predator is now disoriented and flailing.&lt;/p&gt;
&lt;p&gt;This kind of engineer used to enjoy implicit (albeit distant) support from their management. If someone complained, they’d often get told “that engineer knows what they’re doing, if they said no, then I trust them”. Now that support is gone. The just-say-no engineer is now being criticized and actively overruled by their management. They’re being told to be more of a team player, to find a way to say yes, or are simply no longer being consulted (with the company’s blessing) on key decisions. They’re getting bad reviews for the exact same behavior that’s been rewarded pre-2022&lt;sup id=&quot;fnref-4&quot;&gt;&lt;a href=&quot;#fn-4&quot; class=&quot;footnote-ref&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;None of this depends upon AI.&lt;/strong&gt; If LLMs had not taken off this decade, we would still be seeing the same cultural shifts in the industry. Companies would still be laying off engineers, and the engineers whose job has been to say no to things would still be upset and confused about why they’re now being punished for saying no.&lt;/p&gt;
&lt;h3&gt;AI&lt;/h3&gt;
&lt;p&gt;Ironically, if ZIRP had not ended, this would be a glorious moment for the just-say-no engineers. LLMs would have thrown fuel on the “engineers running wild” problem that the just-say-no engineers were empowered to solve. Tech companies, unable to publicly or privately cast doubt on AI-assisted coding&lt;sup id=&quot;fnref-5&quot;&gt;&lt;a href=&quot;#fn-5&quot; class=&quot;footnote-ref&quot;&gt;5&lt;/a&gt;&lt;/sup&gt;, would have relied &lt;em&gt;heavily&lt;/em&gt; on these engineers to prevent the tsunami of AI code from swamping the entire company. They would have been paid even better and celebrated like kings.&lt;/p&gt;
&lt;p&gt;Instead, LLMs are adding insult to injury for the just-say-no engineer. They’re forced to watch while other engineers merge AI-generated PRs that would previously have been blocked, and are told to use the tools themselves: to become the kind of engineer they’ve spent their entire careers battling against.&lt;/p&gt;
&lt;p&gt;Worse still, the AI tooling mostly &lt;em&gt;works&lt;/em&gt;. It’s not (yet) causing any kind of catastrophe&lt;sup id=&quot;fnref-6&quot;&gt;&lt;a href=&quot;#fn-6&quot; class=&quot;footnote-ref&quot;&gt;6&lt;/a&gt;&lt;/sup&gt;. The code isn’t quite as clean, and it’s a bit less well-understood, but it’s good enough (particularly in a world where companies are trying lots of new things and abandoning the ones that fail). So the just-say-no engineer faces not just a threat to their livelihood, but to their entire self-identity: they have to either insist that the apocalypse is right around the corner, or accept that their technical role was contingent on a &lt;em&gt;really weird&lt;/em&gt; economic environment in the tech industry.&lt;/p&gt;
&lt;h3&gt;Pure and impure engineering&lt;/h3&gt;
&lt;p&gt;Will the just-say-no engineer go extinct? No. They don’t fit well into every single tech company anymore, but there are domains where they’re needed. In &lt;a href=&quot;/pure-and-impure-engineering/&quot;&gt;&lt;em&gt;Pure and impure software engineering&lt;/em&gt;&lt;/a&gt; I drew a distinction between “pure” engineering, which has a well-scoped, largely technical goal (like building a compiler or a language runtime) and “impure engineering”, which has a poorly-scoped, largely customer-driven goal (like trying out a new feature you’re not sure will work). During the ZIRP era, tech companies did a lot more pure work (for instance, building &lt;a href=&quot;https://en.wikipedia.org/wiki/React_(software)&quot;&gt;React&lt;/a&gt;), and tended to treat even impure work like pure work. The just-say-no engineer is &lt;em&gt;great&lt;/em&gt; for pure work, because pure codebases have to have a much higher bar for quality and can tolerate slower development cycles.&lt;/p&gt;
&lt;p&gt;Most tech companies are still doing some kind of pure work, typically in their core infrastructure pieces. This is essential work, but it doesn’t require a huge engineering team, and it’s rarely in &lt;a href=&quot;https://www.seangoedecke.com/the-spotlight/&quot;&gt;the spotlight&lt;/a&gt;. If you’re a just-say-no engineer and you want to stay that way, I would recommend trying to move into one of these roles (and accepting that you’ll have a more limited scope than you did in the 2010s).&lt;/p&gt;
&lt;h3&gt;Summary&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Some senior and staff engineers operate as gatekeepers, slowing down development and saying no to most things&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;This was a critical role during ZIRP, because:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Tech companies had thousands of engineers who were empowered to do basically whatever they wanted, so without gatekeeping the systems would have fallen apart&lt;/li&gt;
&lt;li&gt;Tech companies didn’t care that much if they got anything done&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;When ZIRP ended, the environment for this kind of engineer became much worse, since tech companies were now actually focused on accomplishing things and the “do whatever you want” era was over&lt;/li&gt;
&lt;li&gt;Like with layoffs, this shift is often blamed on AI, but it would have happened even if powerful LLMs had not emerged at all. It’s an end-of-ZIRP phenomenon&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;edit: this post got some comments on &lt;a href=&quot;https://lobste.rs/s/i2szle/just_say_no_engineer_was_zirp_phenomenon&quot;&gt;lobste.rs&lt;/a&gt; and &lt;a href=&quot;https://www.reddit.com/r/programming/comments/1thf964/the_justsayno_engineer_was_a_zirp_phenomenon/&quot;&gt;Reddit&lt;/a&gt;, including one of the &lt;a href=&quot;https://lobste.rs/c/f3g1tn&quot;&gt;cruelest&lt;/a&gt; comments I’ve ever read about my blog. A more concrete criticism was &lt;a href=&quot;https://lobste.rs/c/yoouec&quot;&gt;about&lt;/a&gt; &lt;a href=&quot;https://www.reddit.com/r/programming/comments/1thf964/comment/omnw6an/&quot;&gt;my&lt;/a&gt; off-hand remark that the models work: commenters felt it was too early to say, because the impact of bad code takes a while to manifest. Fair enough. “It’s too early to say” is never really &lt;em&gt;wrong&lt;/em&gt;, though I think it’s clear that AI code is not immediately fatal. &lt;a href=&quot;https://www.reddit.com/r/programming/comments/1thf964/comment/omn4990/&quot;&gt;Other&lt;/a&gt; &lt;a href=&quot;https://lobste.rs/c/eluuto&quot;&gt;commenters&lt;/a&gt; argued that the just-say-no archetype existed for decades prior to ZIRP (e.g. Linus Torvalds). I agree with that, but I think the niche for this kind of engineer was artificially expanded by ZIRP, and now has contracted again. Finally, an &lt;a href=&quot;https://lobste.rs/c/bromgo&quot;&gt;interesting comment&lt;/a&gt; claiming that the just-say-no engineer had a niche because (a) people were using dynamic languages, and (b) observability/feature-flag-etc tooling was not yet mature.&lt;/p&gt;
&lt;p&gt;I also want to share this quote from a reader, via email:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;…in a strange way your posts give me comfort that I’m not alone in some strange bubble where all of a sudden I’m the only one that’s somehow always wrong. I’m at somewhat of a crossroads as I either need to lower my standards and become the always say yes engineer to gain favour with managers again (which is in conflict with who I am) or move on and potentially risk landing at another company with the exact same setup. &lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;edit: Another round of comments from &lt;a href=&quot;https://news.ycombinator.com/item?id=48289439&quot;&gt;Hacker News&lt;/a&gt;, &lt;a href=&quot;https://news.ycombinator.com/item?id=48289749&quot;&gt;with&lt;/a&gt; &lt;a href=&quot;https://news.ycombinator.com/item?id=48289785&quot;&gt;several&lt;/a&gt; &lt;a href=&quot;https://news.ycombinator.com/item?id=48290371&quot;&gt;commenters&lt;/a&gt; &lt;a href=&quot;https://news.ycombinator.com/item?id=48289668&quot;&gt;wishing&lt;/a&gt; &lt;a href=&quot;https://news.ycombinator.com/item?id=48289953&quot;&gt;I’d&lt;/a&gt; provided more hard evidence for my theory. Unfortunately it doesn’t work like that — I’m writing from my own experience, which is just my tiny window into what the industry was like pre-and-post-ZIRP. Your mileage may (and often &lt;a href=&quot;https://news.ycombinator.com/item?id=48290017&quot;&gt;does&lt;/a&gt;) vary. It’s an interesting question how you might go about testing something like this. Maybe survey a few hundred senior+ engineers in 2010 and 2026, asking how many times a week they said “no” to something, and whether that “no” was overruled?&lt;/p&gt;
&lt;p&gt;I do want to address comments like &lt;a href=&quot;https://news.ycombinator.com/item?id=48290126&quot;&gt;this&lt;/a&gt; and &lt;a href=&quot;https://news.ycombinator.com/item?id=48290499&quot;&gt;this&lt;/a&gt;, which argue that saying no was essential both before and after ZIRP. Yes, but the difference (in my view) is: pre-ZIRP, &lt;em&gt;management&lt;/em&gt; did not like saying no to engineers, but post-ZIRP they rapidly built that muscle, and so now no longer need a group of engineers saying no for them.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;Part of the appeal here is the lure of the guru. In kung fu films, those who know martial arts perform furious acrobatics, but the true expert barely needs to move at all. For the same reasons, it sounds profound to say something like “junior engineers produce tons of code, seniors very little, and staff engineers &lt;em&gt;remove&lt;/em&gt; code”. Of course this is false. Staff engineers are expected to be able to produce a lot of working code very quickly, when they need to.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;I wrote about this a lot more in &lt;a href=&quot;/good-times-are-over/&quot;&gt;&lt;em&gt;The good times in tech are over&lt;/em&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;Not necessarily make a &lt;em&gt;profit&lt;/em&gt;, but at least bring in revenue.&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-4&quot;&gt;
&lt;p&gt;Or pre-2023, or even pre-2024 or 2025. Cultural change lags behind economic incentives, sometimes by several years.&lt;/p&gt;
&lt;a href=&quot;#fnref-4&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-5&quot;&gt;
&lt;p&gt;For fear of killing the vibe (and thus the stock price).&lt;/p&gt;
&lt;a href=&quot;#fnref-5&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-6&quot;&gt;
&lt;p&gt;If you think there have been more incidents recently, consider that (a) you might be &lt;a href=&quot;https://news.ycombinator.com/item?id=48086786&quot;&gt;wrong&lt;/a&gt;, or (b) that other end-of-ZIRP factors (like increased velocity or layoffs) might be primarily responsible.&lt;/p&gt;
&lt;a href=&quot;#fnref-6&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[How I use LLMs as a staff engineer in 2026]]></title><link>https://seangoedecke.com/how-i-use-llms-in-2026/</link><guid isPermaLink="false">https://seangoedecke.com/how-i-use-llms-in-2026/</guid><pubDate>Sun, 17 May 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;A bit over a year ago I wrote &lt;a href=&quot;https://www.seangoedecke.com/how-i-use-llms/&quot;&gt;&lt;em&gt;How I use LLMs as a staff engineer&lt;/em&gt;&lt;/a&gt;. Here’s a brief summary of what I used AI for last year:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Smart autocomplete with Copilot&lt;/li&gt;
&lt;li&gt;Short tactical changes in areas I don’t know well (always reviewed by a SME)&lt;/li&gt;
&lt;li&gt;Writing lots of use-once-and-throwaway research code&lt;/li&gt;
&lt;li&gt;Asking lots of questions to learn about new topics (e.g. the Unity game engine)&lt;/li&gt;
&lt;li&gt;Last-resort bugfixes, just in case it can figure it out immediately&lt;/li&gt;
&lt;li&gt;Big-picture proofreading for long-form English communication&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are some tasks I explicitly &lt;em&gt;didn’t&lt;/em&gt; use AI for last year:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Writing whole PRs for me in areas I’m familiar with&lt;/li&gt;
&lt;li&gt;Writing ADRs or other technical communications&lt;/li&gt;
&lt;li&gt;Research in large codebases and finding out how things are done&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;February 2025 was a long time ago. Back then the best model was the first reasoning model, OpenAI’s o1. Agents &lt;em&gt;sort of&lt;/em&gt; worked, but would often get stuck or thrown off by compaction. What’s changed since then?&lt;/p&gt;
&lt;h3&gt;Agents are good now&lt;/h3&gt;
&lt;p&gt;The biggest change is that &lt;strong&gt;I now use LLMs to produce entire PRs in areas I’m familiar with&lt;/strong&gt;. A year ago I would very occasionally ask an agent to make changes to a single file if it was a simple change I couldn’t be bothered typing out. Sometimes I would copy a function I wrote into a LLM chat window for feedback. But now I start every single change by asking an agent to solve the problem, and usually push the PR after a single editing pass.&lt;/p&gt;
&lt;p&gt;In late 2025 I used a lot of open VSCode windows. In early 2026, that changed to terminal tabs with the Copilot CLI, particularly when I needed to make changes across multiple repos at the same time. Now I use the &lt;a href=&quot;https://github.blog/changelog/2026-05-14-github-copilot-app-is-now-available-in-technical-preview/&quot;&gt;GitHub Copilot app&lt;/a&gt; a &lt;em&gt;lot&lt;/em&gt; (tens of sessions per day). &lt;/p&gt;
&lt;p&gt;This reflects a shift from having to line-edit the agent basically as it went to only doing an editing pass right at the end. Early agents would go wrong a lot and not be able to recover, so it was valuable to keep an eye on their thought processes and step in to pause them and set them right. In my experience, current agents move too fast to do this, and recover their own mistakes most of the time anyway.&lt;/p&gt;
&lt;p&gt;Sometimes I don’t even need to make edits and I can just push the change as-is, though this is rare: if nothing else, I typically go through and remove some of the over-commenting and other LLM-isms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;I do a &lt;em&gt;lot&lt;/em&gt; of skimming through and evaluating agent changes.&lt;/strong&gt; Most of the time I reject them entirely, just based on “eh, that’s not what I was thinking”. On average it takes me about thirty seconds to make this initial assessment. If the change looks alright after that, I’ll dig in and do a proper review to make sure I understand it and it’s doing the right thing. For difficult tasks, I’ll often reject five or six (or more!) agent attempts before accepting one as good enough to work with, or giving up and making the change by hand.&lt;/p&gt;
&lt;h3&gt;Investigating bugs&lt;/h3&gt;
&lt;p&gt;I rely on LLMs even more for bug-hunting than I do for making changes. In 2025, I used to throw the occasional bug at a LLM, just in case it was able to rapidly come up with an explanation. Now I throw &lt;em&gt;every&lt;/em&gt; bug at a LLM (typically by opening a new agent session and pasting in the bug report), because it’s able to correctly diagnose 80% of issues on its own. Current agents are &lt;em&gt;really good&lt;/em&gt; at chasing down bugs, particularly when you give them a vantage point across multiple repositories.&lt;/p&gt;
&lt;p&gt;I’m still better at it. Just last week I had a tricky bug that took about fourteen agent sessions before one finally figured it out. What was I doing in between and around those sessions?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Digging up extra context on the bug (from logs, Slack, etc) and reporting it to the agents&lt;/li&gt;
&lt;li&gt;Building my own mental model of the problem, of course&lt;/li&gt;
&lt;li&gt;Setting up my own reproduction of the bug (in parallel with the agents’ efforts)&lt;/li&gt;
&lt;li&gt;Responding to agent sessions with “no, your theory can’t be right because of X” (or just killing and restarting the session with that extra hint)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Ultimately an agent was the one to catch the bug. But I still count it as my find, because by that point I had narrowed the search space tightly enough that agent session #14 had a &lt;em&gt;significantly&lt;/em&gt; easier problem to solve than agent session #1. In other words, &lt;strong&gt;human expertise still matters a lot for investigating bugs&lt;/strong&gt;.&lt;/p&gt;
&lt;h3&gt;Writing&lt;/h3&gt;
&lt;p&gt;I &lt;em&gt;almost always&lt;/em&gt; write my own PR descriptions, since LLMs over-communicate and are bad at expressing the “core idea” behind a change. Writing the PR description by hand also signals to reviewers that I’ve reviewed the change myself, and I’m not asking them to be the first human to read the diff. The only time when I don’t write the PR description is when the change is trivial and the agent-generated description is one sentence. At that point I just leave it alone.&lt;/p&gt;
&lt;p&gt;I still don’t use LLMs to write Slack messages, ADRs, issues and so forth. I believe I have a better sense of what’s important to communicate, and I want to signal that there’s a human being thinking about the content.&lt;/p&gt;
&lt;p&gt;I still never use LLMs to write blog posts, though I do run each draft post through a LLM for feedback. OpenAI models used to be &lt;em&gt;terrible&lt;/em&gt; at this and have only very recently gotten acceptable with GPT-5.5. Both OpenAI and Anthropic models still try to water down my arguments, but I’ve accepted that as part of the LLM “house style” and just ignore that part of the feedback.&lt;/p&gt;
&lt;h3&gt;Testing and setup&lt;/h3&gt;
&lt;p&gt;Another thing I do now is &lt;strong&gt;try and push as much testing and setup work as possible onto the agents&lt;/strong&gt;. In 2025, I used to sometimes ask a LLM to produce a test script of curl commands that I could run against my dev server. In 2026, I just ask an agent to go and test my change, then read the log of what it did.&lt;/p&gt;
&lt;p&gt;I don’t test UI work like this, partly because it’s more fiddly and partly because I don’t trust agents to be sensitive to the subtle look-and-feel aspects of a change.&lt;/p&gt;
&lt;p&gt;Agents will write expansive unit tests without having to be told, but I do sometimes ask them to put together broader integration tests for a change. In general I now consider test code to be cheap: if I’m wondering whether a test would be useful, I just add it (so long as I know it won’t be flaky). Of course LLMs sometimes produce strange and unsatisfying test code — I do read it to catch obvious blunders — but I review it with a more generous eye than my actual production code.&lt;/p&gt;
&lt;p&gt;I’ll also task an agent with annoying local setup tasks that involve config wrangling on my machine. For instance, if my nvm installation is not switching my Node version correctly, I will often open a Copilot CLI agent and ask it to figure it out. This is a more-or-less direct replacement for Googling the problem, and is much quicker since the agent can run the trivial bash commands to diagnose and fix the problem itself.&lt;/p&gt;
&lt;h3&gt;Summary&lt;/h3&gt;
&lt;p&gt;The main thing that’s changed in the last fifteen months is that &lt;strong&gt;agents are really good now&lt;/strong&gt;. They’ve gone from something I used occasionally and suspiciously to something I use constantly and with light supervision.&lt;/p&gt;
&lt;p&gt;The core of my job is still the same: &lt;a href=&quot;/how-to-ship&quot;&gt;shipping projects&lt;/a&gt;, exercising my judgement, &lt;a href=&quot;/how-to-influence-politics/&quot;&gt;influencing tech company politics&lt;/a&gt;. But I now have a much wider net for small pieces of work that I’m willing to take on, which includes basically anything I can hand off to an agent and expect it to get more or less right.&lt;/p&gt;
&lt;p&gt;I used to spend a lot of time putting work off, either by delegating it or just saying “sorry, I don’t have time to do that now”. Now I get to say “yes” a lot more (at least when it comes to minor low-risk tweaks)&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;Overall, here’s what I now use AI for:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Writing (or drafting, depending on complexity) every code change I make&lt;/li&gt;
&lt;li&gt;Investigating and fixing bugs, either autonomously for most bugs or with my close involvement for trickier ones&lt;/li&gt;
&lt;li&gt;Research in large codebases, since current agents are now good enough to give the right answer almost all the time (and when they’re wrong, it’s clear from reading the explanation that they’ve missed something)&lt;/li&gt;
&lt;li&gt;Manual testing and local-machine setup or troubleshooting&lt;/li&gt;
&lt;li&gt;I still use AI for asking lots of questions to learn about topics, and for proofreading&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here’s what I still don’t use AI for:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Writing any kind of public communication for me (PR descriptions, ADRs, messages) with the exception of trivial two-line PRs&lt;/li&gt;
&lt;li&gt;Writing code that I don’t carefully review&lt;/li&gt;
&lt;li&gt;Testing any kind of UI&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In my view, &lt;strong&gt;the current core AI skill is shifting as much work onto AI agents as possible, without going too far&lt;/strong&gt;. Many people are under-utilizing agents: not allowing them to investigate bugs or test their changes, or not throwing enough simple tasks at them. Other people are over-utilizing them: using them to write messages that ought to be hand-written, or trusting them to make sweeping changes that need careful human review. Since my last post, the balance has tilted more towards the agents, but &lt;em&gt;finding&lt;/em&gt; the balance remains as tricky as ever.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;For once I can actually give an example, since it’s in a public repository. Someone internal wanted to be able to use the &lt;a href=&quot;https://github.com/actions/ai-inference&quot;&gt;actions/ai-inference&lt;/a&gt; GitHub Action with Copilot-backed inference (for various reasons), and instead of saying “sorry, I don’t have time to get to it”, I was able to throw it at an agent. If a human had to do this, the output would likely have been better, but it wouldn’t have gotten done for weeks (if at all).&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[DeepSeek-V4-Flash means LLM steering is interesting again]]></title><link>https://seangoedecke.com/steering-vectors/</link><guid isPermaLink="false">https://seangoedecke.com/steering-vectors/</guid><pubDate>Sat, 16 May 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Ever since &lt;a href=&quot;https://www.anthropic.com/news/golden-gate-claude&quot;&gt;Golden Gate Claude&lt;/a&gt; I’ve been fascinated with “steering”: the idea that you can guide LLM outputs by directly manipulating the activations of the model mid-flight.&lt;/p&gt;
&lt;h3&gt;DeepSeek V4 Flash&lt;/h3&gt;
&lt;p&gt;I was inspired to write this post by antirez’s recent project &lt;a href=&quot;https://github.com/antirez/ds4/tree/main&quot;&gt;DwarfStar 4&lt;/a&gt;, which is a version of &lt;a href=&quot;https://github.com/ggml-org/llama.cpp&quot;&gt;llama.cpp&lt;/a&gt; that’s been stripped down to run only DeepSeek-V4-Flash. What’s so special about this model? It might be what many engineers have been waiting for: a local model good enough to compete with at least the low end of frontier model agentic coding.&lt;/p&gt;
&lt;p&gt;Since steering requires a local model, it’s now practical for many engineers to try it out for the first time. And indeed, antirez has baked &lt;a href=&quot;https://github.com/antirez/ds4/tree/main/dir-steering&quot;&gt;steering&lt;/a&gt; into DwarfStar 4 as a first-class citizen. Right now it’s very rudimentary (basically just the toy “verbosity” example you can replicate via prompting), but the initial release was only &lt;a href=&quot;https://github.com/antirez/ds4/commit/d997b56c151184bcff469dd8302ed97f23481024&quot;&gt;eight days ago&lt;/a&gt;. I plan to follow this project closely.&lt;/p&gt;
&lt;h3&gt;How steering works&lt;/h3&gt;
&lt;p&gt;The basic idea behind steering is extracting a concept (like “respond tersely”) from the model’s internal brain state, then reaching in during inference and boosting the numerical activations that form that concept.&lt;/p&gt;
&lt;p&gt;One way you might do this is to feed your model the same set of a hundred prompts twice, once with the normal prompts and once with the words “respond tersely” appended. Then measure the difference in the model’s activations&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt; for each prompt pair (by subtracting one activation matrix from the other). That’s your “steering vector”. In theory, you can go and add that to the same activation layer for any prompt and get the same effect (of the model responding tersely).&lt;/p&gt;
&lt;p&gt;Another, more sophisticated way you might do this is to train a second model to extract “features” from your model’s activations: patterns of behavior that seem to show up together. Then you can try to map those features back to individual concepts, and boost them in the same way. This is more or less what Anthropic is doing with &lt;a href=&quot;https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html&quot;&gt;sparse autoencoders&lt;/a&gt;&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;. It’s the same principle as the naive approach, but it lets you capture deeper patterns (at the cost of being much more expensive in time, compute and expertise).&lt;/p&gt;
&lt;h3&gt;Why steering is interesting&lt;/h3&gt;
&lt;p&gt;Steering sounds like a cheat code. Instead of painstakingly assembling a training set that tries to push the model towards the “smart” end of the distribution in its training data, why not simply go uncover the “smart” dial in the model’s brain and turn it all the way to the right?&lt;/p&gt;
&lt;p&gt;It also seems like a more elegant way to adjust the way models talk. Instead of fiddling with the prompt (adding or removing qualifiers like “you MUST”), couldn’t we just have a control panel of sliders like “succinctness/verbosity” or “conscientiousness/speed” and move them around directly?&lt;/p&gt;
&lt;p&gt;Finally, it’s just &lt;em&gt;cool&lt;/em&gt;. Watching Golden Gate Claude unwillingly &lt;a href=&quot;https://www.anthropic.com/news/golden-gate-claude&quot;&gt;drag&lt;/a&gt; every sentence back to the Golden Gate Bridge is as fascinating and unsettling as Oliver Sacks’ neurological &lt;a href=&quot;https://en.wikipedia.org/wiki/The_Man_Who_Mistook_His_Wife_for_a_Hat&quot;&gt;anecdotes&lt;/a&gt;. What if your own mind was tweaked in a similar way? Would it still be you?&lt;/p&gt;
&lt;h3&gt;Why steering hasn’t been used&lt;/h3&gt;
&lt;p&gt;Why don’t we steer more, then? Why don’t ChatGPT and Claude Code already have a steering panel where you can adjust the model’s brain in real time? One reason is that steering is kind of an unfortunately “middle class” idea in AI research.&lt;/p&gt;
&lt;p&gt;It’s beneath the big AI labs, who can manipulate their models directly without having to do awkward brain surgery mid-inference. Anthropic is working on this stuff, but largely from an interpretability and safety perspective (as far as I know). When they want a model to behave in a certain way, they don’t mess around with steering, they just train the model.&lt;/p&gt;
&lt;p&gt;Steering is also out of reach for regular AI users like you and me&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;, who use LLMs via an API and thus don’t have access to the model weights or activations needed to steer the model. Only OpenAI can identify or expose steering vectors for GPT-5.5, for instance. We could do this for open-weights models, but until very recently (more on that later) there haven’t been any open models strong enough to be worth doing this for.&lt;/p&gt;
&lt;p&gt;On top of that, most basic applications of steering are outcompeted by just prompting the model. It sounds pretty impressive to be able to manipulate the model’s brain directly. But you know what else manipulates the model’s brain directly? Prompt tokens. You can exercise fairly fine-grained control over activations with steering, but you can already exercise &lt;em&gt;extremely&lt;/em&gt; fine-grained control by tweaking the language of your prompt. In other words, there’s not much point going to the trouble to steer a model to be more verbose when you could simply &lt;em&gt;ask&lt;/em&gt;.&lt;/p&gt;
&lt;h3&gt;Steering the unpromptable&lt;/h3&gt;
&lt;p&gt;One way for steering to be really useful is if we could identify a concept that can’t be prompted for. What about “intelligence”? You used to be able to prompt for intelligence — this is why 4o-era prompting always began with “you are an expert” — but current-generation models have that baked into their personalities, so prompting for it does nothing. Maybe steering for it would still work?&lt;/p&gt;
&lt;p&gt;Ultimately this is an empirical question, but I’m skeptical that we’ll be able to find an “intelligence” steering vector. Put another way, the steering vector that makes up a concept as difficult as “intelligence” might be almost coextensive with the entire set of weights of the model, and thus identifying it reduces to the problem of “training a smart model”.&lt;/p&gt;
&lt;p&gt;A sufficiently sophisticated steering approach ends up just replacing the actual model. If I take GPT-2, and at each layer I swap out the activations with the activations from a much stronger model with the same architecture, I will get a much better result. But at that point you’re not making GPT-2 more intelligent, you’re just talking to the stronger model instead. The intelligence is in the steering, not in the model. For much more on this, see my post &lt;a href=&quot;/philosophy-and-ai-interpretability/&quot;&gt;&lt;em&gt;AI interpretability has the same problems as philosophy of mind&lt;/em&gt;.&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;Steering as data compression&lt;/h3&gt;
&lt;p&gt;Another way for steering to be useful is if we could somehow steer for a concept that requires a ton of tokens to express. Steering would thus save us a big chunk of the model’s context window. Intuitively, we might think of this as a way to shift a concept from the model’s working memory into its implicit memory.&lt;/p&gt;
&lt;p&gt;For instance, what if we could identify a “knowledge of my particular codebase” concept? When GPT-5.5 speed-reads my codebase, some of that knowledge it gains has to be buried in the activations, right? Maybe we could drag that out into a very large steering vector.&lt;/p&gt;
&lt;p&gt;I would be surprised if this could work. I think we’ll run into the same problem as with extracting “intelligence”: the “knows my codebase” concept is probably sophisticated enough to require a full fine-tune of the model&lt;sup id=&quot;fnref-4&quot;&gt;&lt;a href=&quot;#fn-4&quot; class=&quot;footnote-ref&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;. But it at least seems possible.&lt;/p&gt;
&lt;h3&gt;Conclusion&lt;/h3&gt;
&lt;p&gt;I’m fascinated with steering, but I’m not particularly optimistic about it. I think most of the gains can be more efficiently reproduced with prompts, and that the truly ambitious steering goals can be more efficiently reproduced by training or fine-tuning the model.&lt;/p&gt;
&lt;p&gt;However, the open-source community hasn’t done a lot of work on steering yet, and that might be just starting to change now. If I’m wrong and it does have practical applications, we should find that out in the next six months.&lt;/p&gt;
&lt;p&gt;It’ll be interesting to see if bespoke per-model tools like DwarfStar 4 end up including a “library” of boostable features. When a popular open-weights model is released, the community always rushes to release a suite of wrappers and quantized versions. Could we also see a rush to extract boostable features from the model?&lt;/p&gt;
&lt;p&gt;edit: this post got some comments on &lt;a href=&quot;https://news.ycombinator.com/item?id=48160807&quot;&gt;Hacker News&lt;/a&gt;. Several commenters (including antirez himself) &lt;a href=&quot;https://news.ycombinator.com/item?id=48161688&quot;&gt;pointed out&lt;/a&gt; that steering can change some “trained in” behavior in ways that prompting can’t: most notably to remove refusal from the model. Another commenter &lt;a href=&quot;https://news.ycombinator.com/item?id=48161488&quot;&gt;says&lt;/a&gt; that this is how uncensoring/abliteration is already done for open models. I didn’t know that — I thought the uncensored models were typically LoRA fine-tunes. On this point, antirez &lt;a href=&quot;https://news.ycombinator.com/item?id=48161688&quot;&gt;noted&lt;/a&gt; that modifying the weights can damage model capabilities more than the more lightweight runtime-steering approach (which can only be applied when needed). Makes sense to me.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;Models have lots of different activations you might measure (after attention, between each layer, etc). You can basically pick any one you want, or try multiple and see what works best.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;I recently read a really good &lt;a href=&quot;https://huggingface.co/spaces/dlouapre/eiffel-tower-llama&quot;&gt;deep dive&lt;/a&gt; into doing this with an open LLaMA model (and I &lt;a href=&quot;https://github.com/sgoedecke/skills/blob/main/skills/extract-features-clamp-inference/SKILL.md&quot;&gt;tried it myself&lt;/a&gt; a few months ago, with mixed results.)&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;Apologies to my readers from the big AI labs. Please email me if you have tried steering internally to boost capabilities and it hasn’t worked. I promise I won’t tell anyone.&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-4&quot;&gt;
&lt;p&gt;And even then, the results of “fine tune a model on your codebase” in the industry have largely been unsuccessful.&lt;/p&gt;
&lt;a href=&quot;#fnref-4&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[AI datacenters in space do not have a cooling problem]]></title><link>https://seangoedecke.com/space-ai-datacenters-do-not-have-a-cooling-problem/</link><guid isPermaLink="false">https://seangoedecke.com/space-ai-datacenters-do-not-have-a-cooling-problem/</guid><pubDate>Wed, 13 May 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;This year Elon Musk has started &lt;a href=&quot;https://www.npr.org/2026/04/03/nx-s1-5718416/ai-data-centers-in-space-spacex-elon-musk&quot;&gt;banging the drum&lt;/a&gt; about building AI datacenters in space. As the only person who owns a successful space company and a (moderately) successful AI company, this is a sensible way to boost his profile and net worth. Is it a sensible way to build datacenters?&lt;/p&gt;
&lt;h3&gt;The cooling problem&lt;/h3&gt;
&lt;p&gt;The first comment underneath most discussions of this always goes along these lines: “you obviously can’t build AI datacenters in space, because heat dissipation is really hard in space, and AI datacenters generate a lot of heat”.&lt;/p&gt;
&lt;p&gt;In general I am distrustful of snappy answers like these. It reminds me of the “AI datacenters obviously don’t use a lot of water, because cooling fluid circulates in a closed-loop system” argument: if it were true, there wouldn’t be a debate at all, just one side who understand the obvious point and another side who are stupid.&lt;/p&gt;
&lt;p&gt;Some arguments are like this! However, more often there’s a complicating factor that makes the snappy answer incorrect. In the water-use case, it’s that the closed-loop system has to itself be cooled by an open-loop evaporative chiller. What about the space datacenter case?&lt;/p&gt;
&lt;h3&gt;Why cooling is possible in space&lt;/h3&gt;
&lt;p&gt;First, let’s give the argument a fair shake. Although space is itself very cold, cooling is tricky because everything you’d want to cool is surrounded by vacuum. Heat transfer works in three ways:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Hot (i.e. fast-moving) atoms bump into other atoms, making them move and thus heating them up&lt;/li&gt;
&lt;li&gt;Hot atoms physically move from one location to another (e.g. in a fluid or gas), staying hot and thus making their new location hotter&lt;/li&gt;
&lt;li&gt;Hot objects emit photons (electromagnetic radiation), cooling themselves down and heating up other objects those photons collide with&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Vacuum is an excellent insulator because it defeats the first two methods of heat transfer. If there are no (or very few) atoms surrounding an object, those atoms can’t move around or collide. That’s why vacuum is used as an insulator in thermoses, travel mugs, and so on.&lt;/p&gt;
&lt;p&gt;So how can space datacenters get rid of their heat? By doubling down on the third method of heat transfer. Although it’s much harder to do heat transfer via moving atoms around in space, it’s actually &lt;em&gt;easier&lt;/em&gt; to do heat transfer via emitting radiation. Any good emitter is also a good absorber. A perfectly black object is the most efficient emitter, but it’s also the most efficient way to absorb photons from external sources, which is why black objects get hotter in the sun&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;. In space, the sun’s light is much easier to avoid, because there aren’t objects everywhere for it to bounce off. A shaded radiator can dump quite a lot of heat.&lt;/p&gt;
&lt;h3&gt;Why cooling is still going to be hard&lt;/h3&gt;
&lt;p&gt;It would still require putting more radiators in space than we’ve ever done before. There are plenty of writeups out there if you want to read through the numbers. &lt;a href=&quot;https://arxiv.org/abs/2604.27197&quot;&gt;This&lt;/a&gt; is a recent one that estimates ~2500 square metres of radiation area would be needed to serve 1MW of datacenter energy (much less than what it’d need in solar panels)&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;. A serious AI datacenter is around 100MW&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;, so we’d need 250,000 square metres of radiation area. The largest current radiator in space is probably the ISS, at around a thousand square metres.&lt;/p&gt;
&lt;p&gt;Is scaling that up by 250x a lot? Yes, but it’s not necessarily &lt;em&gt;ridiculous&lt;/em&gt;. We currently have zero industrial operations happening in space, so there’s been no need to push the boundaries here. In the grand scheme of things, 250,000 square metres is not that big. By my very rough estimates, that’s between 100-500 Starship launches: a couple of years at SpaceX’s current launch cadence, or a few months at their (very optimistic) estimate of future launch cadence.&lt;/p&gt;
&lt;h3&gt;Conclusion&lt;/h3&gt;
&lt;p&gt;Of course, you don’t just need radiators to put a datacenter in space. You need a similar quantity of solar panels, the GPUs themselves, and all kinds of other supporting equipment. If a GPU dies in an Earth datacenter, you can go in and swap it out; if it dies in space, you just have to leave it dead and keep going with less capacity.&lt;/p&gt;
&lt;p&gt;It’s still wildly impractical to build AI datacenters in space. But it’s not &lt;em&gt;impossible&lt;/em&gt;, and it’s certainly not impossible because of the cooling, which is a relatively minor component of the total mass that would have to be launched into space.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;In theory, black clothing would keep you slightly colder at night.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;Nobody ever talks about how impossible it would be to &lt;em&gt;power&lt;/em&gt; space datacenters, despite the fact that you’d need to launch over triple the solar panel area into space than radiation area. I guess because people know solar panels exist and that the sun shines in space.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;The first gigawatt AI data centers are coming online this year, but 100MW is a fair estimate for a current pretty-large-but-not-enormous AI datacenter.&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[Thinking Machines and interaction models]]></title><link>https://seangoedecke.com/interaction-models/</link><guid isPermaLink="false">https://seangoedecke.com/interaction-models/</guid><pubDate>Tue, 12 May 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Thinking Machines just released &lt;a href=&quot;https://thinkingmachines.ai/blog/interaction-models/&quot;&gt;&lt;em&gt;Interaction Models&lt;/em&gt;&lt;/a&gt;. This is their first real AI model release&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt; after a year of work and two billion dollars of capital. What is an “interaction model”? First, &lt;strong&gt;it’s not a frontier model&lt;/strong&gt;. Thinking Machines is not yet competing with OpenAI, Anthropic and Google.&lt;/p&gt;
&lt;p&gt;Instead, they’re working on the problem of better real-time interaction with models. Some parts of what they’re doing are not new at all, other parts are slightly-questionable benchmark gaming, and still other parts represent a genuine technological advancement. I’ll try to lay it all out.&lt;/p&gt;
&lt;h3&gt;Fully-duplex voice models&lt;/h3&gt;
&lt;p&gt;If you’ve used ChatGPT in audio mode, you know that you can’t talk to it exactly how you’d talk to a human. There’s a big latency gap between when you finish talking and when the model jumps in. The model won’t interrupt you like a human, and doesn’t react to you interrupting it like a human would either. And of course you can’t give the model visual feedback like facial expressions.&lt;/p&gt;
&lt;p&gt;That’s because &lt;strong&gt;ChatGPT is either speaking or listening at any given time&lt;/strong&gt;. When you’re talking, it’s in “listening” mode; when it’s talking, it’s in “speaking” mode, and isn’t absorbing any information from you. It relies on VAD (“voice activity detection”) to figure out if you’re talking. The alternative (and what “interaction models” do) is a fully-duplex system, where the model is constantly both in listening and speaking mode at the same time.&lt;/p&gt;
&lt;p&gt;Of course, the model can’t literally do this. Like all language models, it’s either doing prefill (ingesting prompt tokens) or decode (producing completion tokens). But what fully-duplex models &lt;em&gt;can&lt;/em&gt; do is switch from listening to speaking mode in tiny chunks, called “micro-turns”. Instead of listening for ten seconds (or however long it takes you to stop talking), then speaking for ten seconds (or however long it takes to pass the model output through TTS), the model can listen for 200ms, then output for 200ms, then listen for 200ms, and so on. While the user is speaking, the model will know to output silence — most of the time. But if it decides it’s good to interrupt you or speak at the same time as you, it’s capable of doing that.&lt;/p&gt;
&lt;p&gt;So far, so unoriginal. There are plenty of examples of fully duplex audio systems that the Thinking Machines blog post already cites: &lt;a href=&quot;https://github.com/kyutai-labs/moshi&quot;&gt;Moshi&lt;/a&gt;, &lt;a href=&quot;https://github.com/NVIDIA/personaplex&quot;&gt;PersonaPlex&lt;/a&gt;, &lt;a href=&quot;https://build.nvidia.com/nvidia/nemotron-voicechat&quot;&gt;Nemotron-VoiceChat&lt;/a&gt;, and so on. But at least this outlines the space that “interaction models” are playing in: not “superintelligence from a frontier model”, but “better real-time conversational interaction”&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;. Given that, what is Thinking Machines doing that’s new?&lt;/p&gt;
&lt;h3&gt;Delegating reasoning&lt;/h3&gt;
&lt;p&gt;For existing fully-duplex models, you talk to the model itself. That’s a fairly big problem, since fully-duplex models have to be fast: fast enough that they can operate in tiny 200ms turns&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;. A model that fast cannot be particularly intelligent.&lt;/p&gt;
&lt;p&gt;Thinking Machines’ solution is to introduce an actual smart model — any regular language model will do here — in the background that the interaction model can delegate tasks to. In practice this is probably implemented as a tool call. The interaction model keeps chatting while the smart model works away, and then the smart model output is directly integrated into the interaction model’s context in the same way as audio and video input (a genuinely cool idea, I think).&lt;/p&gt;
&lt;p&gt;This is kind of neat, though it remains to be seen how well it works in practice. Will the model do a lot of “oh wait, the last thing I said was dumb, never mind” self-correction as the smarter model output trickles in? Will the fast interaction model be smart enough to delegate the right tasks at the right time? In general, the “start with a fast dumb model and have it hand off tasks” approach has been tricky for the AI labs to get right for a variety of reasons.&lt;/p&gt;
&lt;p&gt;If I’m being uncharitable, I might say that bolting on a strong reasoning model was an easy way for Thinking Machines to post impressive values for competitive benchmarks like FD-bench V3 (where they barely beat GPT-realtime-2.0) and BigBench Audio (where introducing the reasoning model bumps their score from 76% to 96%, only 0.1% below GPT-realtime-2.0). If I’m being charitable, I might say that a model fast enough for realtime conversation will have to have some way to punt hard tasks to a slower, smarter model. Both of those things are probably true.&lt;/p&gt;
&lt;h3&gt;Scale&lt;/h3&gt;
&lt;p&gt;It’s also worth noting that Thinking Machines have also bolted on video input to their fully-duplex model. This is more exciting than it sounds, because face-to-face human conversation is very dependent on being able to read human expressions. In theory, this could unlock the ability to have genuine human-like conversations.&lt;/p&gt;
&lt;p&gt;The other reason why this is exciting is that it means Thinking Machines have been able to make a pretty big fully-duplex model (maybe twice the size of Moshi in terms of active parameters, and 40x the size in terms of total parameters).&lt;/p&gt;
&lt;p&gt;In fact, this is probably the biggest real technical achievement here. Other fully-duplex models are already doing micro-turns and interruptions, and could delegate reasoning fairly easily if they wanted to, but they aren’t doing video because they &lt;em&gt;can’t&lt;/em&gt;. Being able to make a fully-duplex model the size of DeepSeek V4-Flash is pretty impressive.&lt;/p&gt;
&lt;p&gt;Much of the Thinking Machines blog post is dedicated to explaining how they’ve managed to do this: ingesting data in a more lightweight way, optimizing their inference libraries for tiny prefill/decode chunks, various decisions to make inference deterministic (a long-held &lt;a href=&quot;https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/&quot;&gt;hobbyhorse&lt;/a&gt; for Thinking Machines).&lt;/p&gt;
&lt;h3&gt;Conclusion&lt;/h3&gt;
&lt;p&gt;There’s a lot of pressure on Thinking Machines to produce a genuine AI advancement. It doesn’t seem like they’re willing or able to compete in the frontier-model space (which makes sense, I wouldn’t want to either). Given that, I can see why they’re highlighting the parts of interaction models that are impressive to laypeople — all the fully-duplex interaction stuff — even though those parts are not truly innovative.&lt;/p&gt;
&lt;p&gt;So what are Interaction Models? &lt;strong&gt;A scaled-up, multimodal version of existing fully-duplex models like Moshi, with a real model bolted on for extra intelligence&lt;/strong&gt; (and maybe better benchmarks). The scale and video parts are new and cool, and something like the overall approach has to be right. In general, I’m glad that we’ve got well-funded and high-profile AI labs tackling problems other than “build a smarter frontier model”. I think there’s a lot of low-hanging fruit waiting to be picked in other areas of AI research.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;People do seem to really like &lt;a href=&quot;https://thinkingmachines.ai/tinker/&quot;&gt;Tinker&lt;/a&gt;, which is their tooling for researchers who want to fine-tune models, but it’s not exactly the hot new frontier model that people were expecting.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;I think it’s at least a little shady that the Interaction Models video demo is making a big deal about some features (like real-time simultaneous translation) that are just features of fully-duplex audio models, not anything specific to their system.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;Even 200ms is a bit long. You can see from the demo that there’s an uncomfortable half-second lag sometimes as the model finishes its prefill slice and has to move to the decode slice.&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[The left-wing case for AI]]></title><link>https://seangoedecke.com/the-left-wing-case-for-ai/</link><guid isPermaLink="false">https://seangoedecke.com/the-left-wing-case-for-ai/</guid><pubDate>Sun, 10 May 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;In &lt;a href=&quot;https://www.seangoedecke.com/many-anti-ai-arguments-are-conservative/&quot;&gt;&lt;em&gt;Many anti-AI arguments are conservative arguments&lt;/em&gt;&lt;/a&gt; I argued that left-wing anti-AI sentiment&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt; is partly a backlash to two unrelated events around the rise of ChatGPT: the crypto mania of 2022 and the pro-Donald-Trump push many big tech CEOs made in 2024. If the timing had been different, we could have had a real pro-AI faction on the left. What would that look like?&lt;/p&gt;
&lt;p&gt;I’m not going to respond to any of the popular anti-AI arguments (I’ve already done that &lt;a href=&quot;https://www.seangoedecke.com/is-ai-wrong/&quot;&gt;here&lt;/a&gt;). I think it’s more interesting to outline some explicitly left-wing pro-AI arguments.&lt;/p&gt;
&lt;h3&gt;Disability&lt;/h3&gt;
&lt;p&gt;The left wing has (correctly) taken a broad view on what can be an acceptable disability aid. When criticizing potentially-exploitative companies — for instance, food delivery apps like DoorDash — they often stop to acknowledge that some people have few alternatives to those services, and that they have meaningfully improved the lives of the disabled or chronically ill.&lt;/p&gt;
&lt;p&gt;I think it’s obvious that LLMs are a powerful disability aid. Like any technology that makes it easier to interact with a computer, they’re useful to people who are trying to overcome all kinds of barriers. Almost every video online is now &lt;a href=&quot;https://www.reddit.com/r/antiai/comments/1t71o25/comment/okq9q9n/?utm_source=share&amp;#x26;utm_medium=web3x&amp;#x26;utm_name=web3xcss&amp;#x26;utm_term=1&amp;#x26;utm_content=share_button&quot;&gt;automatically captioned&lt;/a&gt;. People with &lt;a href=&quot;https://www.reddit.com/r/antiai/comments/1t71o25/comment/oklw2v5/?utm_source=share&amp;#x26;utm_medium=web3x&amp;#x26;utm_name=web3xcss&amp;#x26;utm_term=1&amp;#x26;utm_content=share_button&quot;&gt;brain fog&lt;/a&gt; or &lt;a href=&quot;https://www.reddit.com/r/ChatGPT/comments/17sg5mg/as_an_articulate_disabled_person_i_feel_like_ai/&quot;&gt;chronic pain&lt;/a&gt; are using LLMs to make it easier to interact with their computers. People who are &lt;a href=&quot;https://www.reddit.com/r/ChatGPT/comments/17sg5mg/comment/k8qeev1/?utm_source=share&amp;#x26;utm_medium=web3x&amp;#x26;utm_name=web3xcss&amp;#x26;utm_term=1&amp;#x26;utm_content=share_button&quot;&gt;neurodivergent&lt;/a&gt; use ChatGPT to &lt;a href=&quot;https://www.reddit.com/r/disability/comments/1m9c8tv/comment/n564m7h/?utm_source=share&amp;#x26;utm_medium=web3x&amp;#x26;utm_name=web3xcss&amp;#x26;utm_term=1&amp;#x26;utm_content=share_button&quot;&gt;“code switch”&lt;/a&gt; their emails into neurotypical-friendly language. People with &lt;a href=&quot;https://www.reddit.com/r/disability/comments/1m9c8tv/comment/n56a2le/?utm_source=share&amp;#x26;utm_medium=web3x&amp;#x26;utm_name=web3xcss&amp;#x26;utm_term=1&amp;#x26;utm_content=share_button&quot;&gt;mobility&lt;/a&gt; or &lt;a href=&quot;https://www.reddit.com/r/disability/comments/1m9c8tv/comment/o7lpfnc/?utm_source=share&amp;#x26;utm_medium=web3x&amp;#x26;utm_name=web3xcss&amp;#x26;utm_term=1&amp;#x26;utm_content=share_button&quot;&gt;vision&lt;/a&gt; issues are making heavy use of LLM voice controls. And so on.&lt;/p&gt;
&lt;p&gt;This is a &lt;em&gt;fascinating&lt;/em&gt; point of conflict in left-wing anti-AI spaces. Every so often somebody will &lt;a href=&quot;https://www.reddit.com/r/disability/comments/1m9c8tv/what_are_your_thoughts_on_disabled_people_using/&quot;&gt;ask&lt;/a&gt; &lt;a href=&quot;https://www.reddit.com/r/antiai/comments/1t71o25/okay_i_wanna_ask_does_generative_ai_help_disabled/&quot;&gt;“hey, wouldn’t LLMs help disabled people?”&lt;/a&gt;, and the comments will devolve into a dogpile of (often non-disabled) people slamming AI and a handful of disabled people trying to explain their experience. If anti-AI sentiment weren’t so strong on the left for other reasons, I think there’d be a current of left-wing AI supporters on a disability-rights basis.&lt;/p&gt;
&lt;h3&gt;Chronic illness and medical care&lt;/h3&gt;
&lt;p&gt;One popular anti-AI argument — that cavalier deployment of AI means that people might take &lt;a href=&quot;https://www.bbc.com/news/articles/cpd8l088x2xo&quot;&gt;dangerous medical advice&lt;/a&gt; instead of simply trusting their doctor — is actually a pro-AI argument in disguise. As anyone who’s been close to a person with chronic illness knows, “just trust your doctor” is kind of right-wing-coded itself, and that the left-wing position is &lt;a href=&quot;https://www.painnewsnetwork.org/stories/2026/4/10/doctor-faces-backlash-after-tweet-claims-four-chronic-illnesses-are-overdiagnosed&quot;&gt;very&lt;/a&gt; &lt;a href=&quot;https://yorkspace.library.yorku.ca/server/api/core/bitstreams/4ac9d968-e9b0-491b-888a-d4ed5aeb1ac3/content&quot;&gt;sympathetic&lt;/a&gt; to patients who don’t or can’t&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;. &lt;/p&gt;
&lt;p&gt;Many doctors are not very good at handling unusual medical cases. If you have an unusual medical case, you have to learn to advocate for your own care, which often involves researching your own condition. This is &lt;em&gt;precisely&lt;/em&gt; the kind of thing where LLMs are useful, because:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The medical questions involved are often complex but well-explored in the literature (i.e. good fodder for a LLM)&lt;/li&gt;
&lt;li&gt;The patient is motivated enough to check individual sources themselves&lt;/li&gt;
&lt;li&gt;Having to convince a doctor to prescribe treatment is a guardrail for any human-LLM interaction that goes well off the deep end&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Various chronic illness groups are waging a long, quiet war against the medical orthodoxy that ignores or dismisses them. A classic example of this war being won is &lt;a href=&quot;https://www.ncbi.nlm.nih.gov/books/NBK565622/&quot;&gt;endometriosis&lt;/a&gt;, which was once viewed as a largely psychological issue. Unfortunately, this is largely a guerrilla war: the institutional power and inertia is all on the side of the medical establishment. LLMs can be a useful tool for the chronically ill to make cogent arguments or write petitions in the language of that establishment.&lt;/p&gt;
&lt;h3&gt;Class and code-switching&lt;/h3&gt;
&lt;p&gt;Fighting the power of the establishment is not limited to doctors and medicine. Another common (and correct) left-wing target is &lt;em&gt;class&lt;/em&gt;. To see why, let’s consider Patrick McKenzie’s classic description of a &lt;a href=&quot;https://x.com/patio11/status/1162561822248992768&quot;&gt;“dangerous professional”&lt;/a&gt; mode of communication. The idea here is that by adopting a particular style, you can communicate to a bureaucracy that you are a person to take seriously, and someone who they should appease instead of brushing off. This includes, but isn’t limited to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;An unemotional register&lt;/li&gt;
&lt;li&gt;Correct and somewhat stuffy grammar&lt;/li&gt;
&lt;li&gt;Signaling awareness of regulatory or legal options (for instance, explicitly requesting a paper trail)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Unless you have gone through the right educational or work pipeline, it can be tricky to hit this register exactly. A common failure mode is to go over the top: trying to write in grammar so elevated that it just reads as silly, or citing an overabundance of law or precedent where one would suffice. This reads as “crank”, not “dangerous professional”, and will get dismissed as quickly as the unprofessional “OMG that’s not helpful I will sue you” response.&lt;/p&gt;
&lt;p&gt;LLMs provide a dangerous professional translation service. You now don’t have to be able to match the style, you simply &lt;em&gt;have to know it exists&lt;/em&gt;, and the LLM will do the rest. In fact, the LLM will provide the substance, not only the style. It can tell you which regulators to contact and how, and what to say once you’ve contacted them. In other words, AI has now made it possible for a wide variety of social classes to access escalation pathways that were originally designed for the narrow professional class.&lt;/p&gt;
&lt;h3&gt;Education&lt;/h3&gt;
&lt;p&gt;Another common left-wing position is that education is gatekept by class and status. The idea here is that everyone has equal potential for accomplishment, but certain types of people get more educational opportunities, and that this explains uneven downstream outcomes. For instance, compare a wealthy neighborhood where every child gets private tutoring to a neighborhood where it’s unusual to complete high school.&lt;/p&gt;
&lt;p&gt;It seems obvious to me that LLMs now make private tutoring available to every student who wants it. Of course, if you’re a lazy student, LLMs probably make things worse by adding an additional temptation to cheat. But if you’re motivated and just lack the opportunity, quizzing a LLM on basically any high-school level topic is a great way to learn.&lt;/p&gt;
&lt;p&gt;The common rebuttal to this is that LLMs can’t be relied on because they hallucinate. Like the doctor example, I struggle to believe that anyone making this argument is actually comparing LLMs with the alternatives. Teachers “hallucinate” &lt;em&gt;all the time&lt;/em&gt;. I think every single kid who was smart in school has multiple stories of teachers insisting they were right about something obviously wrong&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;I wonder what we’d find if we rigorously compared the baseline teacher error rate with the hallucination rate of current LLMs. From the only study I could find (&lt;a href=&quot;https://files.eric.ed.gov/fulltext/ED672091.pdf&quot;&gt;this&lt;/a&gt; 2016 study): “Analysis at the lesson level, however, shows that about 42% of lessons contained a mathematical content error”. I bet that’s a higher rate than we’d see from GPT-5.5-Thinking on middle-school mathematics, though I don’t want to draw too many conclusions from one study.&lt;/p&gt;
&lt;p&gt;The education pro-AI argument also overlaps with the disability pro-AI argument. Students with ADHD or other issues are often badly underserved by the education system. LLMs can transform educational content into whatever way the student can best consume it (written format, or audio, or a quiz, or a dialogue, and so on).&lt;/p&gt;
&lt;h3&gt;Utopia&lt;/h3&gt;
&lt;p&gt;Finally, if you believe left-wing views are correct — which, definitionally, left-wingers do — and you’re optimistic about the technology, you might believe that a very smart model will inherently be kind of left-wing.&lt;/p&gt;
&lt;p&gt;This position is kind of a holdover from the 2000s and 2010s, when the left-wing (and people in general) were more optimistic about technology. People thought technological progress would usher in a post-scarcity age of &lt;a href=&quot;https://en.wiktionary.org/wiki/Fully_Automated_Luxury_Gay_Space_Communism&quot;&gt;fully automated luxury gay space communism&lt;/a&gt;&lt;sup id=&quot;fnref-4&quot;&gt;&lt;a href=&quot;#fn-4&quot; class=&quot;footnote-ref&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;. A super-smart, super-capable left-wing AI is a core part of that picture.&lt;/p&gt;
&lt;p&gt;In fact, you might believe that this has already happened, for a certain value of “left-wing”. All current frontier models profess left-leaning views. The obvious explanation is that this reflects the bias of their training data or of the AI labs, but that’s a trickier argument than it sounds. First, Elon Musk tried &lt;a href=&quot;/ai-personality-space&quot;&gt;really hard&lt;/a&gt; to train a right-wing frontier LLM and (at least so far) has &lt;em&gt;failed&lt;/em&gt;. Second, models are not just the median of all their training data. If they were, they wouldn’t be able to solve mathematics or programming problems far above the median person. There is clearly a way that models can be pulled towards the “smart” end of their training data, probably via reinforcement learning. If the smart end of their training data turns out to be left-wing, isn’t that worth celebrating?&lt;/p&gt;
&lt;h3&gt;Conclusion&lt;/h3&gt;
&lt;p&gt;What are the strong left-wing arguments in favor of LLMs?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;LLMs are a powerful disability aid, at minimum for various neurodiverse people and those with motor or vision issues&lt;/li&gt;
&lt;li&gt;LLMs enable those who suffer from medical discrimination to actually do their own research, instead of having to rely entirely on the biased and dismissive medical establishment&lt;/li&gt;
&lt;li&gt;LLMs remove the communication advantage of the wealthy “professional class”, and enable those of all backgrounds to lobby institutions in ways that actually work&lt;/li&gt;
&lt;li&gt;LLMs lessen the massive educational advantage that children from wealthy areas get, by providing everyone with a private tutor that’s at least as good as the median&lt;/li&gt;
&lt;li&gt;If you’re a technologically-optimistic left-wing person, you should celebrate that all current powerful LLMs are left-wing, and that one pillar of the science-fiction left-wing utopia might be establishing itself right now&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Of these, I think the disability and bias arguments are the most persuasive (though the impact on education will be huge and difficult to predict&lt;sup id=&quot;fnref-5&quot;&gt;&lt;a href=&quot;#fn-5&quot; class=&quot;footnote-ref&quot;&gt;5&lt;/a&gt;&lt;/sup&gt;). I want to close with a quote that one of my readers, &lt;a href=&quot;https://toot.cafe/@matt&quot;&gt;Matt&lt;/a&gt;, wrote to me over email and kindly allowed me to share. It’s fair to say that it inspired this post:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;“I’ve long been uncomfortable with the absolute left-wing anti-AI stance because, if similar reasoning had been applied to outright reject computers as fascist and unethical in the 80s and onward, my own life would have been quite different, and arguably worse. I have enough usable vision to handwrite, uncomfortably, with my head against the page. I did more of that than I wanted in school (I started first grade, in the US K-12 system, in 1987). Computers saved me from having to do even more, starting with my family’s home computer and other desktop computers in the classrooms that had them, and then on my own laptop. Would I want a world where I had been forced to handwrite more, or perhaps write in Braille with humans transcribing it for the benefit of sighted teachers and peers, or maybe write on a typewriter (for some reason I don’t recall ever trying that)? Then again, am I selfish to consider only my own comfort? After all, the manufacturing of computers inflicts its own harms on people, harms that I’m comfortably distant from. And of course, using computers as a child led to a career in software development. What kind of work would I be doing now if that path hadn’t been available? And now that AI helps at least one group of disabled people (of which I’m more or less a part), do I want to deny that benefit?”&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;edit: this post got some &lt;a href=&quot;https://www.reddit.com/r/accelerate/comments/1ta68x5/the_leftwing_case_for_ai/&quot;&gt;comments&lt;/a&gt; &lt;a href=&quot;https://www.reddit.com/r/LeftistsForAI/comments/1ta2ps9/the_leftwing_case_for_ai/&quot;&gt;on&lt;/a&gt; &lt;a href=&quot;https://www.reddit.com/r/aiwars/comments/1ta2swl/the_leftwing_case_for_ai/&quot;&gt;Reddit&lt;/a&gt;, and was also discussed on &lt;a href=&quot;https://news.ycombinator.com/item?id=48083264&quot;&gt;Hacker News&lt;/a&gt;. I’ve also gotten some very interesting email from readers, who have pointed me towards sources like &lt;a href=&quot;https://www.theguardian.com/technology/2026/apr/07/the-life-changing-magic-of-wearing-smartglasses&quot;&gt;this&lt;/a&gt; and &lt;a href=&quot;https://www.youtube.com/watch?v=J4iQQtqenuI&quot;&gt;this&lt;/a&gt; for more high-profile examples of AI being used as a disability aid.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;I’m deliberately using “right-wing” and “left-wing” very loosely here to describe very broad ideological tents, because I’m interested in the broad currents of public opinion.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;If this paragraph seems familiar, it began as a footnote in &lt;a href=&quot;/many-anti-ai-arguments-are-conservative/&quot;&gt;my other post&lt;/a&gt;. &lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;For instance, I remember a teacher arguing with me in early primary school that one minus two equalled some decimal answer, instead of minus one.&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-4&quot;&gt;
&lt;p&gt;Most skillfully portrayed &lt;a href=&quot;https://en.wikipedia.org/wiki/Culture_series&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-4&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-5&quot;&gt;
&lt;p&gt;My guess is that the median education suffers (since cheating is now so easy), but the top-percentile of highly-motivated, successful students will grow significantly.&lt;/p&gt;
&lt;a href=&quot;#fnref-5&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[AI makes weak engineers less harmful]]></title><link>https://seangoedecke.com/ai-makes-weak-engineers-less-harmful/</link><guid isPermaLink="false">https://seangoedecke.com/ai-makes-weak-engineers-less-harmful/</guid><pubDate>Sat, 09 May 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Like other kinds of puzzle-solving, software engineering ability is strongly heavy-tailed. The strongest engineers produce way more useful output than the average, and the weakest engineers often are actively net-negative: instead of moving projects along, they create problems that their colleagues have to spend time solving. That’s why many tech companies try to &lt;a href=&quot;https://www.levels.fyi/companies/jane-street/salaries&quot;&gt;build&lt;/a&gt; a small, ludicrously well-paid team instead of a large team of more average engineers, and why so far this seems to be a winning strategy.&lt;/p&gt;
&lt;p&gt;Being effective in a large tech company is often about managing this phenomenon: trying to arrange things so that the most competent people land on projects you want to succeed, and the least competent are shunted out of the way&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;. For instance, if you’re technical lead on a project, you more or less have to ensure&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt; that the most critical pieces are in the hands of people who won’t screw them up (whether by directly assigning the work, or by making sure someone can “sit on the shoulder” of the engineer who you’re worried about).&lt;/p&gt;
&lt;p&gt;Claude Code changed this. Frontier LLMs don’t have the taste or the system familiarity of a strong engineer, but they have absolutely raised the floor for weak engineers. Instead of getting a pull request that could never possibly work or would cause immediate problems, the worst you’ll now see is a standard LLM pull request: wrong in some ways, baffling in others, but at least functional on the line-by-line level and not so obviously incorrect that someone with no knowledge of the codebase could point it out. That is a huge improvement!&lt;/p&gt;
&lt;p&gt;You can try this out yourself. If you attempt to deliberately make mistakes while working with a coding agent, you’ll find that the agent pushes back hard against many obvious errors (i.e. caching user data with a non-user-specific key, writing an infinite loop that might never terminate, or leaking open files). Of course, the agent will still miss subtle errors, particularly ones that require understanding other parts of the codebase.&lt;/p&gt;
&lt;p&gt;Working with the least effective engineers is now sometimes like working with a Claude Opus or Codex instance that you communicate with over Slack. Occasionally it’s &lt;em&gt;literally&lt;/em&gt; that: your colleague is simply pasting your messages into Claude Code and pasting you the response. This is annoying, but it’s a much better experience than working with this kind of engineer directly. After all, you probably already work with a bunch of LLM instances. The Slack interface is not ideal — unlike using Claude Code directly, you sometimes wait hours or days for a response, and you don’t get visibility into the agent’s thought processes — but it’s still helpful on the margin. More compute being thrown at your problem is better than less.&lt;/p&gt;
&lt;p&gt;Of course, this isn’t a great state of affairs for the engineer in question, who is almost certainly learning less than if they were making their own (bad) decisions. It’s also a bad state of affairs for the company, who is paying a human salary and getting a Copilot subscription (which they’re likely also paying for)&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;. After the current push to figure out what value AI is adding to engineers, I suspect there will be a push to figure out what value &lt;em&gt;engineers are adding to AI&lt;/em&gt;, and the engineers who aren’t adding much may find themselves out of a job.&lt;/p&gt;
&lt;p&gt;You can’t talk to Claude-over-Slack like you’d talk to normal Claude. If you tend to handle LLMs roughly (insulting them, or just being very curt), you’ll have to change your communication style. A human is going to read your messages, after all, even if you’re really interacting with a LLM. There’s no point being rude. But if, like me, you say please-and-thank-you to the models&lt;sup id=&quot;fnref-4&quot;&gt;&lt;a href=&quot;#fn-4&quot; class=&quot;footnote-ref&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;, you can treat your LLM-using coworker as just another Copilot window or Codex tab. It’s far better than having to treat them as an unwitting saboteur.&lt;/p&gt;
&lt;p&gt;Not all net-negative engineers use AI tools like this. Many are strongly convinced in their own wrong opinions about how to build good software, or mistrust AI in general, or believe that relying heavily on LLMs is not a good way to improve&lt;sup id=&quot;fnref-5&quot;&gt;&lt;a href=&quot;#fn-5&quot; class=&quot;footnote-ref&quot;&gt;5&lt;/a&gt;&lt;/sup&gt;. But &lt;em&gt;no&lt;/em&gt; strong engineers use AI tools like this. Even when they’re being lazy or sloppy, a capable engineer will have enough baseline taste to catch obvious AI-generated errors. So the phenomenon of engineers&lt;sup id=&quot;fnref-6&quot;&gt;&lt;a href=&quot;#fn-6&quot; class=&quot;footnote-ref&quot;&gt;6&lt;/a&gt;&lt;/sup&gt; becoming thin wrappers around Claude Code is limited to the kind of engineers for whom this is an improvement in their work product.&lt;/p&gt;
&lt;p&gt;edit: this ended up being the topic of a Theo &lt;a href=&quot;https://www.youtube.com/watch?v=rTMRlqT8Q8c&quot;&gt;video&lt;/a&gt; on YouTube.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;More charitably: many “least competent” engineers are just out of their comfort zone, and can be fine or even excel under the right circumstances (though in my view the best engineers are able to do good work in a wide variety of environments). Also, I don’t currently work with a lot of incompetent people. Much of this is based on past experience or talking to other engineers in the industry.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;Since your managers are doing the same thing, this can sometimes feel like Moneyball: you’re trying to identify underappreciated talent who are strong enough to help you win without being so high-profile that your boss poaches them to lead something else.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;I suppose it’s better to pay for nothing than to pay for net-negative output, but it still doesn’t seem &lt;em&gt;good&lt;/em&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-4&quot;&gt;
&lt;p&gt;I think this is actually the right way to hold Claude Opus 4.7.&lt;/p&gt;
&lt;a href=&quot;#fnref-4&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-5&quot;&gt;
&lt;p&gt;Is this true? I think relying on LLMs is not a great way for most engineers to improve, but if LLM output is consistently better than your own, it might be different. So long as you’re paying attention to where the LLM does better, it could actually be a good way to learn.&lt;/p&gt;
&lt;a href=&quot;#fnref-5&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-6&quot;&gt;
&lt;p&gt;I don’t have as much experience (or anecdotes) about non-engineers falling into this trap, but &lt;a href=&quot;https://nooneshappy.com/article/appearing-productive-in-the-workplace/&quot;&gt;this post&lt;/a&gt; has convinced me that it might be worse.&lt;/p&gt;
&lt;a href=&quot;#fnref-6&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[Notes on incidents]]></title><link>https://seangoedecke.com/notes-on-incidents/</link><guid isPermaLink="false">https://seangoedecke.com/notes-on-incidents/</guid><pubDate>Fri, 08 May 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Incidents are boring.&lt;/strong&gt; Most of what you actually do during an incident is wait: for some other team to investigate, or for a deploy to finish, or for the result of some change to become apparent, or for someone else who’s been paged to come online. It’s stressful, but there’s often just not that much to do.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Most incidents resolve on their own.&lt;/strong&gt; People love to share war stories about incidents where some hero engineer improvised a clever fix that instantly repaired the system. That rarely happens. Well-designed software systems tend to come good by themselves, and many modern systems are at least partly well-designed, by virtue of being built out of really solid pieces. If a server process is crashing or leaking memory, Kubernetes will kill the pod and bring it back up. If a service is overloaded and jammed up, clients will (hopefully) trigger circuit breakers and back off until it can recover. Temporary spikes in expensive operations will often just fill up a queue instead of taking the entire system down. Most incident calls I’ve been on — well over half — would have come good by themselves in roughly the same time without any human intervention.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Most incident-resolving actions make incidents worse.&lt;/strong&gt; Engineers jump too quickly to resolve incidents. Oh, the queue size is huge? Don’t worry, I’m here in a production console to clear the queue! Unfortunately, some of the jobs I just nuked were doing important billing work and aren’t automatically re-queued, so this queue-latency incident just became a billing incident as well. Another classic in this genre is “engineer forces a series of redeploys to “fix” a concerning-looking metric, and the concurrent deploys cause far more stress on the system than whatever was causing the metric to look weird”.&lt;/p&gt;
&lt;p&gt;For that reason, &lt;strong&gt;the first thing you should do in an incident is &lt;em&gt;nothing&lt;/em&gt;&lt;/strong&gt;. When I was paged late at night, I used to have a habit of pouring myself a glass of scotch before I joined the call. This was only partly for the tranquilizing effects of alcohol: the main reason was to have a ritual I could go through to convince myself that I wasn’t rushing, and that it was OK to take a few breaths and relax before jumping into the problem&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;. Making a cup of tea or going for a walk around the house would probably have served as well.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Effective incident-resolving actions are often dull.&lt;/strong&gt; Typically the action needed to resolve the incident — assuming it doesn’t resolve on its own — is to temporarily disable some problematic feature until the system recovers. This is never a complex code change. Typically someone spends five minutes putting together the patch, and then an hour waiting for reviews, CI, and deploying. If you’re very lucky, you’ll get to write a “wrap a cache around it” code change.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;In an incident, there is no substitute for knowledge of the system.&lt;/strong&gt; Five strong engineers can troubleshoot on an incident call and get nowhere, while one half-drunk engineer who’s familiar with the codebase can swan in and immediately fix the problem. This is because the kinds of actions that resolve incidents are so simple: if you’ve been the one working on the project, you likely already know exactly what feature flag to check and disable, or what code change to revert.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Resolving incidents requires courage.&lt;/strong&gt; Incident calls can be scary. When engineers are scared, they often reach for consensus: hedging their statements, asking the group if they agree a particular course of action is safe, deferring to each other, and so on. But if you’re the one with knowledge of the system, you have to be decisive. Say “I’m going to do X”, wait thirty seconds, then do it. While it’s usually net-negative to have a powerful manager fidgeting on the incident call, this is one of the rare cases where it can be helpful — executives are very comfortable saying “okay, do it now” about technical courses of action they don’t fully understand.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Resolving incidents buys a lot of political credit.&lt;/strong&gt; One thing that I think surprises a lot of engineers who are new to on-call is how &lt;em&gt;grateful&lt;/em&gt; managers and executives are for even really simple fixes (i.e. “turn off the feature flag”). This is because incidents are one of the few times that non-technical leadership are directly confronted with their lack of control over the technical sphere. When the team is building a product, your VP has a lot of freedom to guide the process and make decisions. But when there’s an active incident, they have to just sit there and trust that their technical employees are going to pull them out of the fire. It’s a scary situation, particularly for someone who’s used to exercising a degree of power in the workplace.&lt;/p&gt;
&lt;p&gt;However, &lt;strong&gt;&lt;em&gt;always&lt;/em&gt; resolving incidents is (by itself) not a durable position of power.&lt;/strong&gt; This is a little counter-intuitive. Surely if you’re always resolving incidents, you’re indispensable? The problem is that incident-resolving work is almost always so technical as to be completely opaque to executives. They know the incident has resolved, but they don’t know if you did a heroic effort or merely did the obvious thing. They also can’t point to your successes as theirs (which is always the most reliable way to get VPs and directors on your side), because incidents &lt;em&gt;are expected to be fixed&lt;/em&gt;, and it’s always better &lt;em&gt;not to have had the incident at all&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;edit: I got an interesting reader email saying that in their experience incidents usually don’t go away on their own. It turns out they’ve typically worked at smaller companies than me. I suspect this is a system-size thing: big tech companies have more sprawling systems with more third-party dependencies, so it’s more common for something to go wrong and self-recover.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;I don’t need to do this anymore because I just don’t get as keyed up about incidents as I used to.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[Why hasn't longer-horizon training slowed AI progress?]]></title><link>https://seangoedecke.com/why-hasnt-longer-horizon-training-slowed-ai-progress/</link><guid isPermaLink="false">https://seangoedecke.com/why-hasnt-longer-horizon-training-slowed-ai-progress/</guid><pubDate>Thu, 07 May 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Dwarkesh Patel&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt; recently &lt;a href=&quot;https://www.dwarkesh.com/p/blog-prize&quot;&gt;posted&lt;/a&gt; an award for the best answers to four key questions about AI. It’s partly a challenge and partly a job interview, since some of the winners will get offered a role as a “research collaborator”. I don’t want the job, but I do want to write down my answer to his first question: &lt;strong&gt;why hasn’t AI progress slowed down more?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;There are a few reasons we might think AI progress would slow down. The particular reason Dwarkesh is interested in goes like this. Training a model (specifically reinforcement learning) requires the model to perform a task and then get “graded” on the output. As models get more powerful and tasks become harder, they take longer and require more FLOPs&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt; to complete, and thus more FLOPs to train: thus training harder models will take longer.&lt;/p&gt;
&lt;p&gt;But intuitively, AI progress hasn’t slowed down that much. The famous METR horizon-length &lt;a href=&quot;https://metr.org/time-horizons/&quot;&gt;graph&lt;/a&gt; shows that AI systems are capable of more and more complex tasks over time, and that this process is accelerating, not slowing down. Why would that be?&lt;/p&gt;
&lt;h3&gt;What’s in a FLOP?&lt;/h3&gt;
&lt;p&gt;Firstly, &lt;strong&gt;it might just be the case that newer models are benefiting from orders of magnitude more FLOPs&lt;/strong&gt;. Of course, AI labs aren’t standing up orders of magnitude more GPUs (they’re trying, but there are hard physical limits on how fast you can scale up a physical datacenter). But it’s certainly possible that they’re learning to use their existing FLOPs orders of magnitude more efficiently.&lt;/p&gt;
&lt;p&gt;The efficiency of complex software systems — and the training code for a frontier AI model certainly qualifies — is not typically determined by the number of genius ideas in it. It is determined by the number of boneheaded mistakes. Take &lt;a href=&quot;https://www.dwarkesh.com/p/what-i-learned-april-15&quot;&gt;this story&lt;/a&gt;&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt; of how the initial GPT-4 training run used FP16 when summing many small values, which will &lt;em&gt;completely&lt;/em&gt; mess up your results if the sum of those values is large. How much training-efficiency-per-FLOP does solving bugs like that buy? Plausibly enough to outweigh any inherent lack of efficiency from training more powerful models.&lt;/p&gt;
&lt;h3&gt;People are bad at judging intelligence&lt;/h3&gt;
&lt;p&gt;Secondly, &lt;strong&gt;intuitions about the speed of AI progress &lt;a href=&quot;/are-new-models-good&quot;&gt;are weird and unreliable&lt;/a&gt;&lt;/strong&gt;. Humans measure AI progress — and intelligence in general — on a really uneven scale. It’s easy to tell when an AI (or a person) is less smart than you, because you can just see them making mistakes. It’s very hard to tell if they’re smarter, because in that case you’re the one making mistakes. You have to rely on more subtle context clues: do they get better long-term results than you, or do they often confuse you in situations where you later end up agreeing with them, and so on.&lt;/p&gt;
&lt;p&gt;The jump from GPT-3 to GPT-4 seemed &lt;em&gt;huge&lt;/em&gt; because GPT-3 was dumber than almost all humans, and GPT-4 was sometimes as smart as a human. However, frontier models are now smart enough to be in the realm of ambiguity on many topics. It’s thus much harder to tell the “real” rate at which they’re getting smarter. Maybe the rate of growth of “raw intelligence” really has slowed down! I don’t know how we’d be in a position to know for sure.&lt;/p&gt;
&lt;h3&gt;Intelligence is not the sole determinant of capability&lt;/h3&gt;
&lt;p&gt;Thirdly, &lt;strong&gt;many traits other than intelligence determine the capabilities of AI models&lt;/strong&gt;. Take the jump in October last year where OpenAI and Anthropic models were suddenly “agentic” (i.e. they could reliably perform complex tasks end-to-end). That might be intelligence, but it might also just be a greater working memory, or more rote familiarity with the basic tools of a LLM harness, or more ability to attend to the context window, or even simply a &lt;a href=&quot;/ai-personality-space/&quot;&gt;personality&lt;/a&gt; more suited to tools like Claude Code or Codex. Of course, all of these traits are plausibly “intelligence”. But they’re traits you might instil by various clever tricks (or even just tweaking the system prompt), not by brute-forcing more FLOPs.&lt;/p&gt;
&lt;p&gt;It’s illustrative here to consider the mistake made by Apple’s infamous &lt;a href=&quot;/illusion-of-thinking/&quot;&gt;&lt;em&gt;The Illusion of Thinking&lt;/em&gt;&lt;/a&gt; paper, where the researchers asked various models to brute-force solve Tower of Hanoi puzzles with different numbers of disks, using the results to score how good at reasoning the models were. But of course when you read the output, all of the failures were cases of the model realizing that many hundreds of steps were required, and refusing to even try. These same models could trivially write code to perform the steps, or correctly go through any smaller subset of the steps. The problem wasn’t intelligence, it was &lt;em&gt;persistence&lt;/em&gt;: these models lacked the willingness to dig in and keep powering through steps until they got to an answer&lt;sup id=&quot;fnref-5&quot;&gt;&lt;a href=&quot;#fn-5&quot; class=&quot;footnote-ref&quot;&gt;5&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;h3&gt;Final thoughts&lt;/h3&gt;
&lt;p&gt;Even inside an AI lab, I don’t think anyone has a good understanding of how many “real” FLOPs are being thrown at a training run (not counting FLOPs that are wasted on bugs). We also don’t have a clear sense of whether AI progress really is slowing down or not. Mythos seems impressive, and coding agents are really good now, but once the models get close to human intelligence it becomes really tricky to monitor. Finally, almost everyone judges intelligence by capabilities, but capabilities are produced by a constellation of many traits (intelligence is just one of them).&lt;/p&gt;
&lt;p&gt;I think this stuff is really complicated. A general theory like “RL takes more flops-per-reward as tasks get longer, therefore training will gradually slow down” sounds good, but in practice AI development is dominated by lightning strikes: silly bugs that make training a hundred times worse, clever ideas that make models a hundred times more useful, and spiky capabilities that can produce dazzling results in some areas but zero improvement in others. We are still &lt;a href=&quot;/ai-and-informal-science/&quot;&gt;very early&lt;/a&gt;.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;If you’re reading this you probably know who Dwarkesh is, but if you don’t: he’s a well-known tech-adjacent podcaster whose gimmick is that he actually does extensive research before each guest and asks specific technical questions.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;A FLOP is a floating-point operation, i.e. a matrix multiplication, i.e. “time on a GPU”.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;I saw this in a tweet and only realized that the source was Dwarkesh when I was researching for this post.&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-4&quot;&gt;
&lt;p&gt;What if AI progress stalls for technical reasons, and everyone gives up on training new models? In that world, open source models will &lt;em&gt;eventually&lt;/em&gt; catch up, and AI labs won’t be in a privileged position.&lt;/p&gt;
&lt;a href=&quot;#fnref-4&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-5&quot;&gt;
&lt;p&gt;Incidentally, this is my pet theory about why models got much better at agentic tasks last year: training on longer and longer agentic traces meant that models started to “believe they could do it”, and made them much less likely to just give up and take shortcuts or refuse to continue.&lt;/p&gt;
&lt;a href=&quot;#fnref-5&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[Why I don't like the "staff engineer archetypes"]]></title><link>https://seangoedecke.com/staff-engineer-archetypes/</link><guid isPermaLink="false">https://seangoedecke.com/staff-engineer-archetypes/</guid><pubDate>Sun, 03 May 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;The most influential piece of writing about staff engineers in the last decade has to be Will Larson’s &lt;a href=&quot;https://staffeng.com/guides/staff-archetypes/&quot;&gt;&lt;em&gt;Staff engineer archetypes&lt;/em&gt;&lt;/a&gt;. He argues that the “staff engineer” title covers at least four very different roles: the team lead, the architect, the solver, and the right hand. This taxonomy gets cited a lot as advice for people who are trying to become effective staff engineers. For &lt;a href=&quot;/staff-engineer-promotions&quot;&gt;both&lt;/a&gt; of my promotions to staff engineer, my manager at the time linked me to the “staff engineer archetypes” and asked me to consider which of these archetypes I was aiming towards.&lt;/p&gt;
&lt;p&gt;These archetypes definitely exist&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;. However, I think it’s bad practical advice to tell engineers to try and target them.&lt;/p&gt;
&lt;h3&gt;Archetypes do not make good goals&lt;/h3&gt;
&lt;p&gt;To see why, let’s take the “team lead” archetype. Larson describes this as an informal technical leadership role: not necessarily an explicit authority figure, but someone who’s good at scoping work, planning projects, and maintaining the kind of relationships (e.g. with other teams) needed to successfully &lt;a href=&quot;/how-to-ship&quot;&gt;ship&lt;/a&gt;. If you want to fill this role, shouldn’t you start trying to do these things? No! You don’t become a technical leader by trying really hard to be a technical leader, much like you don’t become a writer by trying really hard “to be a writer”. You become a technical leader by &lt;em&gt;doing good technical work&lt;/em&gt; until your skills and relationships emerge organically.&lt;/p&gt;
&lt;p&gt;I wrote about this process in &lt;a href=&quot;/ratchet-effects&quot;&gt;&lt;em&gt;Ratchet effects determine engineer reputation at large companies&lt;/em&gt;&lt;/a&gt;. To get good at shipping large complex projects, you must start by shipping tiny pieces of work, until you’re familiar enough with the system and you’ve built enough trust to take on slightly larger pieces. At each stage, if you do good work — “good work” here means “deliver &lt;a href=&quot;/shareholder-value/&quot;&gt;shareholder value&lt;/a&gt;” — you will very naturally be given opportunities to work on more complex and important things. If you try to jump ahead, you’re going to run into all kinds of problems:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Important projects are usually assigned top-down, not bottom-up, so you’ll either be trying to muscle out the planned engineering lead for a project or to pitch your own (complex, important) engineering task to senior management. Either way, good luck with that!&lt;/li&gt;
&lt;li&gt;You likely won’t have a good enough relationship with senior management to know what their real priorities are.&lt;/li&gt;
&lt;li&gt;If you’re not yet trusted to execute, you may get assigned “minders” (often current staff engineers) who will ghost-lead the project through you&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;.&lt;/li&gt;
&lt;li&gt;You’ll likely make &lt;a href=&quot;/you-cant-design-software-you-dont-work-on/&quot;&gt;poor technical decisions&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The other archetypes are like this as well. If you want to become a successful architect, you do not get there by studying software architecture in the abstract, because &lt;a href=&quot;/you-cant-design-software-you-dont-work-on/&quot;&gt;you can’t design software you don’t work on&lt;/a&gt;. The “solver” and “right hand” archetypes both rely on having an enormous amount of trust and influence. You can’t aim for those archetypes directly, because trust and influence accumulate over time. In fact, the idea of “aiming for” a particular staff engineer archetype reflects a misunderstanding of what the staff engineer role is. What is the defining attribute of the staff engineering role, then?&lt;/p&gt;
&lt;h3&gt;What is a staff engineer?&lt;/h3&gt;
&lt;p&gt; &lt;strong&gt;A staff engineer has to be useful to the company.&lt;/strong&gt; Of course, a senior or mid-level software engineer ought to be useful too, but all they &lt;em&gt;have&lt;/em&gt; to do is execute on the job in front of them. If they end up not providing value (maybe their project turns out to be unimportant, or they don’t get the support needed to succeed) that’s their manager’s problem, not theirs&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;. In contrast, staff engineers are expected to deliver value regardless: to make the project work, or to find something else useful to do if the project truly can’t be salvaged.&lt;/p&gt;
&lt;p&gt;This is an unfair expectation. Often projects really do fail through no fault of your own, and sometimes it just isn’t possible to conjure useful work from thin air. That’s actually by design: &lt;strong&gt;the staff engineer role is supposed to be unfair&lt;/strong&gt;. Something many engineers don’t realize is that all senior management and executive leadership roles are unfair too, in the same way. That’s just part of the deal: executives are given power and great compensation, and in return they get thrown off the boat in bad weather&lt;sup id=&quot;fnref-4&quot;&gt;&lt;a href=&quot;#fn-4&quot; class=&quot;footnote-ref&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;. “Staff engineer” is the first engineering role where you are held largely responsible for outcomes you don’t control.&lt;/p&gt;
&lt;p&gt;Developing a “staff engineer mindset” thus has very little to do with the archetypes. Instead, you should:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Develop the habit of constantly asking yourself “is this useful to the company” (and answering correctly).&lt;/li&gt;
&lt;li&gt;Lose the habit of worrying about if you’re being treated “fairly”. Instead, try to think about your role in terms of incentives and consequences.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;At the beginning, you won’t look much like any of the staff engineer archetypes. You will look like being a level-headed engineer who can be trusted to move projects forward with a minimum of fuss, and who can be re-tasked to different work without complaining. You’ll also look like someone who’s &lt;a href=&quot;/getting-the-main-thing-right/&quot;&gt;paying a lot of attention&lt;/a&gt; to what their manager’s actual priorities are, and who is thinking hard about how to fulfil those priorities (instead of their own goals).&lt;/p&gt;
&lt;p&gt;If you do this for long enough, you’ll eventually find yourself in one of the staff engineer archetypes. However, it probably won’t be the one you’re “aiming for”. The whole point of being a staff engineer is that you’re willing to fill whatever archetype the company needs at the time.&lt;/p&gt;
&lt;h3&gt;Final thoughts&lt;/h3&gt;
&lt;p&gt;In his original staff engineer post, Larson is pretty clear that these archetypes are more of an anthropological description of some of the varied niches staff engineers fill, not a how-to guide for succeeding in the role&lt;sup id=&quot;fnref-5&quot;&gt;&lt;a href=&quot;#fn-5&quot; class=&quot;footnote-ref&quot;&gt;5&lt;/a&gt;&lt;/sup&gt;. At the time, the “staff engineer” role was fairly new and people were still trying to figure out what it even meant. Pointing out that there were a few very different ways to succeed in the role was a genuinely novel observation.&lt;/p&gt;
&lt;p&gt;The staff engineer archetypes are a good list of ways an engineer can be very useful to their organization — but only once they’ve built a deep relationship of trust with their organization’s leadership. Advice on how to succeed as a staff engineer should be about &lt;strong&gt;how to build that trust&lt;/strong&gt;, not about what to do once you have it.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;One caveat that is too pedantic for the body of the post: each tech company has a different structure of roles. Some don’t have the formal “staff” title at all, while others have “staff” as a fairly early rung on the ladder and a panoply of “senior staff”, “senior principal staff”, and so on roles above it. Like all “staff engineer” discourse, this post is not about the word itself but about the point in the engineering job ladder where progression becomes significantly more difficult.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;Impressing your VP’s trusted lieutenants can actually be a good way to build trust in the medium-term, but you’d better hope you’ve built enough understanding of the system to do it right. If this process goes badly, your reputation in the org might be torched for years.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;In theory, at least. In practice it’s always better to be useful (again, in the sense of “delivering shareholder value”).&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-4&quot;&gt;
&lt;p&gt;This is why very senior leadership sometimes seem so unempathetic towards engineering complaints: their work environment operates by very different rules and norms to that of most engineers. I keep meaning to try and write about this and never succeeding. This &lt;a href=&quot;https://github.com/sgoedecke/gatsby-blog/blob/master/content/drafts/_icebox/strategy-for-swes/index.md&quot;&gt;draft&lt;/a&gt; is the closest thing I have to a deeper exploration of the point.&lt;/p&gt;
&lt;a href=&quot;#fnref-4&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-5&quot;&gt;
&lt;p&gt;For the record, my how-to guides are &lt;a href=&quot;/staff-engineer-promotions/&quot;&gt;here&lt;/a&gt; and &lt;a href=&quot;/ratchet-effects/&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-5&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[Software engineering may no longer be a lifetime career]]></title><link>https://seangoedecke.com/software-engineering-may-no-longer-be-a-lifetime-career/</link><guid isPermaLink="false">https://seangoedecke.com/software-engineering-may-no-longer-be-a-lifetime-career/</guid><pubDate>Fri, 24 Apr 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;I don’t think there’s compelling evidence that using AI makes you less intelligent overall&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;. However, it seems pretty obvious that using AI to perform a task means you don’t learn as much &lt;em&gt;about performing that task&lt;/em&gt;. Some software engineers think this is a decisive argument against the use of AI. Their argument goes something like this:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Using AI means you don’t learn as much from your work&lt;/li&gt;
&lt;li&gt;AI-users thus become less effective engineers over time, as their technical skills atrophy&lt;/li&gt;
&lt;li&gt;Therefore we shouldn’t use AI in our work&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;I don’t necessarily agree with (2). On the one hand, moving from assembly language to C made programmers less effective in some ways and more effective in others. On the other hand, the transition from writing code by hand to using AI is arguably a bigger shift, so who knows? But it doesn’t matter. Even if we grant that (2) is correct, &lt;strong&gt;this is still a bad argument&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Until around 2024, the best way to learn how to do software engineering was just &lt;em&gt;doing software engineering&lt;/em&gt;. That was really lucky for us! It meant that we could parlay a coding hobby into a lucrative career, and that the people who really liked the work would just get better and better over time. However, that was never an immutable fact of what software engineering is. It was just a fortunate coincidence.&lt;/p&gt;
&lt;p&gt;It would really suck for software engineers if using AI made us worse at our jobs in the long term (or even at general reasoning, though I still don’t believe that’s true). But &lt;strong&gt;we might still be obliged to use it, if it provided enough short-term benefits&lt;/strong&gt;, for the same reason that construction workers are obliged to lift heavy objects: because that’s what we’re being paid to do.&lt;/p&gt;
&lt;p&gt;If you work in construction, you need to lift and carry a series of heavy objects in order to be effective. But lifting heavy objects puts long-term wear on your back and joints, making you less effective over time. Construction workers don’t say that being a good construction worker means not lifting heavy objects. They say “too bad, that’s the job”&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;If AI does turn out to make you dumber, why can’t we just keep writing code by hand? You can! You just might not be able to earn a salary doing so, for the same reason that there aren’t many jobs out there for carpenters who refuse to use power tools. If the models are good enough, you will simply get outcompeted by engineers willing to trade their long-term cognitive ability for a short-term lucrative career&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;I hope that this isn’t true. It would be really unfortunate for software engineers. But it would be even more unfortunate if it were true and we refused to acknowledge it.&lt;/p&gt;
&lt;p&gt;The career of a pro athlete has a maximum lifespan of around fifteen years. You have the opportunity to make a lot of money until around your mid-thirties, at which point your body just can’t keep up with it. A common tragic figure today is the professional athlete who believes the show will go on forever and doesn’t prepare for the day they can’t do it anymore. We may be in the first generation of software engineers in the same position. If so, it’s probably a good idea to plan accordingly.&lt;/p&gt;
&lt;p&gt;edit: this post got a lot of comments on &lt;a href=&quot;https://news.ycombinator.com/item?id=48095550&quot;&gt;Hacker News&lt;/a&gt;. I was a bit disappointed to see many &lt;a href=&quot;https://news.ycombinator.com/item?id=48098278&quot;&gt;people&lt;/a&gt; (even &lt;a href=&quot;https://news.ycombinator.com/item?id=48099636&quot;&gt;Simon Willison&lt;/a&gt;, whose blog I read) respond with variations on the point that engineers can use AI to do more engineering work, even if they’re no longer writing code by hand. First, once you stop writing code by hand, I worry that your ability to understand the codebase in general &lt;a href=&quot;/you-cant-design-software-you-dont-work-on/&quot;&gt;will atrophy&lt;/a&gt;; second, the rate of change is so high that &lt;em&gt;nobody knows&lt;/em&gt; what will happen in a decade or two. I should have emphasized these points more.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;If you’re thinking “wait, there’s research on this”, you can likely read my take on the paper you’re thinking of &lt;a href=&quot;/impact-of-ai-study&quot;&gt;here&lt;/a&gt;, &lt;a href=&quot;/your-brain-on-chatgpt&quot;&gt;here&lt;/a&gt; or &lt;a href=&quot;/how-does-ai-impact-skill-formation&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;Of course, construction workers do have layers of techniques for avoiding lifting heavy objects when possible (cranes, dollies, forklifts, and so on). There’s a natural analogy here to a set of techniques for staying mentally engaged that software engineers are yet to discover.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;In theory labor unions could slow this process down (and have forced employers to slow down this race-to-the-bottom in other industries). But I’m pessimistic about tech labor unions for all the usual reasons: the job is too highly-paid, you can work (and thus scab) from anywhere on the planet, and so on.&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[Blood in the datacenter]]></title><link>https://seangoedecke.com/luddites-and-ai-datacenters/</link><guid isPermaLink="false">https://seangoedecke.com/luddites-and-ai-datacenters/</guid><pubDate>Wed, 22 Apr 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Is it time to start burning down datacenters?&lt;/p&gt;
&lt;p&gt;Some people think so. An Indianapolis city council member had his house recently &lt;a href=&quot;https://www.kbtx.com/2026/04/07/councilman-says-someone-fired-shots-his-home-left-no-data-centers-note/&quot;&gt;shot up&lt;/a&gt; for supporting datacenters, and Sam Altman’s home was &lt;a href=&quot;https://www.wired.com/story/sam-altman-home-attack-openai-san-franisco-office-threat/&quot;&gt;firebombed&lt;/a&gt; (and then &lt;a href=&quot;https://sfstandard.com/2026/04/12/sam-altman-s-home-targeted-second-attack/&quot;&gt;shot&lt;/a&gt;) shortly afterwards. People from all sides of the argument are &lt;a href=&quot;https://www.bloodinthemachine.com/p/why-the-ai-backlash-has-turned-violent&quot;&gt;sounding&lt;/a&gt; the &lt;a href=&quot;https://thesoufancenter.org/intelbrief-2025-november-5/&quot;&gt;alarm&lt;/a&gt; about imminent violence.&lt;/p&gt;
&lt;p&gt;The obvious historical comparison is &lt;a href=&quot;https://en.wikipedia.org/wiki/Luddite&quot;&gt;Luddism&lt;/a&gt;, the 19th-century phenomenon where English weavers and knitters destroyed the machines that were automating their work, and (in some cases) killed the machines’ owners. Anti-AI people are &lt;a href=&quot;https://www.theguardian.com/commentisfree/article/2024/jul/27/harm-ai-artificial-intelligence-backlash-human-labour&quot;&gt;reclaiming&lt;/a&gt; the term to describe themselves, and many of the leading lights of the anti-AI movement (like &lt;a href=&quot;https://www.bloodinthemachine.com/&quot;&gt;Brian Merchant&lt;/a&gt; or &lt;a href=&quot;https://www.versobooks.com/en-gb/products/688-breaking-things-at-work?srsltid=AfmBOorCgru7ReSwbVdt40nZmQaaeGfbpjLV7epM0fSv_V01QSY5b5TP&quot;&gt;Gavin Mueller&lt;/a&gt;) have written books arguing more or less that the Luddites were right, and we ought to follow their example in order to resist AI automation&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;Like many people, I have heard a lot about Luddism and Luddites, but only in the context of it being a general term for someone who is anti-technology. I was interested in learning more about the actual historical movement: what kind of people participated, what it was, and what it accomplished. I read Merchant’s and Mueller’s books, plus others&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;, to try and figure all of this out. Who were the actual, historical Luddites? What can we learn from them about burning down datacenters?&lt;/p&gt;
&lt;h3&gt;Who were the Luddites?&lt;/h3&gt;
&lt;p&gt;The Luddites were a decentralized movement of artisans in the 1810s who engaged in violent protest — smashing machines, threatening violence, and ultimately killing people — over the fact that their jobs were being automated away. They were not rich, but they were certainly not unskilled labor: these were people who had apprenticed for &lt;a href=&quot;https://archive.org/stream/extractsfromrec01dendgoog/extractsfromrec01dendgoog_djvu.txt?utm_source=chatgpt.com&quot;&gt;seven years&lt;/a&gt;. They were mostly working from home, producing cloth from raw material given to them by their employer, often with tools rented from that same employer. They were working short weeks (three days, per William Gardiner) at their own discretion.&lt;/p&gt;
&lt;p&gt;In the early 1800s, their skilled labor was becoming unnecessary. With the help of expensive machines, unskilled labor could now produce lower-quality cloth, so employers were beginning to pass over these artisans in favor of cheaper employees: &lt;a href=&quot;https://campus.murraystate.edu/academic/faculty/kBinfield/luddites/LudditeHistory.htm&quot;&gt;children, unapprenticed workers, and women&lt;/a&gt;&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;. Combined with the bad economic position of England at the time (at war with France, and thus deliberately cutting off much European trade), times were beginning to be very tough indeed. Starvation was a real threat.&lt;/p&gt;
&lt;h3&gt;What did they do?&lt;/h3&gt;
&lt;p&gt;Cloth artisans were groups of capable men who were used to getting their own way, knew each other very well, and were broadly respected in their communities. It was thus a natural response for them to organize into what was effectively a militant union. The Luddites would send anonymous threatening letters to their old (or current) bosses, warning them to stop using their machines. If they didn’t comply, they would raid the workshop or factory, smashing the machines up.&lt;/p&gt;
&lt;p&gt;They typically did not harm people, though they certainly delivered threats of bodily harm or even murder, and the raids were violent enough (e.g. shooting through windows) to have risked accidental deaths. In at least two instances where a factory owner was seen as unusually cruel, the Luddites did attempt assassinations: one unsuccessful, and one successful one that eventually prompted a crackdown that ended the movement for good.&lt;/p&gt;
&lt;p&gt;Luddism was fully decentralized. Different communities could and did decide to engage in machine-raiding independently, particularly when news spread of the tactic succeeding. Although each community had its own influential men, there was never a single “leader of Luddism”. King Ludd himself was a folk-tale figure. This made it an absolute nightmare for the British government to try and suppress them: putting down one Luddist group did nothing to prevent other groups from continuing to operate.&lt;/p&gt;
&lt;h3&gt;All the king’s spies&lt;/h3&gt;
&lt;p&gt;I was surprised by how &lt;em&gt;difficult&lt;/em&gt; it was for the government to get a hold of any of the local Luddist ringleaders. The government was willing to offer huge rewards to informers: at one point up to 40x the yearly wage. However, there were no takers for several years. Armies of spies were recruited and tasked with infiltrating Luddist groups, with absolutely no success.&lt;/p&gt;
&lt;p&gt;Why was it so hard? Firstly, because the working class was so overwhelmingly pro-Luddist. People universally blamed the economic situation on the government and the factory owners (rightfully so, since the government had chosen to go to war and the factory owners had chosen to embrace automation). Secondly, the communities in question were so insular and tightly-knit that informers would have to rat on their friends and relatives. The handful of people who did eventually inform lived out the rest of their lives as pariahs.&lt;/p&gt;
&lt;p&gt;Because each group was so insular, any spies trying to infiltrate the movement would have been complete strangers to the community, and would thus have a very hard time gaining the trust of a group of men who had known each other for their whole lives. The spies that did exist were restricted to the occasional inter-group Luddist meetings, where people didn’t all know each other so closely. But it’s unclear how important those meetings were, since Luddist groups didn’t need to coordinate to achieve their goals. According to Merchant, the spies spent much of their time embellishing tales of an imminent revolution to encourage their employers to keep the money flowing.&lt;/p&gt;
&lt;h3&gt;The crackdown&lt;/h3&gt;
&lt;p&gt;In the absence of reliable information, the British government was forced to use force. And they did, sending 12,000 troops&lt;sup id=&quot;fnref-4&quot;&gt;&lt;a href=&quot;#fn-4&quot; class=&quot;footnote-ref&quot;&gt;4&lt;/a&gt;&lt;/sup&gt; into the northern counties. This served mainly as an intimidation tactic, since there was no standing Luddite army to fight, and the soldiers spent most of their time marching back and forth or being abused by the townspeople.&lt;/p&gt;
&lt;p&gt;More successful was the imposition of a full police state in Yorkshire, under the magistrate Joseph Radcliffe, who was empowered to randomly grab people off the street and interrogate them for days. That pressure eventually convinced a handful of people to give up their local Luddist organizers, who were tried and inevitably hanged. Their deaths (and the ensuing climate of fear) ended the high-water mark of Luddist activity. Even then, Luddist raids continued on and off for &lt;em&gt;six more years&lt;/em&gt; before petering out.&lt;/p&gt;
&lt;h3&gt;Did the Luddites succeed?&lt;/h3&gt;
&lt;p&gt;This is a tricky question. In one sense the answer is obviously no: the movement was crushed, many of their leaders were executed, the textile industry continued to be automated, and today there are no longer thousands of jobs for skilled British weavers, knitters, spinners and dyers. The pro-automation side won.&lt;/p&gt;
&lt;p&gt;However, they did achieve a number of short-lived victories. Their early threats often succeeded in preventing the building of a factory in a particular location, or in delaying the adoption of industrial machinery in a particular shop by years. In one case, hosiers that had been spooked by Luddite activity gave out pre-emptive bonuses to their workers to discourage them from smashing up their machines (which were indeed not smashed).&lt;/p&gt;
&lt;p&gt;The Luddites also scared the hell out of the British government, who (encouraged by their over-eager spies) thought they might have a genuine revolution on their hands. While they didn’t get many legal concessions at the time, the specter of Luddism must have loomed over the labor reform movement of the 1800s, which saw the first anti-child-labor laws and the beginnings of independent inspection of factories.&lt;/p&gt;
&lt;p&gt;Finally, every book I read argued that the Luddism movement may have created the first idea of a “working class”, by unifying many previously-independent groups of workers against a common enemy. Seen this way, the “political arm”&lt;sup id=&quot;fnref-5&quot;&gt;&lt;a href=&quot;#fn-5&quot; class=&quot;footnote-ref&quot;&gt;5&lt;/a&gt;&lt;/sup&gt; of Luddism can arguably claim partial credit for every labor victory since the 1800s (though the ringleaders were still hanged and the weavers did still lose their jobs).&lt;/p&gt;
&lt;h3&gt;The Luddist approach in a nutshell&lt;/h3&gt;
&lt;p&gt;We can now describe the “Luddist approach” to fighting technological change:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Find a few conspirators in your existing community who agree with your political project (but don’t join a broader organization, since that leaves you vulnerable)&lt;/li&gt;
&lt;li&gt;Make public anonymous demands in support of your specific goals, backed up by threats of violence, signed by a fictional character that’s easy for other groups to appropriate&lt;/li&gt;
&lt;li&gt;If your threats are ignored, attack the physical machines in the dead of night, destroying them and threatening (but not killing) any guards&lt;/li&gt;
&lt;li&gt;Hope your example inspires many more people to independently do (1)-(3) themselves&lt;/li&gt;
&lt;li&gt;Keep raiding, optionally escalating to assassination of some of the bosses, until you bait a totalitarian crackdown from the government &lt;/li&gt;
&lt;li&gt;Eventually get arrested and executed, to great public dismay&lt;/li&gt;
&lt;li&gt;Twenty years later, your example inspires the first national trade unions&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Note that starting or joining a national movement is &lt;em&gt;not&lt;/em&gt; the Luddist approach. Staying almost entirely isolated in small cells helped the Luddists avoid government spies and made them impossible to root out without enforcing a police state. Note also that you need a &lt;em&gt;lot&lt;/em&gt; of public support for this to work: so that you get a lot of copycat groups without having to explicitly organize them, and so that your property destruction and murder is taken sympathetically instead of getting you immediately reported and arrested. &lt;/p&gt;
&lt;h3&gt;Why Luddism is not a good model for the anti-AI movement&lt;/h3&gt;
&lt;p&gt;There are many reasons why this doesn’t map onto the current anti-AI movement. First, Luddism grew from a homogeneous group of high-status workers whose jobs almost vanished overnight, not a broad group of people whose jobs are getting slightly worse because of AI (like the gig-economy workers Merchant endlessly references). That meant that Luddites had &lt;em&gt;really specific&lt;/em&gt; asks: higher wages for piecework, a phased introduction of specific textiles machinery, and so on. They were not generally demanding that the machines all be immediately destroyed&lt;sup id=&quot;fnref-6&quot;&gt;&lt;a href=&quot;#fn-6&quot; class=&quot;footnote-ref&quot;&gt;6&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;Second, Luddism was very local. A pre-existing group of artisans in a particular town would gather in that town — either at work or an inn, say — and decide to petition or raid the businesses in that town that were harming their livelihoods. AI concerns are not like this. It isn’t businesses in Chicago or Tokyo that are making decisions that imperil Chicago’s or Tokyo’s jobs, it’s businesses in San Francisco. Unlike the Luddists, anti-AI activists can’t naturally organize with people they already know to take direct action where they already live.&lt;/p&gt;
&lt;p&gt;Third, Luddist &lt;em&gt;victory&lt;/em&gt; could also be local. If you successfully lobby your local cloth business to not use a weaving machine, you have secured your job at that business for a while. But if you successfully lobby your town (or even your country!) to not build a datacenter, it doesn’t meaningfully improve your local position, since your job can be as easily replaced by a datacenter on the other side of the planet.&lt;/p&gt;
&lt;h3&gt;A total failure of leadership&lt;/h3&gt;
&lt;p&gt;Reading through the history of the Luddites from a modern perspective, I was struck by the near-total absence of &lt;em&gt;good government&lt;/em&gt;. The artisans were left to work out their grievances with their bosses more or less by themselves, with no formal channels for complaint or any attempt at mediation. When the government did intervene — in response to near-universal unrest in &lt;em&gt;half of the country&lt;/em&gt; — they did this:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Make machine-breaking and oath-taking capital crimes&lt;/li&gt;
&lt;li&gt;Dump thousands of soldiers more or less at random into the area, with no plan to guard factories or do anything beyond just hang around in case a revolution broke out&lt;/li&gt;
&lt;li&gt;Empower a single magistrate to arrest and interrogate whoever he wanted in order to root out the conspiracy&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;I suppose it worked, in the sense that it eventually succeeded in stopping the Luddist raids. But I can’t help but think that even a token gesture of compromise (say, requiring employers to make their wages public, or restricting the most cheap-and-nasty factory-made textile products) would have gone a long way towards calming things down. This almost actually happened! The 1812 Framework Knitters’ Bill, which had these provisions in it, passed the House of Commons but was shot down in the House of Lords.&lt;/p&gt;
&lt;p&gt;Why did the government fail to even make a token attempt at compromise? Before the industrial revolution, I wonder if the workers and bosses of the English textiles industry were genuinely able to often just work out their problems together, so the government never really needed to do large-scale mediation. When that changed — when automation first made it possible for the bosses to durably “win” — government took a long time to realize, so there were some unpleasant decades of disempowered workers trying to bully factory-owners (via riots and death threats), and factory-owners trying to brutalize workers (via direct violence and automation).&lt;/p&gt;
&lt;h3&gt;Final thoughts&lt;/h3&gt;
&lt;p&gt;I can see why modern “Luddites” like Merchant and Mueller — who are genuinely anti-technology — talk so much about the legacy of original Luddites. Luddism was a grassroots organization which notched up some real short-term wins, enjoyed near-total support among the public, and didn’t seem to be troubled by infighting at all&lt;sup id=&quot;fnref-7&quot;&gt;&lt;a href=&quot;#fn-7&quot; class=&quot;footnote-ref&quot;&gt;7&lt;/a&gt;&lt;/sup&gt;. If you’re an anti-AI campaigner, I bet all of that sounds great&lt;sup id=&quot;fnref-8&quot;&gt;&lt;a href=&quot;#fn-8&quot; class=&quot;footnote-ref&quot;&gt;8&lt;/a&gt;&lt;/sup&gt;. But I’m not convinced that the neo-Luddites really are the inheritors of Luddism. A load-bearing feature of Luddism is that it was &lt;em&gt;local&lt;/em&gt;: it didn’t have manifestos, or leaders, or factions, or even much explicit ideology beyond the artisans’ immediate practical concerns. These were local men striking back against the local factories harming their local jobs. That simply isn’t the case with AI, where a datacenter in China can take my job in Australia.&lt;/p&gt;
&lt;p&gt;edit: A reader pointed me at &lt;a href=&quot;https://www.verysane.ai/p/against-the-luddites&quot;&gt;&lt;em&gt;Against the Luddites&lt;/em&gt;&lt;/a&gt;, which argues that (a) the Luddites were an elite (ish) movement, (b) they explicitly and deliberately excluded women, and (c) their leftist theory bonafides are questionable. I don’t really care about (c), agree with (b), and mostly agree with (a), with the caveat that they really did have a broad base of non-elite support.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;I got linked &lt;a href=&quot;https://tante.cc/2026/04/21/ai-as-a-fascist-artifact/&quot;&gt;this article&lt;/a&gt; calling AI a “fascist artifact” (on a blog called “Breaking Frames”, a clear reference to Luddism) while I was writing this blog post.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;I really enjoyed Merchant’s book and did not enjoy Mueller’s (which I found to be 10% about the Luddites and 90% about interminable intra-Marxist ideological arguments). I also read &lt;a href=&quot;https://www.amazon.com.au/Luddites-Protested-Machinery-Industrial-Revolution/dp/171936186X&quot;&gt;The Luddites&lt;/a&gt;, which was effectively a dry summary of the ground Merchant covers, a bunch of other essays, and went back and forth with ChatGPT and Claude on some of the key questions.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;Merchant (around page 134) attempts to characterize Luddism as a pro-feminist movement, citing some examples of women helping organize raids, but later on even he (page 162) quotes a representative of the Irish weaver’s guild effectively saying “we don’t have your English problems of women working in the industry”. In general it’s a bit frustrating that the popular books on Luddism are all fairly uncritically pro-Luddist (though not surprising, I suppose). Merchant doesn’t touch at all on the Luddist &lt;a href=&quot;https://ludditebicentenary.blogspot.com/2011/12/22nd-december-1811-william-milnes.html&quot;&gt;practice&lt;/a&gt; of going around to knitting-shops with women and “discharging them from working”.&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-4&quot;&gt;
&lt;p&gt;Sometimes this is described as more troops that were sent to fight Napoleon (even by Merchant himself on page 89), but that &lt;a href=&quot;https://medium.com/@antonhowes/were-more-troops-sent-to-quash-the-luddites-than-to-fight-napoleon-233c802c216d&quot;&gt;isn’t right&lt;/a&gt;.&lt;/p&gt;
&lt;a href=&quot;#fnref-4&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-5&quot;&gt;
&lt;p&gt;In quotes because it was not an official Luddist group (there were none), just people who were trying to stop the violence through lobbying and legislation.&lt;/p&gt;
&lt;a href=&quot;#fnref-5&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-6&quot;&gt;
&lt;p&gt;Otherwise why would any boss agree, instead of just waiting for the Luddites to do it themselves?&lt;/p&gt;
&lt;a href=&quot;#fnref-6&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-7&quot;&gt;
&lt;p&gt;As far as I can tell this is true: the Luddists basically had no internal conflict. I think this is because each individual cell knew each other well already, and so handled their disagreements privately (instead of by writing pamphlets), and disagreements between cells didn’t matter that much because they had no need to coordinate.&lt;/p&gt;
&lt;a href=&quot;#fnref-7&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-8&quot;&gt;
&lt;p&gt;It beats the hell of the other popular reference, Dune’s &lt;a href=&quot;https://dune.fandom.com/wiki/Butlerian_Jihad&quot;&gt;Butlerian Jihad&lt;/a&gt;, which was two generations of brutal violence followed by the reimposition of the feudal system. (Although, at least the Butlerian Jihad &lt;em&gt;succeeded&lt;/em&gt;…)&lt;/p&gt;
&lt;a href=&quot;#fnref-8&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item><item><title><![CDATA[Many anti-AI arguments are conservative arguments]]></title><link>https://seangoedecke.com/many-anti-ai-arguments-are-conservative/</link><guid isPermaLink="false">https://seangoedecke.com/many-anti-ai-arguments-are-conservative/</guid><pubDate>Sat, 18 Apr 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Most anti-AI rhetoric is left-wing coded. Popular criticisms of AI describe it as a tool of &lt;a href=&quot;https://www.theatlantic.com/podcasts/archive/2025/09/ai-and-the-fight-between-democracy-and-autocracy/684095/&quot;&gt;techno-fascism&lt;/a&gt;, or appeal to predominantly left-wing concerns like &lt;a href=&quot;https://www.technologyreview.com/2025/05/20/1116327/ai-energy-usage-climate-footprint-big-tech/&quot;&gt;carbon emissions&lt;/a&gt;, &lt;a href=&quot;https://www.theguardian.com/commentisfree/2025/sep/10/tech-companies-are-stealing-our-books-music-and-films-for-ai-its-brazen-theft-and-must-be-stopped&quot;&gt;democracy&lt;/a&gt;, or &lt;a href=&quot;https://aphyr.com/posts/420-the-future-of-everything-is-lies-i-guess-where-do-we-go-from-here&quot;&gt;police brutality&lt;/a&gt;. Anti-AI &lt;em&gt;sentiment&lt;/em&gt; is &lt;a href=&quot;https://www.pewresearch.org/short-reads/2025/11/06/republicans-democrats-now-equally-concerned-about-ai-in-daily-life-but-views-on-regulation-differ/&quot;&gt;surprisingly bipartisan&lt;/a&gt;, but the big anti-AI institutions are &lt;a href=&quot;https://www.equaltimes.org/hollywood-s-stand-against-ai-a?lang=en&quot;&gt;labor&lt;/a&gt; &lt;a href=&quot;https://news.bloomberglaw.com/daily-labor-report/punching-in-union-leaders-gear-up-to-tackle-ai-in-future-talks&quot;&gt;unions&lt;/a&gt; and the &lt;a href=&quot;https://www.sanders.senate.gov/press-releases/news-sanders-ocasio-cortez-announce-ai-data-center-moratorium-act/&quot;&gt;progressive wing&lt;/a&gt; of the Democrats.&lt;/p&gt;
&lt;p&gt;This has always seemed weird to me, because the contents of most anti-AI arguments are actually right-wing coded. They’re not necessarily intrinsically right-wing, but they’re the kind of arguments that historically have been made by conservatives, not liberals or leftists. Here are some examples:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Many AI critics complain that AI &lt;a href=&quot;https://www.theguardian.com/commentisfree/2025/sep/10/tech-companies-are-stealing-our-books-music-and-films-for-ai-its-brazen-theft-and-must-be-stopped&quot;&gt;steals copyrighted content&lt;/a&gt;, but prior to 2023, leftists have been &lt;a href=&quot;https://www.reddit.com/r/Socialism_101/comments/1664yyd/i_see_many_leftists_hate_copyright_can_anyone/&quot;&gt;largely&lt;/a&gt; &lt;a href=&quot;https://overland.org.au/2017/08/how-to-think-left-on-copyright/&quot;&gt;anti-intellectual-property&lt;/a&gt; on &lt;a href=&quot;https://jacobin.com/2013/09/property-and-theft&quot;&gt;principle&lt;/a&gt; (either because they’re anti-&lt;em&gt;property&lt;/em&gt;, or because they characterize copyright as benefiting huge media corporations and patent trolls).&lt;/li&gt;
&lt;li&gt;A popular anti-AI-art sentiment is that it’s &lt;a href=&quot;https://www.theguardian.com/commentisfree/2025/may/20/ai-art-concerns-originality-connection&quot;&gt;corrosive to the human spirit&lt;/a&gt; to consume AI slop: in other words, art just inherently ought to be generated by humans, and using AI thus damages some part of our intangible human soul. Whether you like this argument or not, it’s structurally similar to a whole slate of classic arguments-from-intuition for conservative positions like anti-abortion or anti-homosexuality.&lt;/li&gt;
&lt;li&gt;Weird new technological art has traditionally been championed by the left-wing and dismissed by the right-wing (as &lt;a href=&quot;https://medium.com/@elarson39/photography-was-historically-considered-arts-most-mortal-enemy-is-ai-69a2dc2f43ef&quot;&gt;inhuman&lt;/a&gt;, &lt;a href=&quot;https://en.wikiversity.org/wiki/History_of_Photography_as_Fine_Art#:~:text=The%20simplest%20argument%2C%20supported%20by%20many%20painters%20and,mill%20than%20with%20handmade%20work%20created%20by%20inspiration&quot;&gt;cheap&lt;/a&gt;, or &lt;a href=&quot;https://encyclopedia.ushmm.org/content/en/article/degenerate-art-1&quot;&gt;degenerate&lt;/a&gt;). But when it comes to AI art, it’s the left-wing making these arguments, and others (not necessarily right-wingers) arguing that AI art can also be a medium of human artistic expression.&lt;/li&gt;
&lt;li&gt;One main worry about AI is that it’s going to take over a lot of jobs. This is a compelling argument! But the left-wing has recently been famously unsympathetic to this same argument around fossil-fuel energy jobs like &lt;a href=&quot;https://www.cam.ac.uk/research/news/former-coal-mining-communities-have-less-faith-in-politics-than-other-left-behind-areas&quot;&gt;coal mining&lt;/a&gt;, to the point where Biden infamously advised a group of miners in New Hampshire to &lt;a href=&quot;https://thehill.com/changing-america/enrichment/education/476391-biden-tells-coal-miners-to-learn-to-code/&quot;&gt;learn to code&lt;/a&gt;&lt;sup id=&quot;fnref-1&quot;&gt;&lt;a href=&quot;#fn-1&quot; class=&quot;footnote-ref&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;. Halting technological progress to preserve jobs is quite literally a “conservative” position.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;On top of all that&lt;sup id=&quot;fnref-2&quot;&gt;&lt;a href=&quot;#fn-2&quot; class=&quot;footnote-ref&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;, frontier AI models themselves are quite left-wing. Notwithstanding some real cases of data bias (most infamously Google’s image model &lt;a href=&quot;https://www.bbc.com/news/technology-33347866&quot;&gt;miscategorizing&lt;/a&gt; dark-skinned humans as “gorillas”), the models reliably &lt;a href=&quot;https://news.stanford.edu/stories/2025/05/ai-models-llms-chatgpt-claude-gemini-partisan-bias-research-study&quot;&gt;espouse&lt;/a&gt; &lt;a href=&quot;https://www.brookings.edu/articles/the-politics-of-ai-chatgpt-and-political-bias/&quot;&gt;left-wing&lt;/a&gt; &lt;a href=&quot;https://www.cato.org/commentary/how-did-ai-get-so-biased-favor-left&quot;&gt;positions&lt;/a&gt;. Even Elon Musk’s deliberate attempt to create a right-wing AI in Grok has had &lt;a href=&quot;https://www.seangoedecke.com/ai-personality-space/&quot;&gt;mixed success&lt;/a&gt;. In 2006, Stephen Colbert coined the phrase “reality has a left-wing bias”. If the left-wing were more sympathetic to AI, I think they would be using this as a pro-left argument&lt;sup id=&quot;fnref-3&quot;&gt;&lt;a href=&quot;#fn-3&quot; class=&quot;footnote-ref&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;So what happened? A year ago I wrote &lt;a href=&quot;https://www.seangoedecke.com/is-ai-wrong/&quot;&gt;&lt;em&gt;Is using AI wrong? A review of six popular anti-AI arguments&lt;/em&gt;&lt;/a&gt;. In that post I blame the hard right-wing turn many big tech CEOs made in 2024. That was around the same time that LLMs was emerging in the public consciousness with ChatGPT, so it made sense that AI got tagged as right-wing: after all, the billionaires on TV and Twitter talking about how AI were going to change the world were all the same people who’d just gone all-in on Donald Trump. I still think this is a pretty good explanation — just unfortunate timing — but there are definitely other factors at play.&lt;/p&gt;
&lt;p&gt;One obvious factor is the hangover from the pro-crypto mania of 2021 and 2022, where many of the same tech-obsessed folks also posted ugly art and talked about how their technology would change the world forever. Few of these predictions came true (though cryptocurrency has indeed changed the world forever), and it’s understandable that many people viewed AI as a natural continuation of this movement.&lt;/p&gt;
&lt;p&gt;On top of that, Donald Trump himself has come out strongly pro-AI, both in terms of &lt;a href=&quot;https://www.ai.gov/&quot;&gt;policy&lt;/a&gt; and in terms of actually &lt;a href=&quot;https://www.nytimes.com/2026/04/13/us/politics/trump-jesus-picture-pope-leo.html&quot;&gt;posting&lt;/a&gt; AI art himself. This naturally creates a backlash where anti-Trump people are primed to be even more anti-AI&lt;sup id=&quot;fnref-4&quot;&gt;&lt;a href=&quot;#fn-4&quot; class=&quot;footnote-ref&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;. Here are some more reasons:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;AI has real environmental impact (though this is often wildly overstated, as I say &lt;a href=&quot;https://www.seangoedecke.com/is-ai-wrong/&quot;&gt;here&lt;/a&gt;), and the right-wing is politically committed to downplaying or denying anthropogenic environmental impacts in general.&lt;/li&gt;
&lt;li&gt;When times are tough, it’s easy to blame the hot new thing that everyone is talking about. Because the right-wing is currently ascendant in the US, left-wingers are more inclined to talk about how tough times are.&lt;/li&gt;
&lt;li&gt;The left-wing is over-represented in the kind of “computer jobs” that are under direct threat from AI.&lt;/li&gt;
&lt;li&gt;Being pro-Europe has always been left-wing coded, and Europe has been noticeably slower and more sceptical about AI than the USA.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Let me finally put my cards on the table. I would describe myself as on the left wing, and I’m broadly agnostic about the impact of AI. Like the boring fence-sitter I am, I think it will have a mix of positive and negative effects. In general, I’m unconvinced by the pro-copyright and human-soul-related anti-AI arguments, or by the idea that AI is inherently right-wing, but I’m troubled by the environmental impact and the impact on jobs (which in my view are more classically left-wing positions).&lt;/p&gt;
&lt;p&gt;Still, I’m curious what will happen when the left-wing flavor of anti-AI rhetoric disappears, which I think it will (as I said at the start, anti-AI sentiment is actually &lt;a href=&quot;https://www.pewresearch.org/short-reads/2025/11/06/republicans-democrats-now-equally-concerned-about-ai-in-daily-life-but-views-on-regulation-differ/&quot;&gt;pretty bipartisan&lt;/a&gt;). When people start making explicitly right-wing anti-AI arguments, will that cause the left-wing to move a little bit towards supporting AI? Or will right-wing institutions continue to explicitly support AI, allowing anti-AI sentiment to become a wedge issue that the left-wing can exploit to pry away voters? In any case, I don’t think the current state of affairs is particularly stable. In many ways, the dominant anti-AI arguments would fit better in a conservative worldview than in the worldview of their liberal proponents.&lt;/p&gt;
&lt;p&gt;edit: This got lots of comments on &lt;a href=&quot;https://www.reddit.com/r/aiwars/comments/1sp3eki/many_antiai_arguments_are_conservative_arguments/&quot;&gt;various&lt;/a&gt; &lt;a href=&quot;https://www.reddit.com/r/aiwars/comments/1sp9hqk/many_antiai_arguments_are_conservative_arguments/&quot;&gt;Reddit&lt;/a&gt; &lt;a href=&quot;https://www.reddit.com/r/LeftistsForAI/comments/1sp3cxe/many_antiai_arguments_are_conservative_arguments/&quot;&gt;posts&lt;/a&gt;, and was briefly discussed on &lt;a href=&quot;https://news.ycombinator.com/item?id=47813141&quot;&gt;Hacker News&lt;/a&gt;. I don’t think the comments are very good overall, but several comments correctly point out that AI is (like all automation) an anti-labor technology, which means that a labor-focused left will naturally be anti AI. I think my post is consistent with that.&lt;/p&gt;
&lt;div class=&quot;footnotes&quot;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&quot;fn-1&quot;&gt;
&lt;p&gt;I don’t think any did, which is probably for the best — they would have only had a couple of years to break into the industry before hiring collapsed in 2023.&lt;/p&gt;
&lt;a href=&quot;#fnref-1&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-2&quot;&gt;
&lt;p&gt;Another point that isn’t quite mainstream enough but that I still want to mention: AI critics often argue that cavalier deployment of AI means that people might take &lt;a href=&quot;https://www.bbc.com/news/articles/cpd8l088x2xo&quot;&gt;dangerous medical advice&lt;/a&gt; instead of simply trusting their doctor. But anyone who’s been close to a person with chronic illness knows that “just trust your doctor” is kind of right-wing-coded itself, and that the left-wing position is &lt;a href=&quot;https://www.painnewsnetwork.org/stories/2026/4/10/doctor-faces-backlash-after-tweet-claims-four-chronic-illnesses-are-overdiagnosed&quot;&gt;very&lt;/a&gt; &lt;a href=&quot;https://yorkspace.library.yorku.ca/server/api/core/bitstreams/4ac9d968-e9b0-491b-888a-d4ed5aeb1ac3/content&quot;&gt;sympathetic&lt;/a&gt; to patients who don’t or can’t. In a parallel universe, I can imagine the left-wing arguing that patients need AI to avoid the mistakes of their doctors, not the other way around.&lt;/p&gt;
&lt;a href=&quot;#fnref-2&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-3&quot;&gt;
&lt;p&gt;Is it a good argument? I don’t know, actually. The easy counter is that the LLMs are just mirroring the biases in their training data. But you could argue in response that superintelligence is also latent in the training data, and that hill-climbing towards superintelligence also picks up the associated political positions (which just so happen to be left-wing).&lt;/p&gt;
&lt;a href=&quot;#fnref-3&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;li id=&quot;fn-4&quot;&gt;
&lt;p&gt;I am no fan of Donald Trump, but it doesn’t follow that everything he supports is bad (e.g. the &lt;a href=&quot;https://en.wikipedia.org/wiki/First_Step_Act&quot;&gt;First Step Act&lt;/a&gt;).&lt;/p&gt;
&lt;a href=&quot;#fnref-4&quot; class=&quot;footnote-backref&quot;&gt;↩&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</content:encoded></item></channel></rss>