{"componentChunkName":"component---src-templates-blog-post-js","path":"/how-to-read-code/","result":{"data":{"site":{"siteMetadata":{"title":"sean goedecke"}},"markdownRemark":{"id":"a1b1e237-3cc6-5f03-ac2c-009eb8064646","excerpt":"Everyone knows how to read a book. Beginning at the first page, you read each word in order, stopping periodically to think, until you arrive at the last word…","html":"<p>Everyone knows how to read a book. Beginning at the first page, you read each word in order<sup id=\"fnref-1\"><a href=\"#fn-1\" class=\"footnote-ref\">1</a></sup>, stopping periodically to think, until you arrive at the last word on the last page. That’s how you read a newspaper article, or a poem, or an email. Why would reading code be any different?</p>\n<h3 id=\"why-code-is-different\" style=\"position:relative;\">Why code is different<a href=\"#why-code-is-different\" aria-label=\"why code is different permalink\" class=\"heading-anchor after\"><svg aria-hidden=\"true\" focusable=\"false\" height=\"16\" version=\"1.1\" viewBox=\"0 0 16 16\" width=\"16\"><path fill-rule=\"evenodd\" d=\"M4 9h1v1H4c-1.5 0-3-1.69-3-3.5S2.55 3 4 3h4c1.45 0 3 1.69 3 3.5 0 1.41-.91 2.72-2 3.25V8.59c.58-.45 1-1.27 1-2.09C10 5.22 8.98 4 8 4H4c-.98 0-2 1.22-2 2.5S3 9 4 9zm9-3h-1v1h1c1 0 2 1.22 2 2.5S13.98 12 13 12H9c-.98 0-2-1.22-2-2.5 0-.83.42-1.64 1-2.09V6.25c-1.09.53-2 1.84-2 3.25C6 11.31 7.55 13 9 13h4c1.45 0 3-1.69 3-3.5S14.5 6 13 6z\"></path></svg></a></h3>\n<p>English text is designed to be read in order. In fact, there are almost no constraints on the order of a text aside from how you want the reader to consume it. In writing this post, I could put the ideas I want to convey in any order I like. Code, on the other hand, is designed to be <em>run</em> by a computer. The order is thus primarily determined by non-human factors. I cannot simply move<sup id=\"fnref-2\"><a href=\"#fn-2\" class=\"footnote-ref\">2</a></sup> a line of code to the beginning of a file or function because I think it provides a better introduction to the program for human readers.</p>\n<p>The other big difference is that English text is always read as a final product, while code is usually read as a <em>diff</em>. We software engineers spend most of our time reading subtle changes to existing code, not brand-new programs. Imagine if reading this post was like that. You would read each successive draft in the order I wrote them, consuming the post as a changed sentence here and a new sentence there. It would be easy to lose track of the overall flow.</p>\n<p>The third reason — and I say this as a lover of literature and poetry — is that code is much more structurally complex than English texts. Even famously difficult books are <em>syntactically</em> simpler than most computer programs (for instance, grammatical dependencies are largely bounded by a single paragraph, while code dependencies can stretch across the entire codebase). Their primary difficulty lies in understanding the nuances of human nature being discussed, not in understanding what each word’s grammatical function is<sup id=\"fnref-3\"><a href=\"#fn-3\" class=\"footnote-ref\">3</a></sup>. Large codebases are also just longer: <em>War and Peace</em> contains around 600,000 words, while most large modern codebases have that many <em>lines</em>.</p>\n<h3 id=\"people-are-bad-at-reading-code\" style=\"position:relative;\">People are bad at reading code<a href=\"#people-are-bad-at-reading-code\" aria-label=\"people are bad at reading code permalink\" class=\"heading-anchor after\"><svg aria-hidden=\"true\" focusable=\"false\" height=\"16\" version=\"1.1\" viewBox=\"0 0 16 16\" width=\"16\"><path fill-rule=\"evenodd\" d=\"M4 9h1v1H4c-1.5 0-3-1.69-3-3.5S2.55 3 4 3h4c1.45 0 3 1.69 3 3.5 0 1.41-.91 2.72-2 3.25V8.59c.58-.45 1-1.27 1-2.09C10 5.22 8.98 4 8 4H4c-.98 0-2 1.22-2 2.5S3 9 4 9zm9-3h-1v1h1c1 0 2 1.22 2 2.5S13.98 12 13 12H9c-.98 0-2-1.22-2-2.5 0-.83.42-1.64 1-2.09V6.25c-1.09.53-2 1.84-2 3.25C6 11.31 7.55 13 9 13h4c1.45 0 3-1.69 3-3.5S14.5 6 13 6z\"></path></svg></a></h3>\n<p>As I’ve said <a href=\"/nobody-knows-how-software-products-work/\">many</a> <a href=\"/in-defense-of-not-understanding-your-codebase/\">times</a>, large computer programs are simply too complex for a single person to fully understand. Reading code in a large program is thus a process of <em>compromise</em>: of deciding which parts to thoroughly grasp and which parts to gloss over; or of portioning out your finite mental capacity across the codebase.</p>\n<p>Because of all this, <strong>most people read code very badly</strong>. They struggle through it front-to-back, like a book, and lose track of the execution flow. Or they just read through the diff and miss the significance of un-edited parts of the code<sup id=\"fnref-4\"><a href=\"#fn-4\" class=\"footnote-ref\">4</a></sup>. Or they simply are defeated by the complexity, give up and just guess what it means. How can you do better?</p>\n<h3 id=\"how-i-read-code\" style=\"position:relative;\">How I read code<a href=\"#how-i-read-code\" aria-label=\"how i read code permalink\" class=\"heading-anchor after\"><svg aria-hidden=\"true\" focusable=\"false\" height=\"16\" version=\"1.1\" viewBox=\"0 0 16 16\" width=\"16\"><path fill-rule=\"evenodd\" d=\"M4 9h1v1H4c-1.5 0-3-1.69-3-3.5S2.55 3 4 3h4c1.45 0 3 1.69 3 3.5 0 1.41-.91 2.72-2 3.25V8.59c.58-.45 1-1.27 1-2.09C10 5.22 8.98 4 8 4H4c-.98 0-2 1.22-2 2.5S3 9 4 9zm9-3h-1v1h1c1 0 2 1.22 2 2.5S13.98 12 13 12H9c-.98 0-2-1.22-2-2.5 0-.83.42-1.64 1-2.09V6.25c-1.09.53-2 1.84-2 3.25C6 11.31 7.55 13 9 13h4c1.45 0 3-1.69 3-3.5S14.5 6 13 6z\"></path></svg></a></h3>\n<p>The best article I’ve read about reading code is <a href=\"https://radimentary.wordpress.com/2020/05/16/of-math-and-memory-part-3-final/\">this piece</a> about reading mathematics papers, which have a similar structure. The author describes a technique called “dyadic scanning”. Instead of reading slowly and sequentially, you make several passes: first to figure out the overall structure, then the sub-structure, and then finally the details.</p>\n<p>For code<sup id=\"fnref-5\"><a href=\"#fn-5\" class=\"footnote-ref\">5</a></sup>, this means reading out-of-order. I like to pick an important path (say, the happy path for the feature introduced in the diff) and trace through which functions are calling which other functions, just to get a sense of the flow. Only once I’ve got a good sense of that do I pay close attention to what those functions are actually doing.</p>\n<p>Usually I do multiple passes, each following a different thread. I’ll take a function or a piece of data and try to figure out how it’s used, fanning out to multiple call-sites (including ones outside the diff) as I go. For small diffs, I just ctrl+f for the function name to jump around; for large diffs, I open it in-editor and ctrl+click for easier navigation. I try to be ruthlessly focused on just the thing I’m looking at right now: everything else gets treated as a black box.</p>\n<p>Once I’m confident I understand the diff, only then will I sit down and carefully read it end-to-end. The purpose of that read is less to learn about the structure — which I should already know by this point — than to catch any weird bits of code I hadn’t noticed in previous out-of-order passes. If I do see anything unusual, I then go back to doing passes.</p>\n<p>This might sound slow. But in fact each pass is very fast, since I’m not painstakingly puzzling through each line of code.</p>\n<h3 id=\"cant-ai-just-do-it-for-you\" style=\"position:relative;\">Can’t AI just do it for you?<a href=\"#cant-ai-just-do-it-for-you\" aria-label=\"cant ai just do it for you permalink\" class=\"heading-anchor after\"><svg aria-hidden=\"true\" focusable=\"false\" height=\"16\" version=\"1.1\" viewBox=\"0 0 16 16\" width=\"16\"><path fill-rule=\"evenodd\" d=\"M4 9h1v1H4c-1.5 0-3-1.69-3-3.5S2.55 3 4 3h4c1.45 0 3 1.69 3 3.5 0 1.41-.91 2.72-2 3.25V8.59c.58-.45 1-1.27 1-2.09C10 5.22 8.98 4 8 4H4c-.98 0-2 1.22-2 2.5S3 9 4 9zm9-3h-1v1h1c1 0 2 1.22 2 2.5S13.98 12 13 12H9c-.98 0-2-1.22-2-2.5 0-.83.42-1.64 1-2.09V6.25c-1.09.53-2 1.84-2 3.25C6 11.31 7.55 13 9 13h4c1.45 0 3-1.69 3-3.5S14.5 6 13 6z\"></path></svg></a></h3>\n<p>Many people are now saying that you don’t have to read code anymore: either because LLMs now produce reliably high-quality code without oversight, or because you can simply ask a reviewer LLM to read the code for you. I think both of these ideas are false.</p>\n<p>Obviously the quality of LLM code is context-dependent. As I wrote in <a href=\"/pure-and-impure-engineering/\"><em>Pure and impure software engineering</em></a>, some software fields (game development, libraries, tools like databases) have wildly different engineering standards, practices and values to other software fields (say, distributed systems at big tech companies). If you’re just making a tool for yourself, you probably don’t have to read the code if you don’t want to. But having read a bunch of AI-generated code this year, I can say that you definitely still have to read it.</p>\n<p>I <em>routinely</em> find massive errors in AI-generated code. These are not bugs — the code typically does what the AI wanted it to do — so much as they’re problems of <a href=\"/human-ai-partnerships-are-for-alignment-not-capability/\">alignment</a>. As a recent example, a small change to thread an extra value through some existing code ballooned out into a complex three-thousand-line diff, because the agent noticed a race condition and built a complex machinery to “fix” it. In fact, this race condition was harmless by design: two pieces of unrelated data could become briefly out of sync, with no customer impact.</p>\n<p>Can LLMs just read the code for you? No, for the same reason: even if they make no mistakes, their technical values will not match yours or those of your company. You can still use LLMs to help you read code, but you have to carefully read and review that LLM output, and you should also be carefully reading the code itself.</p>\n<div class=\"footnotes\">\n<hr>\n<ol>\n<li id=\"fn-1\">\n<p>This is in fact not how many people read in practice. It’s common to skip words or even whole paragraphs due to inattention, or to read an entire book without taking the time to think about it. But at least in theory everyone agrees this is how you should read.</p>\n<a href=\"#fnref-1\" class=\"footnote-backref\">↩</a>\n</li>\n<li id=\"fn-2\">\n<p>Of course, you can do some of this some of the time (more so in some programming languages than others). But the <em>primary</em> constraint is the computer. It is more important that code compiles (or runs without syntax error) than it is for that code to be readable.</p>\n<a href=\"#fnref-2\" class=\"footnote-backref\">↩</a>\n</li>\n<li id=\"fn-3\">\n<p>I don’t wish to understate the importance of syntactically parsing great literature, which I think is both difficult and an underrated skill. For example, I read through the first scene of <em>Hamlet</em> until I saw: “Sit down awhile / And let us once again assail your ears / That are so fortified against our story / What we have two nights seen.” Even a single sentence like this takes work to determine that “That” refers to “ears” and “What” refers to “story” (or, equivalently, has an implied “with” preceding it). And just as one function can have many scattered callers, one sentence in a text can alter the significance of many others.</p>\n<a href=\"#fnref-3\" class=\"footnote-backref\">↩</a>\n</li>\n<li id=\"fn-4\">\n<p>In a previous post, I labeled this the <a href=\"/good-code-reviews/\">main mistake</a> that most engineers make in code review.</p>\n<a href=\"#fnref-4\" class=\"footnote-backref\">↩</a>\n</li>\n<li id=\"fn-5\">\n<p>Here I’m talking about a meaningful diff: something that touches a few hundred lines or more. Trivial diffs can be read end-to-end.</p>\n<a href=\"#fnref-5\" class=\"footnote-backref\">↩</a>\n</li>\n</ol>\n</div>","fields":{"discussionLinks":[]},"frontmatter":{"title":"How to read code","description":null,"date":"October 7, 2026","tags":["good engineers","explainers","ai"]}}},"pageContext":{"slug":"/how-to-read-code/","previous":{"slug":"/shipping-is-the-foundation/","title":"Shipping is the foundation"},"next":null,"preview":{"slug":"/good-code-reviews/","title":"Mistakes I see engineers making in their code reviews","snippetHtml":"<p>In the last two years, code review has gotten much more important. Code is now easy to generate using LLMs, but it’s still just as hard to review. Many software engineers now spend as much (or more) time reviewing the output of their own AI tools than their colleagues’ code.</p><p>I think a lot of engineers don’t do code review correctly. Of course, there are lots of different ways to do code review, so this is largely a statement of my <a href=\"/taste\">engineering taste</a>.<br /><a href=\"/good-code-reviews/\">Continue reading...</a></p>"}}},"staticQueryHashes":["1146911855","3764592887"]}