{"componentChunkName":"component---src-templates-blog-post-js","path":"/deckard/","result":{"data":{"site":{"siteMetadata":{"title":"sean goedecke"}},"markdownRemark":{"id":"2c91389e-0d67-5c11-966f-56a80b1709a8","excerpt":"Automated AI text detection is currently an underserved niche. The only game in town is Pangram, which does an excellent job but desperately needs more…","html":"<p>Automated AI text detection is currently an underserved niche. The only game in town is <a href=\"https://www.pangram.com/\">Pangram</a>, which does an excellent job but desperately needs more competition. In a few years, I would be surprised if every major social network doesn’t scan new posts<sup id=\"fnref-1\"><a href=\"#fn-1\" class=\"footnote-ref\">1</a></sup> and comments for AI content in order to tag them (or simply remove them).</p>\n<p>I like that I can rely on Pangram to confirm my suspicions when I read <a href=\"https://arxiv.org/abs/2609.03344\">something</a> that sounds like AI. But it’d be much better if I could choose to avoid AI-generated text in the first place. What I want is something that runs in the background and automatically scans text on websites I visit, without me having to ask for it. I could build something like this on top of Pangram, but it’d <a href=\"https://www.pangram.com/pricing\">cost money</a>, and in general I don’t like the idea of sending every piece of text my browser sees to a third-party service. What about local models?</p>\n<p>The open-source models available for AI text detection are <em>fine</em>. Pangram <a href=\"https://www.pangram.com/blog/pangram-4-technical\">claims</a> a 99.66% detection rate with a 0.004% false positive rate. I benchmarked<sup id=\"fnref-2\"><a href=\"#fn-2\" class=\"footnote-ref\">2</a></sup> a bunch of small local models against a combination of AI-detection datasets and got these results:</p>\n<table>\n<thead>\n<tr>\n<th>Model / variant</th>\n<th align=\"right\">Human falsely flagged</th>\n<th align=\"right\">AI-involved text caught</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><a href=\"https://huggingface.co/ShantanuT01/gradient-ai-text-detector\">Gradient — MLX 4-bit</a></strong></td>\n<td align=\"right\"><strong>2.712%</strong></td>\n<td align=\"right\"><strong>52.35%</strong></td>\n</tr>\n<tr>\n<td><a href=\"https://huggingface.co/benreeve/editlens-roberta-large-onnx-int8\">EditLens RoBERTa-large — community INT8</a></td>\n<td align=\"right\">2.484%</td>\n<td align=\"right\">56.06%</td>\n</tr>\n<tr>\n<td><a href=\"https://huggingface.co/ShantanuT01/vanguard-ai-text-detector\">Vanguard</a></td>\n<td align=\"right\">2.267%</td>\n<td align=\"right\">44.92%</td>\n</tr>\n<tr>\n<td><a href=\"https://huggingface.co/desklib/ai-text-detector-v1.01\">Desklib</a></td>\n<td align=\"right\">3.008%</td>\n<td align=\"right\">45.04%</td>\n</tr>\n<tr>\n<td><a href=\"https://huggingface.co/rasbt/ai-text-detector-distilbert\">Raschka DistilBERT</a></td>\n<td align=\"right\">2.598%</td>\n<td align=\"right\">39.01%</td>\n</tr>\n<tr>\n<td><a href=\"https://huggingface.co/rasbt/ai-text-detector-qwen3-0.6b-variable\">Raschka Qwen3-0.6B</a></td>\n<td align=\"right\">2.028%</td>\n<td align=\"right\">28.67%</td>\n</tr>\n<tr>\n<td><a href=\"https://huggingface.co/rasbt/ai-text-detector-modernbert\">Raschka ModernBERT</a></td>\n<td align=\"right\">1.698%</td>\n<td align=\"right\">21.58%</td>\n</tr>\n<tr>\n<td><a href=\"https://huggingface.co/onnx-community/tmr-ai-text-detector-ONNX\">TMR / Oxidane — INT8</a></td>\n<td align=\"right\">1.595%</td>\n<td align=\"right\">19.35%</td>\n</tr>\n</tbody>\n</table>\n<p>I’m not surprised these are so much worse. I didn’t even benchmark Pangram’s own EditLens 3B model, since that’s too big to keep running in the background on my laptop, and the real production Pangram model is likely one or two orders of magnitude bigger than that. But these models are still good enough to be useful to someone who understands their limitations. If you want to flag an AI-written article, you don’t need to flag all of it, just enough to be suspicious. And so long as you’re aware that the false-positive rate is ~2%, you can avoid treating a single flag as solid proof of AI use.</p>\n<p>Encouraged by this, I vibed up <a href=\"https://github.com/sgoedecke/deckard\">Deckard</a>: a Chrome extension that talks to a locally-running model (the bolded one in the table above) on your Mac. One nice thing is that I didn’t have to start a web server: the Chrome extension is happy to start the model as-needed and can talk with it over <a href=\"https://developer.chrome.com/docs/extensions/develop/concepts/native-messaging\">native messaging</a>. It uses about 400MB-1.2GB of memory while active (so it’s like having five or six extra Chrome tabs open), and it turns itself off if you go five minutes without using the model.</p>\n<p>I was pleasantly surprised to see Deckard successfully mark text I knew was AI-generated, such as the built-in YouTube AI summary or the AI <a href=\"/ai-research-with-codex/\">snippets</a> in my own posts:</p>\n<p><span\n      class=\"gatsby-resp-image-wrapper\"\n      style=\"position: relative; display: block; margin-left: auto; margin-right: auto; max-width: 590px; \"\n    >\n      <a\n    class=\"gatsby-resp-image-link\"\n    href=\"/static/d9c4176cc00bde9a34ef2c29ad9e54cb/c549b/example2.png\"\n    style=\"display: block\"\n    target=\"_blank\"\n    rel=\"noopener\"\n  >\n    <span\n    class=\"gatsby-resp-image-background-image\"\n    style=\"padding-bottom: 26.351351351351354%; position: relative; bottom: 0; left: 0; background-image: url('data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAAFCAYAAABFA8wzAAAACXBIWXMAABYlAAAWJQFJUiTwAAAA1UlEQVQY042O206DQBRF+QuFDjPUDhcpMNBaaZQp7UsjGmP//2eWpzyZqNGHdfa5ZJ/soCgzjuuK/nGLLSx6qTF3muXKiMYyx/NOmYWgvpOo+WbzFVmZEsQqxtt7dt0Gm6YYrdEm4VYl3ITyZCEGpYjCiCj6mTAMsdZijCGo65rn05F+3zMMA4eDZ7vdUFQtRbkmzzLyPP+TVMJcNcjKmoenE93eCyNd72kah6srXNPgnPsXbdtyDRe0Ut7HkUnSTd7z6gfeZv2d6WsvvpfzmY/LBS/zJ8b1l6Hf2hQ/AAAAAElFTkSuQmCC'); background-size: cover; display: block;\"\n  ></span>\n  <img\n        class=\"gatsby-resp-image-image\"\n        alt=\"youtube\"\n        title=\"youtube\"\n        src=\"/static/d9c4176cc00bde9a34ef2c29ad9e54cb/fcda8/example2.png\"\n        srcset=\"/static/d9c4176cc00bde9a34ef2c29ad9e54cb/12f09/example2.png 148w,\n/static/d9c4176cc00bde9a34ef2c29ad9e54cb/e4a3f/example2.png 295w,\n/static/d9c4176cc00bde9a34ef2c29ad9e54cb/fcda8/example2.png 590w,\n/static/d9c4176cc00bde9a34ef2c29ad9e54cb/efc66/example2.png 885w,\n/static/d9c4176cc00bde9a34ef2c29ad9e54cb/c83ae/example2.png 1180w,\n/static/d9c4176cc00bde9a34ef2c29ad9e54cb/c549b/example2.png 2128w\"\n        sizes=\"(max-width: 590px) 100vw, 590px\"\n        style=\"width:100%;height:100%;margin:0;vertical-align:middle;position:absolute;top:0;left:0;\"\n        loading=\"lazy\"\n      />\n  </a>\n    </span></p>\n<p><span\n      class=\"gatsby-resp-image-wrapper\"\n      style=\"position: relative; display: block; margin-left: auto; margin-right: auto; max-width: 590px; \"\n    >\n      <a\n    class=\"gatsby-resp-image-link\"\n    href=\"/static/9ddabc445cc492c847dc6b178801cd3c/07d7d/example1.png\"\n    style=\"display: block\"\n    target=\"_blank\"\n    rel=\"noopener\"\n  >\n    <span\n    class=\"gatsby-resp-image-background-image\"\n    style=\"padding-bottom: 71.62162162162163%; position: relative; bottom: 0; left: 0; background-image: url('data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAAOCAYAAAAvxDzwAAAACXBIWXMAABYlAAAWJQFJUiTwAAACL0lEQVQ4y1VUW3LaQBDkEs6HAS2PSpCQQC8M6An+wDgVx64yBZTtC/j+F+h0LxJ2PprunZmdnZ1Z0RkOhyiKArvdDmVZIssyrNdr5HmOKIoQxzHCMESaptanWGnxeDyG9o9Goys6w8EALy8vOBwO+PP0ZPnt7Q2fn594f3/H6+urxfl8xsfHh+X9fm/xPaFY6Gix3W5xOp1soufnvzbp/uEBvx8fcTwesd1sLHuuix83N+h1u+je3mJgDAYsSLhWOGLWaRAgTBeYRYlFEEbkGLM4gT8PEVCL53GKeULQLlaM2uH7/rcK+bO4u0NWVsirGkW9QbW9x5o9qu/vrS2nb8WeiuUrWXFBe1aUtt8uK2+v3pGQwZ9O4XmevdZk8gtTancysTZpz3MvcN1rnNhtuO2jTdhObblc2ukuVyu71lRX1IKdPG3SGWNan+IXi8X/FcpQVxUybtLTqesaG6JNWnBTqRYwJm901TwxccQ+Gg7omjDke4tSNTxByipD6gU54jrhYQl7HJMF2cUuW9TndA33m2bC1ysPmT0kIuNgLnYc+OQpN8zIQQPpacM+fXMiaDBUQlVotOj3UTNJ6fRRWHaQEXkD6bW58Hd7i4L7fqowJbQPU4EDg8xoo7EojGOTV/KZC3RQ0cRI59R3Nt75SqgrO7oiTwlYqS9w7ZED8kxo7f0egl7P6qDBvPGN24Rqpl57xD8BDSEmpGN+CRqIBmZBW8JhpbTJ367FWrff9T+aVLCghwPhQQAAAABJRU5ErkJggg=='); background-size: cover; display: block;\"\n  ></span>\n  <img\n        class=\"gatsby-resp-image-image\"\n        alt=\"snippets\"\n        title=\"snippets\"\n        src=\"/static/9ddabc445cc492c847dc6b178801cd3c/fcda8/example1.png\"\n        srcset=\"/static/9ddabc445cc492c847dc6b178801cd3c/12f09/example1.png 148w,\n/static/9ddabc445cc492c847dc6b178801cd3c/e4a3f/example1.png 295w,\n/static/9ddabc445cc492c847dc6b178801cd3c/fcda8/example1.png 590w,\n/static/9ddabc445cc492c847dc6b178801cd3c/efc66/example1.png 885w,\n/static/9ddabc445cc492c847dc6b178801cd3c/c83ae/example1.png 1180w,\n/static/9ddabc445cc492c847dc6b178801cd3c/07d7d/example1.png 1478w\"\n        sizes=\"(max-width: 590px) 100vw, 590px\"\n        style=\"width:100%;height:100%;margin:0;vertical-align:middle;position:absolute;top:0;left:0;\"\n        loading=\"lazy\"\n      />\n  </a>\n    </span></p>\n<p>It’s lightweight enough that I have it running all the time. I haven’t noticed my MacBook Pro get hot at all or any decrease in battery life, though your mileage may vary on different machines.</p>\n<p>Is Deckard good yet? That depends. It’s good enough that I’m planning to use it, and I recommend it to anyone who’s interested in automatic AI checking. It’s way, way worse than Pangram, and way worse than I think tooling like this is going to be in the next few years.</p>\n<p>Way back in November 2023, I <a href=\"https://www.seangoedecke.com/llm-driven-agents/\">wrote</a> that AI-driven agents were going to be a really big deal. I recommended starting to develop harnesses early, so you can be ready when the models get good enough:</p>\n<blockquote>\n<p>As with most modern language model engineering, a ReAct agent can also see massive sudden improvements by swapping out the underlying model for a better one. … I think this is another reason to invest in agents like this early, in order to take advantage of more powerful models as they come out.</p>\n</blockquote>\n<p>I was right about that, and I (although it’s lower-stakes) think I’m also right about this. AI detection models are only going to get better<sup id=\"fnref-3\"><a href=\"#fn-3\" class=\"footnote-ref\">3</a></sup> over time: Pangram is not going to be the only game in town forever, and we’re eventually going to see small local models that do a good-enough job at identifying AI-written text. I look forward to swapping out the local model in <a href=\"https://github.com/sgoedecke/deckard\">Deckard</a> with something that’s 2x or 10x better.</p>\n<div class=\"footnotes\">\n<hr>\n<ol>\n<li id=\"fn-1\">\n<p>Substack <a href=\"https://support.substack.com/hc/en-us/articles/50891130623508-How-can-I-detect-AI-on-Substack\">kind of has this</a> already, although you have to click a button to scan the post.</p>\n<a href=\"#fnref-1\" class=\"footnote-backref\">↩</a>\n</li>\n<li id=\"fn-2\">\n<p>Well, me and Astra. Overall my experience vibecoding this was very pleasant: I was able to make a bunch of top-level decisions, I could choose programming languages I was less familiar with but were better choices (like doing inference in C++ instead of Python), and the LLM made me aware of choices I would not have thought of by myself (e.g. using native messaging instead of local HTTP).</p>\n<a href=\"#fnref-2\" class=\"footnote-backref\">↩</a>\n</li>\n<li id=\"fn-3\">\n<p>Is this true, given that AI models will also be getting more human-like over time? That’s a subject for a whole other post, but I think so. First, the AI labs aren’t really incentivized to defeat tools like Pangram (if anything it’s the reverse). Second, I don’t see any way around the fact that AI models have a distinct writing style that’s RL-ed into them.</p>\n<a href=\"#fnref-3\" class=\"footnote-backref\">↩</a>\n</li>\n</ol>\n</div>","frontmatter":{"title":"Automatically detecting AI text in my browser","description":null,"date":"September 8, 2026","tags":["ai","projects"]}}},"pageContext":{"slug":"/deckard/","previous":{"slug":"/radical-responsibility-means-treating-people-like-tools/","title":"Radical responsibility means treating people like tools"},"next":null,"preview":{"slug":"/weird-projects-i-shipped-with-ai/","title":"Weird projects I shipped with AI","snippetHtml":"<p>Where are all the AI-generated projects? This is a <a href=\"https://news.ycombinator.com/item?id=46262545\">common question</a> from AI skeptics: if LLMs are so good at writing code, where is the tsunami of new AI-generated apps, services and games?</p><p>I personally don’t find this to be much of a paradox. Writing code is only one of the bottlenecks involved in actually <a href=\"/how-to-ship\">shipping</a> a new product, after all. It’s also impossible to talk about the paid work I’ve done with AI (you’ll simply have to take my word that it’s increased my productivity). But one thing I can do is share a list of personal projects I’ve built with AI in the last twelve months.<br /><a href=\"/weird-projects-i-shipped-with-ai/\">Continue reading...</a></p>"}}},"staticQueryHashes":["1146911855","3764592887"]}