<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[TEP: Technology, Education and Policy: Under The Hood]]></title><description><![CDATA[The engineering behind things that work. Architecture patterns, system design, and technical deep-dives that cross industries — from AI to blockchain to web. For builders who want to know why it works, not just that it works.]]></description><link>https://www.thewhyman.blog/s/under-the-hood</link><image><url>https://substackcdn.com/image/fetch/$s_!Z0Ez!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb866b3c9-2a56-41b2-8864-3239eb0ef170_1280x1280.png</url><title>TEP: Technology, Education and Policy: Under The Hood</title><link>https://www.thewhyman.blog/s/under-the-hood</link></image><generator>Substack</generator><lastBuildDate>Thu, 01 Oct 2026 03:41:47 GMT</lastBuildDate><atom:link href="https://www.thewhyman.blog/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[The Why Man]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[thewhyman@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[thewhyman@substack.com]]></itunes:email><itunes:name><![CDATA[The Why Man]]></itunes:name></itunes:owner><itunes:author><![CDATA[The Why Man]]></itunes:author><googleplay:owner><![CDATA[thewhyman@substack.com]]></googleplay:owner><googleplay:email><![CDATA[thewhyman@substack.com]]></googleplay:email><googleplay:author><![CDATA[The Why Man]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Cost, Latency, Precision: How Much of the AI Tradeoff Is Real?]]></title><description><![CDATA[The Economics of Shipping AI #2 &#183; Measure first, then decide. The 18&#162; piece measured the cost bar. This one measures the latency bar. The goal never moves.]]></description><link>https://www.thewhyman.blog/p/cost-latency-precision-how-much-of</link><guid isPermaLink="false">https://www.thewhyman.blog/p/cost-latency-precision-how-much-of</guid><dc:creator><![CDATA[The Why Man]]></dc:creator><pubDate>Fri, 11 Sep 2026 01:05:58 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xL41!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca308694-14d5-4c59-9adb-fadd0f405265_2400x1260.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>THE ECONOMICS OF SHIPPING AI &#8212; #2</strong> &#183; <em>Measure first, then decide.</em></p><p><em>A follow-up to <a href="https://www.thewhyman.blog/p/only-18-of-your-ai-dollar-reaches">the 18&#162; piece</a>. That one measured the cost bar. This one measures the latency bar. The goal &#8212; precision &#8212; never moves.</em></p><div><hr></div><p>I was at a dinner for AI leaders at a large bank this week. I asked one of the engineering VPs the question I always ask: <strong>what is actually in the way?</strong> Not the roadmap &#8212; what stops the thing from shipping.</p><p>She answered immediately, without having to think about it. <strong>Cost against latency.</strong></p><p>I asked what models they run. She declined, which is the right answer to that question from someone she met an hour earlier, and I did not press. But the constraint itself she stated plainly, as the thing her organisation is up against.</p><p>That stayed with me, because <strong>it is not the argument I made in the 18&#162; piece.</strong> That one was about where the money goes and had nothing to say about time. I had a gap between what I had published and what someone actually running one of these platforms told me was hard.</p><p>So I measured it. Not next quarter &#8212; <strong>the next day</strong>, on the only system where I control every variable.</p><h3>What I am actually measuring</h3><p>There is an AI concierge on <a href="https://thewhyman.com">thewhyman.com</a>. A recruiter or hiring manager lands on the page and asks it questions about me &#8212; how large a team have you run, what is your reliability record, where has your judgment failed &#8212; and it answers from a knowledge base I maintain.</p><p><strong>It is a small system with a serious job</strong>, and that combination is why it is a useful proxy:</p><ul><li><p><strong>Wrong is expensive.</strong> If it drops a number or invents one, that is my credibility with someone deciding whether to interview me. There is no "eh, close enough."</p></li><li><p><strong>Slow is visible.</strong> Nobody waits ten seconds on a personal site. They leave, and I never learn they were there.</p></li><li><p><strong>The answers are numbers-heavy</strong> &#8212; team sizes, portfolio scale, availability figures &#8212; which makes precision measurable instead of a matter of taste.</p></li></ul><p>The stack, so you can map it onto yours:</p><ul><li><p><strong>Serving</strong> &#8212; Cloudflare Workers AI, streaming responses at the edge.</p></li><li><p><strong>Shape</strong> &#8212; retrieval into a prompt. The same shape as most enterprise RAG.</p></li><li><p><strong>Knowledge base</strong> &#8212; about 15,000 tokens total, of which a typical answer needs a fraction.</p></li><li><p><strong>Retrieval</strong> &#8212; deterministic keyword selection over labelled blocks, not vector search, so I can tell exactly what the model saw for any given answer.</p></li><li><p><strong>Volume</strong> &#8212; low. One site. &#9888;&#65039; Nothing here is a scale claim.</p></li></ul><p>&#11088; <strong>What transfers is not my numbers &#8212; it is the method, and the size of the error I found by re-running it.</strong> A bank at volume has different absolutes and the same structure: retrieval into a prompt, a quality bar that cannot move, a cost line somebody reviews monthly, and a latency budget a human is sitting inside.</p><p>Six models, same knowledge base, same question, five runs each.</p><div><hr></div><h2>It was never a triangle</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xL41!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca308694-14d5-4c59-9adb-fadd0f405265_2400x1260.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xL41!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca308694-14d5-4c59-9adb-fadd0f405265_2400x1260.png 424w, https://substackcdn.com/image/fetch/$s_!xL41!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca308694-14d5-4c59-9adb-fadd0f405265_2400x1260.png 848w, https://substackcdn.com/image/fetch/$s_!xL41!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca308694-14d5-4c59-9adb-fadd0f405265_2400x1260.png 1272w, https://substackcdn.com/image/fetch/$s_!xL41!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca308694-14d5-4c59-9adb-fadd0f405265_2400x1260.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xL41!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca308694-14d5-4c59-9adb-fadd0f405265_2400x1260.png" width="1200" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ca308694-14d5-4c59-9adb-fadd0f405265_2400x1260.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:573424,&quot;alt&quot;:&quot;Maximize precision, subject to cost and latency bars&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Maximize precision, subject to cost and latency bars" title="Maximize precision, subject to cost and latency bars" srcset="https://substackcdn.com/image/fetch/$s_!xL41!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca308694-14d5-4c59-9adb-fadd0f405265_2400x1260.png 424w, https://substackcdn.com/image/fetch/$s_!xL41!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca308694-14d5-4c59-9adb-fadd0f405265_2400x1260.png 848w, https://substackcdn.com/image/fetch/$s_!xL41!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca308694-14d5-4c59-9adb-fadd0f405265_2400x1260.png 1272w, https://substackcdn.com/image/fetch/$s_!xL41!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca308694-14d5-4c59-9adb-fadd0f405265_2400x1260.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Cost, latency and precision get drawn as a triangle. Three things in tension, pick two. That framing is in most vendor decks and it is wrong in a way that costs you the decision.</p><p><strong>They are not three of a kind. Precision is the goal. Cost and latency are bars.</strong></p><p>A bar is a threshold. You clear it or you do not, and clearing it by more buys you nothing &#8212; nobody thanks you for 200ms when the SLA was two seconds. Precision does not work that way. It is the thing you are maximising, and there is no point at which more of it stops being worth having.</p><p>Everything else is a dial.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!J76A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4a04371-15cc-4c78-bc5f-b438f9328bd3_2400x1350.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!J76A!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4a04371-15cc-4c78-bc5f-b438f9328bd3_2400x1350.png 424w, https://substackcdn.com/image/fetch/$s_!J76A!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4a04371-15cc-4c78-bc5f-b438f9328bd3_2400x1350.png 848w, https://substackcdn.com/image/fetch/$s_!J76A!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4a04371-15cc-4c78-bc5f-b438f9328bd3_2400x1350.png 1272w, https://substackcdn.com/image/fetch/$s_!J76A!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4a04371-15cc-4c78-bc5f-b438f9328bd3_2400x1350.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!J76A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4a04371-15cc-4c78-bc5f-b438f9328bd3_2400x1350.png" width="1200" height="675" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a4a04371-15cc-4c78-bc5f-b438f9328bd3_2400x1350.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:675,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:954503,&quot;alt&quot;:&quot;The constraint equation&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The constraint equation" title="The constraint equation" srcset="https://substackcdn.com/image/fetch/$s_!J76A!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4a04371-15cc-4c78-bc5f-b438f9328bd3_2400x1350.png 424w, https://substackcdn.com/image/fetch/$s_!J76A!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4a04371-15cc-4c78-bc5f-b438f9328bd3_2400x1350.png 848w, https://substackcdn.com/image/fetch/$s_!J76A!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4a04371-15cc-4c78-bc5f-b438f9328bd3_2400x1350.png 1272w, https://substackcdn.com/image/fetch/$s_!J76A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4a04371-15cc-4c78-bc5f-b438f9328bd3_2400x1350.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That line replaces the triangle, and it tells you what to do when a model disappoints you. In the results below there is a model that cleared both bars comfortably and still lost, and a model that hit the goal and blew a bar. Neither is a tradeoff. <strong>One violated a constraint. The other missed the point.</strong> You fix those differently, and a triangle cannot tell you which you are looking at.</p><div><hr></div><h2>What the numbers said</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OTRx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8f1ec22-499d-4efe-9a70-a860b6fa84e0_2400x1260.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OTRx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8f1ec22-499d-4efe-9a70-a860b6fa84e0_2400x1260.png 424w, https://substackcdn.com/image/fetch/$s_!OTRx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8f1ec22-499d-4efe-9a70-a860b6fa84e0_2400x1260.png 848w, https://substackcdn.com/image/fetch/$s_!OTRx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8f1ec22-499d-4efe-9a70-a860b6fa84e0_2400x1260.png 1272w, https://substackcdn.com/image/fetch/$s_!OTRx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8f1ec22-499d-4efe-9a70-a860b6fa84e0_2400x1260.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OTRx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8f1ec22-499d-4efe-9a70-a860b6fa84e0_2400x1260.png" width="1200" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f8f1ec22-499d-4efe-9a70-a860b6fa84e0_2400x1260.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:373788,&quot;alt&quot;:&quot;Cost per answer across six models, median of five runs&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Cost per answer across six models, median of five runs" title="Cost per answer across six models, median of five runs" srcset="https://substackcdn.com/image/fetch/$s_!OTRx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8f1ec22-499d-4efe-9a70-a860b6fa84e0_2400x1260.png 424w, https://substackcdn.com/image/fetch/$s_!OTRx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8f1ec22-499d-4efe-9a70-a860b6fa84e0_2400x1260.png 848w, https://substackcdn.com/image/fetch/$s_!OTRx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8f1ec22-499d-4efe-9a70-a860b6fa84e0_2400x1260.png 1272w, https://substackcdn.com/image/fetch/$s_!OTRx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8f1ec22-499d-4efe-9a70-a860b6fa84e0_2400x1260.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Median of five runs per model, after a warm-up call I throw away. Failed responses counted separately, never averaged in.</p><p>&#11088;&#11088; <strong>The model I had in production was the worst of the four that worked, on all three axes.</strong> Most expensive. Slowest. No more accurate. Something in that table is <strong>80% cheaper, 1.4 seconds faster, and drops exactly as many figures.</strong></p><p>Nothing about that is a tradeoff. It is an unmeasured default, which is a different and more embarrassing thing.</p><p><strong>And the reasoning model did not lose on latency &#8212; it never finished.</strong> Five of five runs failed against a forty-five-second ceiling. For this workload it cannot clear the bar at any price. It was buying deliberation on a task with almost nothing to deliberate about: pull the right figures out of supplied context and say them plainly.</p><p>Three method notes, because each exists where the naive version lied to me:</p><ul><li><p><strong>Throw away the first call.</strong> A cold start recorded 137 seconds and made one model look like the cheapest winner at two samples. It was failing two runs in three.</p></li><li><p><strong>Medians, not means.</strong> Cost varied <strong>0.7%</strong> across five runs; precision swung between 3/4 and 4/4 with nothing changed. <strong>Cost is stable enough to publish from few samples. Accuracy is not.</strong></p></li><li><p><strong>Count failures separately.</strong> A request returning nothing in four tenths of a second is the fastest number in the set. Worse, I was still billed for it. <strong>Failure is not the absence of cost. It is cost with no product.</strong></p></li></ul><p>&#128204; Which points at the thing underneath: <strong>I did not benchmark six models. I benchmarked one platform's serving of six models, on one afternoon.</strong> That is the honest reason to run your own numbers &#8212; public benchmarks answer a different question than the one you have.</p><div><hr></div><h2>The drift nobody named</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!AFil!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec2b1502-1b67-44d4-9473-615b4de76deb_2400x1260.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!AFil!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec2b1502-1b67-44d4-9473-615b4de76deb_2400x1260.png 424w, https://substackcdn.com/image/fetch/$s_!AFil!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec2b1502-1b67-44d4-9473-615b4de76deb_2400x1260.png 848w, https://substackcdn.com/image/fetch/$s_!AFil!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec2b1502-1b67-44d4-9473-615b4de76deb_2400x1260.png 1272w, https://substackcdn.com/image/fetch/$s_!AFil!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec2b1502-1b67-44d4-9473-615b4de76deb_2400x1260.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!AFil!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec2b1502-1b67-44d4-9473-615b4de76deb_2400x1260.png" width="1200" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ec2b1502-1b67-44d4-9473-615b4de76deb_2400x1260.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:464242,&quot;alt&quot;:&quot;Model drift versus cost drift, and who owns each&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Model drift versus cost drift, and who owns each" title="Model drift versus cost drift, and who owns each" srcset="https://substackcdn.com/image/fetch/$s_!AFil!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec2b1502-1b67-44d4-9473-615b4de76deb_2400x1260.png 424w, https://substackcdn.com/image/fetch/$s_!AFil!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec2b1502-1b67-44d4-9473-615b4de76deb_2400x1260.png 848w, https://substackcdn.com/image/fetch/$s_!AFil!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec2b1502-1b67-44d4-9473-615b4de76deb_2400x1260.png 1272w, https://substackcdn.com/image/fetch/$s_!AFil!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec2b1502-1b67-44d4-9473-615b4de76deb_2400x1260.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Here is the part I did not expect, and the one I would put in front of an executive.</p><p><strong>Model drift</strong> is named, watched and tooled. Everyone knows the shape: the world moves, inputs shift, answers get worse. There are dashboards, alerts, vendors.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-OjT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cf16310-592b-49c6-a30d-90b9ddf8e0ae_2400x1350.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-OjT!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cf16310-592b-49c6-a30d-90b9ddf8e0ae_2400x1350.png 424w, https://substackcdn.com/image/fetch/$s_!-OjT!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cf16310-592b-49c6-a30d-90b9ddf8e0ae_2400x1350.png 848w, https://substackcdn.com/image/fetch/$s_!-OjT!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cf16310-592b-49c6-a30d-90b9ddf8e0ae_2400x1350.png 1272w, https://substackcdn.com/image/fetch/$s_!-OjT!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cf16310-592b-49c6-a30d-90b9ddf8e0ae_2400x1350.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-OjT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cf16310-592b-49c6-a30d-90b9ddf8e0ae_2400x1350.png" width="1200" height="675" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1cf16310-592b-49c6-a30d-90b9ddf8e0ae_2400x1350.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:675,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:954618,&quot;alt&quot;:&quot;Model drift versus cost drift&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Model drift versus cost drift" title="Model drift versus cost drift" srcset="https://substackcdn.com/image/fetch/$s_!-OjT!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cf16310-592b-49c6-a30d-90b9ddf8e0ae_2400x1350.png 424w, https://substackcdn.com/image/fetch/$s_!-OjT!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cf16310-592b-49c6-a30d-90b9ddf8e0ae_2400x1350.png 848w, https://substackcdn.com/image/fetch/$s_!-OjT!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cf16310-592b-49c6-a30d-90b9ddf8e0ae_2400x1350.png 1272w, https://substackcdn.com/image/fetch/$s_!-OjT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cf16310-592b-49c6-a30d-90b9ddf8e0ae_2400x1350.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Look at what that does to your safeguards. Quality evals pass &#8212; output unchanged. SLOs green &#8212; nothing got slower. Users content &#8212; they cannot tell. <strong>Every instrument reports healthy while your cost per unit of work climbs for a quarter.</strong></p><p>&#11088; <strong>And notice who owns each.</strong> Model drift is an engineering alarm with a team that knows what to do. <strong>Cost drift arrives as a line item that grew without explanation, in a room where nobody has an engineering dashboard.</strong> The more senior problem, with the less mature tooling.</p><p>I want to be precise about my own case, because the obvious version of this story is not what happened. <strong>I did measure when I chose the model.</strong> I ran the comparison, picked the winner on evidence, shipped it. The embarrassing part is not that I skipped the work &#8212; <strong>it is that the answer expired and I did not notice.</strong> Two of the models that beat it this week did not exist when I chose.</p><p>That is a new failure mode, not dependency drift renamed. In conventional software your dependency graph changes only when you change it, and <strong>pinning is a complete defence.</strong> Here it is not &#8212; <strong>because pinning protects correctness, and what decayed was optimality.</strong> My pinned model kept working: same accuracy, no errors, nothing in any log. It quietly became the most expensive way to get that result, because <em>correct</em> is absolute and <em>best</em> is relative to a menu somebody else controls.</p><p>There is no lockfile for that. No test fails.</p><div><hr></div><h2>Continuous Evaluation</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uhQP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6f2b081-8a2a-47b8-8b37-b12c097d71ed_2400x1260.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uhQP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6f2b081-8a2a-47b8-8b37-b12c097d71ed_2400x1260.png 424w, https://substackcdn.com/image/fetch/$s_!uhQP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6f2b081-8a2a-47b8-8b37-b12c097d71ed_2400x1260.png 848w, https://substackcdn.com/image/fetch/$s_!uhQP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6f2b081-8a2a-47b8-8b37-b12c097d71ed_2400x1260.png 1272w, https://substackcdn.com/image/fetch/$s_!uhQP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6f2b081-8a2a-47b8-8b37-b12c097d71ed_2400x1260.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uhQP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6f2b081-8a2a-47b8-8b37-b12c097d71ed_2400x1260.png" width="1200" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b6f2b081-8a2a-47b8-8b37-b12c097d71ed_2400x1260.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:377125,&quot;alt&quot;:&quot;CI, CD, CI2 and Continuous Evaluation, and what triggers each&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="CI, CD, CI2 and Continuous Evaluation, and what triggers each" title="CI, CD, CI2 and Continuous Evaluation, and what triggers each" srcset="https://substackcdn.com/image/fetch/$s_!uhQP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6f2b081-8a2a-47b8-8b37-b12c097d71ed_2400x1260.png 424w, https://substackcdn.com/image/fetch/$s_!uhQP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6f2b081-8a2a-47b8-8b37-b12c097d71ed_2400x1260.png 848w, https://substackcdn.com/image/fetch/$s_!uhQP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6f2b081-8a2a-47b8-8b37-b12c097d71ed_2400x1260.png 1272w, https://substackcdn.com/image/fetch/$s_!uhQP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6f2b081-8a2a-47b8-8b37-b12c097d71ed_2400x1260.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Years ago I gave a talk at SXSW about continuous integration, delivery and innovation &#8212; that shipping is not an event you schedule but a capability you maintain. This is the same argument wearing different clothes.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5gZd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e00fbcc-cec4-4329-9010-759b9f3b6336_2400x1350.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5gZd!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e00fbcc-cec4-4329-9010-759b9f3b6336_2400x1350.png 424w, https://substackcdn.com/image/fetch/$s_!5gZd!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e00fbcc-cec4-4329-9010-759b9f3b6336_2400x1350.png 848w, https://substackcdn.com/image/fetch/$s_!5gZd!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e00fbcc-cec4-4329-9010-759b9f3b6336_2400x1350.png 1272w, https://substackcdn.com/image/fetch/$s_!5gZd!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e00fbcc-cec4-4329-9010-759b9f3b6336_2400x1350.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5gZd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e00fbcc-cec4-4329-9010-759b9f3b6336_2400x1350.png" width="1200" height="675" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4e00fbcc-cec4-4329-9010-759b9f3b6336_2400x1350.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:675,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:949561,&quot;alt&quot;:&quot;Continuous evaluation&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Continuous evaluation" title="Continuous evaluation" srcset="https://substackcdn.com/image/fetch/$s_!5gZd!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e00fbcc-cec4-4329-9010-759b9f3b6336_2400x1350.png 424w, https://substackcdn.com/image/fetch/$s_!5gZd!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e00fbcc-cec4-4329-9010-759b9f3b6336_2400x1350.png 848w, https://substackcdn.com/image/fetch/$s_!5gZd!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e00fbcc-cec4-4329-9010-759b9f3b6336_2400x1350.png 1272w, https://substackcdn.com/image/fetch/$s_!5gZd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e00fbcc-cec4-4329-9010-759b9f3b6336_2400x1350.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>So a one-time bake-off is the wrong shape. It answers <em>which model is best today</em>, and the honest shelf life on that is somewhere between a quarter and a month.</p><p>&#11088;&#11088; <strong>If you already run continuous evals &#8212; and most serious teams do &#8212; this is not new infrastructure.</strong> Your suite runs on a schedule and gates releases. <strong>It almost certainly measures quality alone.</strong> Put cost and latency in the harness you already own. Same fixtures, same schedule, same gate. Three columns instead of one.</p><p>That small change makes two things visible that were structurally invisible:</p><ul><li><p><strong>Regressions you ship blind.</strong> A prompt change that adds context is a cost increase, and nothing in a quality-only suite flags it. Mine moved <strong>2.5&#215; on identical infrastructure</strong> purely from how much context the router loaded.</p></li><li><p><strong>Savings nobody is hunting.</strong> An 80% reduction does not announce itself. It shows up as a new row in a table you were already generating.</p></li></ul><p>&#128204; <strong>And it inverts the objection.</strong> The reflex is that evaluation is overhead. The harness I built to answer a latency complaint is the same harness that found an 80% cost reduction on a system I had already optimised. <strong>Evaluation is not the tax you pay for shipping AI. It is the instrument that finds the money.</strong></p><div><hr></div><h2>What I would tell a platform team</h2><ol><li><p><strong>Precision is the goal; cost and latency are bars.</strong> When a model disappoints, first ask which one it failed &#8212; you fix those differently.</p></li><li><p><strong>Price per completed answer, not per million tokens.</strong> Nobody consumes a million tokens, and a failed request still bills.</p></li><li><p><strong>Put a correctness metric in the comparison.</strong> "Figures kept" was invented for this problem and it is the only thing that stopped me shipping a model that answers fast and leaves the numbers out.</p></li><li><p><strong>Warm up, take medians, count failures separately.</strong> All three exist because the version without them handed me a confident wrong answer.</p></li><li><p>&#11088; <strong>Keep the harness and re-run it.</strong> Continuous integration for code, <strong>continuous evaluation for models</strong> &#8212; because the right answer moves without asking your permission.</p></li></ol><div><hr></div><p><em>The 18&#162; piece found the cost bar breached by rework, because verification was missing. This one finds the latency bar breached by capability the task never asked for. Same goal, different bar &#8212; and neither is safely reclaimed without the same instrument underneath.</em></p><p><em>One site, one workload, five runs per model. If you have measured this at real scale and got something different, I want to hear it.</em></p>]]></content:encoded></item><item><title><![CDATA[Only 18¢ of Your AI Dollar Reaches the Product. Here's How to Fix It.]]></title><description><![CDATA[Rework has always been most of the work. What changed is that generation outran review, and the two things everyone is buying more of are making it worse.]]></description><link>https://www.thewhyman.blog/p/only-18-of-your-ai-dollar-reaches</link><guid isPermaLink="false">https://www.thewhyman.blog/p/only-18-of-your-ai-dollar-reaches</guid><dc:creator><![CDATA[The Why Man]]></dc:creator><pubDate>Sat, 15 Aug 2026 18:59:09 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/63daaa3b-7999-413d-9eee-cc6a100a576f_1200x800.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Paper 1 in the Reliable AI Systems series. Companion to <a href="https://www.thewhyman.blog">Defense in Depth for AI Agents</a>.</em></p><p><em>This is #1 in <strong>The Economics of Shipping AI</strong>. It measures the cost bar. <strong><a href="https://www.thewhyman.blog/p/cost-latency-precision-how-much-of">#2 measures the latency bar</a></strong> &#8212; six models on one production system, and the one I was running turned out to be the most expensive and the slowest of them.</em></p><div><hr></div><p><em>Written as the companion to my lightning talk at ClawCamp SF, Saturday 15 August 2026. If you just scanned the code from the room: this is the whole argument, with the sources.</em></p><p><em>Slides from the talk: <a href="https://xteamos.exponentialos.io/deck/">https://xteamos.exponentialos.io/deck/</a> </em></p><div><hr></div><p>A survey of 2,444 companies put a number on something most engineering teams already feel.</p><p>For every dollar spent on AI tokens: <strong>44 cents goes to fixing bugs the AI created. 27 cents to rewriting its code. 11 cents to review and merge delays.</strong></p><p>Eighteen cents reaches production.</p><p>That number went around as &#8220;82% of your AI spend is wasted,&#8221; and I want to deal with the objection before I use it, because the objection is good.</p><h2>The number is directionally right and numerically dubious</h2><p>The study comes from <a href="https://research.entelligence.ai/">Entelligence</a>, a company that sells tooling for the exact problem the statistic describes. The methodology is self-reported. The 2,444 companies likely include free-tier accounts. A <a href="https://jakecuth.com/work/ai-rework-lab/">forensic teardown</a> of the claim landed on the right verdict: <em>directionally right, numerically dubious.</em></p><p>The teardown&#8217;s sharpest point is the one nobody quoting the stat mentions: <strong>pre-AI software development already spent 75&#8211;80% of effort on maintenance and debugging.</strong> So 82% isn&#8217;t a shocking new tax. It&#8217;s roughly the historical baseline wearing a new hat.</p><p>Which means the interesting question isn&#8217;t &#8220;how much is wasted.&#8221; It&#8217;s <strong>what actually changed.</strong></p><h2>What changed: generation outran review</h2><p>Rework has always been most of the work. What&#8217;s new is the ratio between how fast you can produce code and how fast anyone can check it.</p><p>CodeRabbit measures AI-generated changes carrying roughly <strong>1.7&#215; more issues</strong> than human-written ones. Lightrun&#8217;s 2026 report found <strong>43% of AI-generated code still requires manual debugging in production &#8212; after it passed quality checks.</strong></p><p>Sit with that second one. Not &#8220;before review caught it.&#8221; After.</p><p>So the bottleneck was never cost per token. It&#8217;s <strong>review throughput</strong>, and every model upgrade makes the imbalance worse, because generation scales and human review does not.</p><p>That reframes what a quality gate is. A gate isn&#8217;t a tax on velocity. <strong>A gate is the only thing that raises the share of your spend that reaches the product.</strong> Yield, not savings.</p><p>Three things are quietly eating that yield. The first you can fix before you write a line, and the last is the one nobody says out loud.</p><div><hr></div><h2>Leak zero: you specified it wrong</h2><p>Before routing and before context, there is a cheaper failure. The most expensive bug is the one you specified wrong, because everything downstream then executes the wrong thing faithfully and on budget.</p><p>You know what you want. Getting it into words precise enough for a model to act on is the actual work, and it is where the time and the tokens go. A vague request does not fail loudly. It returns something plausible, you read it, you realise it is not what you meant, and you ask again in slightly different words. Do that three times and you have paid for four answers to get one. A meaningful share of the 44 cents starts there, before a single routing or context decision.</p><p>Neither of the other two fixes helps you here. You can send a badly specified task to the cheapest capable model, with a perfectly curated window, and still get back code you have to rewrite.</p><p>The fix is unglamorous. State the outcome rather than the steps: what must be true when it is done, who it is for, and what would make it wrong. Then leave the how to the model. That is most of it.</p><p>I automated this for myself, because I was never going to do it by hand on every request. The prompt gets rewritten before it runs and I see both versions, so the sharpening happens inside the ten seconds I was going to spend anyway rather than in a fourth attempt.</p><h2>Leak one: you route every task as if it were the hardest one</h2><p>You get access to a frontier model. It&#8217;s good at everything, so you send it everything &#8212; file reads, git operations, deploys, browser automation, boilerplate tests, and the two decisions a day that actually need judgment.</p><p>Those are not the same task. You are paying judgment prices for clerical work, and you hit the rate ceiling by mid-afternoon because you spent the budget on <code>git status</code>.</p><p>The fix is routing by <strong>task class</strong> rather than by habit:</p><p>1) <strong>Deterministic work</strong> &#8212; deploys, git, file operations &#8212; goes to a small fast model. On current public pricing that&#8217;s roughly <strong>20&#215; cheaper</strong> than the frontier tier.</p><p>2) <strong>Browser automation and large reads</strong> go to a cheap high-context model, roughly <strong>50&#215; cheaper</strong>. This one matters more than it sounds: a single page snapshot can dump an entire DOM into your context, and you pay for those tokens on every subsequent turn.</p><p>3) <strong>Code generation</strong> goes to a code-specialised model.</p><p>4) <strong>The expensive model does one thing: judgment.</strong> Specs, gates, synthesis. That&#8217;s it.</p><p>None of that is controversial. It&#8217;s the second leak that people argue with.</p><div><hr></div><h2>Leak two: a bigger context window makes your agent worse</h2><p>Everyone has been told to start a new session when things go sideways. Almost nobody explains why.</p><p>It&#8217;s a cure prescribed without a diagnosis, and once you have the diagnosis, the advice turns out to be the crudest possible intervention.</p><h3>Why a fresh session feels smarter</h3><p>Not because it&#8217;s emptier. <strong>Because the wrong turns are gone.</strong></p><p>Every dead end you explored is still sitting in that window. The three incorrect things you told it about your schema two hours ago are still there, still being conditioned on. A fresh session doesn&#8217;t give the model more room to think &#8212; it removes the accumulated wrong answers.</p><p>That distinction kills the naive fix. <strong>Compressing by volume doesn&#8217;t help if you compress the wrong turns along with the right ones.</strong> You end up with a tidy summary of your own mistakes.</p><h3>The research says it&#8217;s worse than dilution</h3><p>Two findings, both well-established, both counterintuitive.</p><p><strong>Position matters, badly.</strong> Liu et al., <em><a href="https://arxiv.org/abs/2307.03172">Lost in the Middle: How Language Models Use Long Contexts</a></em> (Stanford, published in <a href="https://direct.mit.edu/tacl/article/doi/10.1162/tacl_a_00638/119630/Lost-in-the-Middle-How-Language-Models-Use-Long">TACL</a>), found performance is highest when relevant information sits at the beginning or end of the context and <strong>degrades significantly when the model must retrieve from the middle</strong>. Put the important thing in the middle of a long window and you have hidden it.</p><p><strong>Irrelevant context isn&#8217;t neutral &#8212; it competes.</strong> Chroma&#8217;s context-rot work and the <a href="https://arxiv.org/abs/2505.18761">GSM-DC benchmark</a> both find that models are meaningfully degraded by distractors, and that <strong>semantic similarity drives decay more than length does.</strong> A distractor that looks like the answer costs you far more than one that obviously doesn&#8217;t. A single distractor already degrades baseline; several compound it.</p><p>And then the finding that should change how you think about this entirely:</p><p><strong>Models performed better on shuffled context than on logically coherent context</strong> &#8212; across 18 models. Coherent documents share terminology and structure, which makes them <em>better</em> distractors.</p><p>Read that again, because it inverts the instinct everyone has. <strong>The more relevant-looking the material you stuff into the window, the more damage it does.</strong> &#8220;Just give it all the related files, it&#8217;s all connected&#8221; is precisely the wrong move.</p><p>&#8220;It fits&#8221; is not the same as &#8220;it helps.&#8221;</p><div><hr></div><h2>The architecture: the smallest thing that does the job</h2><p>The other two leaks have the same shape, and the same fix.</p><p><strong>Sharp focus.</strong> Keep the working context deliberately small. Not as a limitation you&#8217;re working around &#8212; as a design choice.</p><p><strong>Long-term memory.</strong> Everything else lives in an index <em>outside</em> the window. The agent doesn&#8217;t need to hold it. It needs to know where it is.</p><p><strong>Just-in-time injection.</strong> Hydrate only the relevant slice, at the moment the task needs it. When the session ends, distil what was learned back into the index.</p><p>That&#8217;s what I run daily. The agent doesn&#8217;t carry everything. It carries the right thing at the right moment and knows where the rest lives.</p><p>Which makes &#8220;start a new session&#8221; look like what it is: <strong>amputation.</strong> It works. It also throws away everything you learned getting there. That&#8217;s only a rational trade if you have no memory layer &#8212; which, for most people giving the advice, is true.</p><p><strong>Two better techniques:</strong></p><p><strong>Self-compress on your own schedule.</strong> The critical word is <em>your</em>. Auto-compaction fires when the system hits a limit, which is precisely the moment you have least control over what survives. Compress deliberately, at a natural boundary, and you choose what carries forward &#8212; by relevance, not by volume.</p><p><strong>Write the handoff as insurance.</strong> You don&#8217;t get to pick when a session dies. This is the case you can&#8217;t schedule, which is exactly why it&#8217;s load-bearing.</p><p>Both convert a reset from a loss into a checkpoint.</p><div><hr></div><h2>The gates, and why they&#8217;re the yield story</h2><p>Cheap routing and a small context window let you generate faster. That&#8217;s only an improvement if what you generate is right. Otherwise you&#8217;ve optimised your way into producing defects more efficiently.</p><p>So: gates. Acceptance criteria and evals fixed <em>before</em> code is written. Ships only if it beats baseline.</p><p>If you take one gate from this piece, take this one.</p><p><strong>Never let the model that wrote the code be the model that reviews it.</strong></p><p>Not a different prompt. Not a different persona. A <strong>different family</strong>. Claude reviewing Claude shares the training, so it shares the blind spot &#8212; and it will confidently approve its own mistake.</p><p>I didn&#8217;t want to believe that on vibes, so I tested it: 50 artifacts seeded with known flaws from a taxonomy of five failure modes, four review conditions, 240 review runs, hypothesis registered before a single run so I couldn&#8217;t move the goalposts afterward.</p><p>But the production evidence is blunter than the experiment.</p><p><strong>A cross-family review panel caught a data-integrity bug and a silent cost-tracking regression that the CI gates &#8212; including 100% SonarCloud coverage &#8212; had waved through.</strong> Perfect coverage. Static analysis green. Two real defects shipped anyway, one of them quietly burning money.</p><p>That is the Lightrun 43% statistic happening in a single repository.</p><p>Two more from the same period:</p><p><strong>A single judge passed an attribution defect that the cross-family panel flagged RED at 97% confidence.</strong> One event keyed a raw session-row id while its pair used a different identity, so conversion credit would never have joined back to the referrer. Same code, same moment &#8212; one reviewer missed it, a panel of different families caught it.</p><p><strong>One model caught a corrupt-state bug that two models from the same other family both missed</strong> &#8212; corrupt state silently wiping configuration to defaults.</p><p>I&#8217;m not going to give you a catch rate. I don&#8217;t have a denominator, and a made-up percentage would undo everything above. What I have is a dated record of gates catching defects that coverage, unit tests, and single-model review all passed.</p><div><hr></div><h2>The tax nobody prices</h2><p>Rework doesn&#8217;t only burn tokens. It burns the thing you can&#8217;t buy back.</p><p>Debugging plausible-looking code you didn&#8217;t write is a genuinely harder cognitive task than debugging your own. You have no mental model to fall back on, so you reconstruct intent from scratch. Every defect pulls someone out of build mode, and the context switch costs more than the fix.</p><p>And then the compounding one: <strong>every defect that ships through your gates teaches the developer to stop trusting the output.</strong> Once that happens, they re-read everything &#8212; and you have paid for speed and lost it.</p><p>That isn&#8217;t a soft cost. It&#8217;s the mechanism by which the entire investment decays, and it appears in none of the studies above, because nobody is measuring it.</p><div><hr></div><h2>What to change on Monday</h2><p>Zero. Before you send your next non-trivial request, write one sentence saying what must be true when it is done and what would make it wrong. Notice how often you cannot, because that is the answer.</p><p>One. Take the three most repetitive things you send to a frontier model this week and route them to a cheap one. Measure the difference. You will be annoyed at how much you were spending.</p><p><strong>Two.</strong> Take whatever you stuff into context by default and move it behind a retrieval step. Inject it when the task needs it, not before. Then add one review pass from a <em>different model family</em> before anything ships.</p><p>The model isn&#8217;t your bottleneck. Your request and your routing are.</p><div><hr></div><h2>What I built, and what it costs you</h2><p>Every fix in this piece is implemented in one place, and I did not write the piece and then go looking for a product to attach to it. The product came first.</p><p>Leak one, leak two and the gate are the <strong>Development pack</strong> in xTeamOS: cost routing, staged quality gates, and review by a model family that did not write the code. It is live and running in production today, because it is the pipeline that ships my own work.</p><p>Leak zero is the <strong>Ideation pack</strong>: fixing the thinking before the code. That one is still coming, but the prompt-sharpening half of it is available today as Co-Dialectic, which is free and open source.</p><p><strong>Early access to xTeamOS is free.</strong> Everyone who joins during beta gets 50 percent off the subscription for their first year when it reaches production.</p><p>What the subscription buys is maintenance you do not have to do. Bug fixes and upgrades as the model landscape shifts underneath all of us, which it does constantly. You focus on shipping your product; the research and the plumbing stay my problem.</p><p><strong>The cost difference is the part worth reading twice. You bring your own LLM subscriptions.</strong> There is no API key to fund, unlike tools that meter you through their key and bill you for tokens on top of a seat. You are already paying for those subscriptions. Use them.</p><p>And the direction of travel: support for low-cost open models, on the order of five times cheaper, for the work that does not need frontier judgment. Which is this entire argument applied to the tool itself.</p><h2>Four things, in order of how little they cost you</h2><p><strong>Install Co-Dialectic.</strong> Free, open source, about a minute: <a href="https://codi.exponentialos.io/">codi.exponentialos.io</a>. It is leak zero, fixed, today.</p><p><strong>Join the xTeamOS beta.</strong> Also free, seats are limited, and beta members keep 50 percent off the first year: <a href="https://exponentialos.io/">exponentialos.io</a>.</p><p><strong>Follow Exponential OS on LinkedIn</strong> for the build as it happens: <a href="https://www.linkedin.com/company/exponentialos/">linkedin.com/company/exponentialos</a>.</p><p><strong>Share this</strong> with whoever on your team is burning the other 82 cents. That is the one I cannot do myself.</p><p><em>This is Paper 1 in the Reliable AI Systems series, which argues that AI systems must be governed at multiple layers simultaneously to be reliable in deployment. Each piece takes one layer and contributes a primitive. The companion piece, <a href="https://www.thewhyman.blog">Defense in Depth for AI Agents</a>, covers the operational layer: structured evaluation pipelines, drift monitoring, and guardrail-as-architecture.</em></p><p><em>I&#8217;m Anand Vallamsetla &#8212; Exponential OS, ex-Google, ex-AI Fund (Andrew Ng&#8217;s venture studio). I build the systems that build AI products.</em></p><h2>Sources</h2><ul><li><p>Entelligence AI &#8212; token spend breakdown across 2,444 companies &#8212; https://research.entelligence.ai/</p></li><li><p>Lightrun &#8212; 2026 State of AI-Powered Engineering &#8212; 43% of AI code needs manual debugging after passing quality checks</p></li><li><p>CodeRabbit &#8212; AI-generated changes carry ~1.7&#215; more issues than human-written</p></li><li><p>Jake Cuthbertson &#8212; <em>The Rework Tax</em>, forensic teardown of the 82% claim &#8212; https://jakecuth.com/work/ai-rework-lab/</p></li><li><p>Liu, Lin, Hewitt, Paranjape, Bevilacqua, Petroni, Liang &#8212; <em>Lost in the Middle: How Language Models Use Long Contexts</em>, TACL 2024 &#8212; https://arxiv.org/abs/2307.03172</p></li><li><p><em>How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark</em> (GSM-DC) &#8212; https://arxiv.org/abs/2505.18761</p></li><li><p>Chroma &#8212; context rot research on distractor semantics and shuffled-vs-coherent haystacks</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Why your site is invisible to ChatGPT (even when Google loves you)]]></title><description><![CDATA[Ranking on Google and being citable by AI are not the same problem &#8212; and I asked four engines the same question to prove it.]]></description><link>https://www.thewhyman.blog/p/why-your-site-is-invisible-to-chatgpt</link><guid isPermaLink="false">https://www.thewhyman.blog/p/why-your-site-is-invisible-to-chatgpt</guid><dc:creator><![CDATA[The Why Man]]></dc:creator><pubDate>Fri, 24 Jul 2026 23:55:08 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/74c1b5a9-4a43-4612-9d1d-a841af560321_1280x720.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>An accountant asked a friend of mine a simple question: <em>"If someone asks ChatGPT for an accountant near me, do I show up?"</em> He had no idea. So he checked 15 small-business sites. Two showed up. Thirteen were invisible &#8212; most of them ranking perfectly well on Google.</p><p>That gap is the whole story. Ranking on Google and being citable by AI are <strong>not the same problem</strong>, and treating them as one is why so many sites have gone quiet in AI answers without noticing.</p><p>Here's what's actually happening under the hood.</p><p><strong>An AI answer comes from one of two places.</strong> Either the model answers from memory &#8212; what it absorbed in training, which you can't change and which is months stale &#8212; or it runs a live retrieval, fetches pages, and cites them. Only the second path is one you can influence. And that path has two gates almost nobody checks.</p><h2>Gate 1: Can the crawler even fetch you?</h2><p>The retrieval bots behind ChatGPT, Claude, and Perplexity &#8212; OAI-SearchBot, ClaudeBot, PerplexityBot &#8212; <strong>do not run JavaScript.</strong> A large-scale study of roughly a billion requests found none of them render JS. So if your site builds its content in the browser (any SPA &#8212; React, Vue, most modern stacks), the crawler sees an almost-empty shell where a human sees a full page. You can rank on Google, which does render JS, and be a blank page to ChatGPT.</p><p>The other half of Gate 1 is your own security layer. A WAF or Cloudflare rule that challenges "bot-like" traffic will 403 the AI crawler while serving humans normally. You never see it. You just quietly disappear.</p><p>Here's the ten-second test. Run this on your own site:</p><pre><code><code>curl -A "ClaudeBot/1.0 (+https://www.anthropic.com/claude-bot)" https://yoursite.com</code></code></pre><p>If you get your real content back, good. If you get an empty shell, a redirect loop, or a 403 &#8212; that's why you're not cited, and no amount of content strategy fixes it until you fix this.</p><h2>Gate 2: Once fetched, the index decides &#8212; and every engine's index is different</h2><p>I ran one experiment that made this vivid. I asked four assistants the same open question &#8212; <em>"best API to extract structured data from PDFs"</em> &#8212; each on its own search index.</p><ul><li><p><strong>ChatGPT (Bing index):</strong> surfaced exact-match-domain microsites &#8212; frompdf.dev, transpdf.ai, extractocr.com. When I actually fetched frompdf.dev, it returned <strong>zero bytes</strong> &#8212; no readable content at all &#8212; yet it was cited the most. These sites have no measurable traffic and zero presence on Reddit. They won on one thing: their domain name is the query.</p></li><li><p><strong>Gemini (Google index):</strong> Google Document AI, AWS Textract, Azure, LlamaParse, Unstructured &#8212; plus LandingAI, a legitimate, well-built player with a fraction of the incumbents' traffic.</p></li><li><p><strong>Claude (Brave index):</strong> cloud incumbents and real tools; Parseur, Klippa, LlamaParse.</p></li></ul><p>Same question. Four different worlds. On Bing/ChatGPT, keyword-stuffed domains with no audience beat a readable, funded product. On Google and Brave, legitimacy won. <strong>"Getting cited by AI" isn't one target &#8212; it's four, and they reward different things.</strong></p><p>There's a concentration story on top of this. One synthesis of 680M+ AI citations (Everything-PR / 5WPR, 2026) found the <strong>top 15 domains capture ~68% of all citations</strong>, with Reddit alone near 40% and #1 on every engine. (It's a third-party aggregation, not an audited study &#8212; treat it as directional.) Much of AI citation flows through a handful of high-trust intermediaries you don't own &#8212; a different fight from making your own site citable.</p><h2>What doesn't work, despite the hype</h2><p><strong>llms.txt.</strong> One audit of 1,500 sites found <strong>0.2%</strong> use it, and real crawler consumption today is near zero. Schema/JSON-LD helps a parser understand you but doesn't get you retrieved. Both are fine hygiene; neither is the lever people think it is.</p><h2>So, in order of impact</h2><ol><li><p>Make sure AI crawlers can <em>fetch</em> you &#8212; server-render or pre-generate your main content, and stop challenging verified bots at the edge. (The curl test above.)</p></li><li><p>Put concrete, quotable facts in the first few hundred words, not boilerplate.</p></li><li><p>Optimize per engine, not in general &#8212; the index that feeds your buyers' assistant is the one that matters.</p></li><li><p>llms.txt last, if at all.</p></li></ol><p>None of this requires a rebrand. It requires knowing what an AI agent actually sees when it reads you &#8212; which, it turns out, is almost never what you see in a browser.</p><p><em>(I got tired of running that curl test by hand, so I built a free scanner that fetches your page as ChatGPT and Claude's crawlers do and shows the exact gaps &#8212; no signup: <a href="https://getagentview.com">getagentview.com</a>. But the curl one-liner above will get you 80% of the way for free.)</em></p><p><em>Caveats, because they matter: the four-engine test is a single snapshot, not a controlled study &#8212; results drift week to week. The citation-share figures are third-party aggregations. And correlation isn't causation &#8212; reachability tracks closely with citation in my testing, but the clean proof is a before/after on a site you actually fix.</em></p>]]></content:encoded></item></channel></rss>