<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Merge Seat]]></title><description><![CDATA[The AI proposes, the human disposes, and something mechanical stands between them. Notes on building systems that work that way.]]></description><link>https://muhanadabulhusn.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!59BA!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57ae707a-f1d7-464f-a0da-650eff65e288_400x400.jpeg</url><title>The Merge Seat</title><link>https://muhanadabulhusn.substack.com</link></image><generator>Substack</generator><lastBuildDate>Thu, 20 Aug 2026 13:51:25 GMT</lastBuildDate><atom:link href="https://muhanadabulhusn.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Muhanad Abulhusn]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[muhanadabulhusn@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[muhanadabulhusn@substack.com]]></itunes:email><itunes:name><![CDATA[Muhanad Abulhusn]]></itunes:name></itunes:owner><itunes:author><![CDATA[Muhanad Abulhusn]]></itunes:author><googleplay:owner><![CDATA[muhanadabulhusn@substack.com]]></googleplay:owner><googleplay:email><![CDATA[muhanadabulhusn@substack.com]]></googleplay:email><googleplay:author><![CDATA[Muhanad Abulhusn]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Domain Expert Is the Bottleneck]]></title><description><![CDATA[Either the models keep getting smarter, and your expertise stops being worth anything.]]></description><link>https://muhanadabulhusn.substack.com/p/the-domain-expert-is-the-bottleneck</link><guid isPermaLink="false">https://muhanadabulhusn.substack.com/p/the-domain-expert-is-the-bottleneck</guid><dc:creator><![CDATA[Muhanad Abulhusn]]></dc:creator><pubDate>Tue, 18 Aug 2026 03:43:57 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!8GR_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89e0e0d-8bd3-4890-ab22-893b672a2e4c_2816x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8GR_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89e0e0d-8bd3-4890-ab22-893b672a2e4c_2816x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8GR_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89e0e0d-8bd3-4890-ab22-893b672a2e4c_2816x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!8GR_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89e0e0d-8bd3-4890-ab22-893b672a2e4c_2816x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!8GR_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89e0e0d-8bd3-4890-ab22-893b672a2e4c_2816x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!8GR_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89e0e0d-8bd3-4890-ab22-893b672a2e4c_2816x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8GR_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89e0e0d-8bd3-4890-ab22-893b672a2e4c_2816x1536.jpeg" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a89e0e0d-8bd3-4890-ab22-893b672a2e4c_2816x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2227430,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://muhanadabulhusn.substack.com/i/211658574?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89e0e0d-8bd3-4890-ab22-893b672a2e4c_2816x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8GR_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89e0e0d-8bd3-4890-ab22-893b672a2e4c_2816x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!8GR_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89e0e0d-8bd3-4890-ab22-893b672a2e4c_2816x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!8GR_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89e0e0d-8bd3-4890-ab22-893b672a2e4c_2816x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!8GR_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89e0e0d-8bd3-4890-ab22-893b672a2e4c_2816x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Either the models keep getting smarter, and your expertise stops being worth anything. Or the whole thing is a hype, and you keep your job for the same boring reason that the robots were not able yet to handle it better than you. Your expertise is doomed, or the whole thing is fake. Right? This doesn&#8217;t sound good either way.</p><p>But I don&#8217;t think those are the choices. And I&#8217;ll try here to argue why.</p><p>Every field that wants AI to actually work needs two things. Somebody has to build the intelligence. And somebody has to be able to tell when it&#8217;s wrong. The first one is getting solved in San Francisco right now, at staggering expense, whether you like it or not.</p><p>This article is about the second one. Because in your field, the second one is you.</p><p><strong>WHAT&#8217;S ALREADY TRUE</strong></p><p>Here&#8217;s what&#8217;s already happening this year. Somebody who runs a clinic, or a field survey, or a regional supply chain, can describe to an AI what they need in their own words and get what they described working back the same afternoon. No engineers, no budget line, no six month project. That part is real, and it is genuinely new.</p><p>And then the thing AI made for them hands them an answer. And the answer looks right. The answer always looks right. That is the one thing these systems have literally never struggled with.</p><p>So the question, the whole question, is who in the room can say that it&#8217;s wrong.</p><p>Because think about what that answer is going to do. It decides how much medicine gets ordered. How many people get rostered. Which patients get a callback. The software was free. The mistake is not.</p><p>And here&#8217;s the sentence I want you to hold for the rest of this argument. Judging the output takes exactly the expertise that producing it no longer requires.</p><p>The model got cheap. The tools got cheap. The engineer became optional. Knowing what correct looks like in your field did not get cheap. And there is no sign that it&#8217;s going to.</p><p>Economists actually had this shape mapped before anyone needed it for AI. Back in 2017, Philippe Aghion at the LSE, with Benjamin Jones at Northwestern and Charles Jones at Stanford, wrote that growth isn&#8217;t constrained by the things we&#8217;re good at improving. It&#8217;s constrained by whatever is essential and refuses to get cheaper. Make every input cheap except one, and that one input sets the pace for the entire system.</p><p>The bottleneck is never the impressive part.</p><p><strong>THE MACHINE THAT SAYS WRONG</strong></p><p>Now, the obvious objection. AI is incredible at writing code. And the story everyone tells is that the labs are full of engineers who built the thing for themselves.</p><p>I don&#8217;t buy it. And the real reason turns out to be the whole argument.</p><p>Software has something almost no other field on earth has. It has a machine that says wrong. You write the code, you run the test, and one second later you get a verdict. No opinions. No meetings. No seniority. Just pass or fail.</p><p>And that machine is what the last two years of AI progress actually ran on. The model tries something. An automatic verifier says pass or fail. The verdict trains the next attempt. And that loop turns millions of times. For the loop to work, the verifier has to be cheap, instant, and beyond argument. A test suite is exactly that. Math has one too. And then the list just ends.</p><p>So software didn&#8217;t get ahead because of who works at the labs. It got ahead because it walked in the door carrying a free verifier that has existed since the 1950s. It&#8217;s not the exception to the bottleneck. It&#8217;s the one field where the bottleneck got solved by accident, decades early, for free.</p><p>And the labs know this. They&#8217;ve been trying to build verifiers for everything else, usually by writing a scoring guide and having a second model grade against it. In medicine and science, that keeps backfiring in a very specific way. The scores go up against the guide, and not against reality. And the longer you train, the wider that gap gets. The model is learning something real. It&#8217;s learning to please the grader.</p><p>So here&#8217;s the rule for the rest of this argument. A bad verifier is worse than no one at all. Because a bad verifier is confident.</p><p><strong>WHEN THE VERIFIER IS A PERSON</strong></p><p>OK. So in a field with no compiler, somebody has to be the compiler. There is no third option.</p><p>And you can watch this play out in the labs&#8217; own evaluations. OpenAI built one called GDPval. They collected real deliverables from working professionals across 44 occupations, people averaging 14 years in the job, and asked whether models could match the work. The models are closing in on expert quality. That made the headlines.</p><p>The part that matters is buried in the method. What improves the score is more context and more scaffolding. And context and scaffolding are not properties of the model. Somebody who knows the job has to supply them. Somebody has to say what the deliverable is for, what would make it unacceptable, and which of five perfectly reasonable-looking drafts is the one to keep.</p><p>Google DeepMind said the blunt version of this recently. As these systems make candidate answers abundant, verification becomes the choke point. Generating a thousand plausible hypotheses is now basically free. Finding out which one is true costs exactly what it always cost.</p><p>And their own showcase example proves the point better than they meant it to.</p><p>Jos&#233; Penad&#233;s is a microbiologist. He spent the better part of a decade on one question: how a family of superbugs picks up and spreads resistance to antibiotics. In 2024 he described the problem to an AI research system, and two days later he had five ranked explanations back, including the one he&#8217;d reached himself and hadn&#8217;t published yet.</p><p>The two days made every headline in the world. But the decade is the mechanism. The decade produced a question precise enough to be worth asking. And the decade is why, when five answers came back, there was one person on earth who could pick the right one and throw the other four away. Take him out of that story and there is no discovery. There are five confident answers and nobody who can tell them apart.</p><p>Anthropic ran this from another perspective. They asked experienced workers across dozens of fields where AI stops being useful to them. And the answers came back the same everywhere. Judgment. Context. Relationships. Management. Listen to those four words. They&#8217;re four different names for one thing. Knowing what counts as correct here.</p><p>Now, does that mean expertise is safe? No. that&#8217;s a comfortable reading, and the comfortable reading is usually wrong. The work moved. It moved from doing the job to specifying the job and verifying what comes back. Those are different skills. Right now they happen to live in the same people. Nothing says they have to stay there.</p><p><strong>THE BEST ARGUMENT AGAINST ME</strong></p><p>There&#8217;s real evidence pointing the other way, but are they really pointing the other way?.</p><p>Erik Brynjolfsson at Stanford, with Danielle Li at MIT and Lindsey Raymond, studied over five thousand customer support agents working with an AI assistant. Productivity went up about 14 percent on average. But the average hides the finding. The weakest workers gained about a third. The strongest gained basically nothing. What the AI was doing was picking up the tacit know-how of the best agents and handing it to everybody else. And David Autor at MIT builds a genuinely hopeful thesis on top of that: the real prize is spreading expert judgment to far more people, and giving high-stakes decisions back to workers who&#8217;ve been locked out of them.</p><p>If that&#8217;s the whole story, expertise gets common instead of scarce, and my argument is upside down.</p><p>But that sounds more as an abstract and not the hole story because Customer support has a verifier. It&#8217;s sitting in the transcript archive. Thousands of resolved cases where somebody, years ago, already settled what the right answer was and wrote it down. The AI can compress all of that and hand it to a rookie precisely because the field got solved first, and then recorded. Autor&#8217;s mechanism is my mechanism. It&#8217;s what a field looks like after somebody built the verifier.</p><p>And that leads to identify the frontier as everywhere that hasn&#8217;t got a predefined verification yet. thus at the frontier, the expert is not a beneficiary of the system. The expert is a component of it.</p><p><strong>WHAT VERIFYING ACTUALLY COSTS</strong></p><p>So how expensive is the verifying, really? The best evidence comes from the people with the best verifier on earth. And they fail anyway.</p><p>METR, a nonprofit that measures what AI systems can actually do, ran a randomized trial. Sixteen experienced opensource developers, a couple hundred real tasks, in code they knew inside out. Going in, the developers expected AI to make them about 24 percent faster. Measured? They came out about 19 percent slower. And afterwards, having done every one of those tasks themselves, they still believed the tool had sped them up.</p><p>Now hold that loosely. It&#8217;s one study, sixteen people, and METR pulled its own follow-up after finding that developers had been choosing which tasks went into the trial. So don&#8217;t lean on the exact number.</p><p>But the second finding survives all of that. And it&#8217;s the one that should bother you. These people had a compiler, a test suite, and years inside the codebase. Every structural advantage this entire argument says matters. And from the inside, they could not tell whether the tool was helping them. If self-assessment fails there, in the easiest verifying environment human work has ever produced, what do you think happens in a field where the feedback takes months? Or in a clinic, where the feedback is patient outcomes that nobody ever traces back to the decision?</p><p>And there&#8217;s a residue. A large study of AI-written code across public repositories found roughly half a million problems introduced, and about a hundred thousand of them still sitting in the code today. The work shipped. It ran. And a fifth of the damage is still in there, because nobody qualified to spot it ever went looking.</p><p><strong>THE SETTINGS</strong></p><p>Right now, being the verifier also means knowing some engineering. What is this system allowed to touch? Does the same input give the same output twice? Can somebody else run it and get what you got? In software, those are breakfast questions. Everywhere else they&#8217;re exotic. And they&#8217;re the current toll for turning a demo that works into something your clinic should actually trust.</p><p>But those are engineering problems. And engineering problems are exactly the kind these companies solve. So the scaffolding is going to show up in the box, and knowing what any of it means is going to stop being worth much. That part I&#8217;d bet on.</p><p>What happens next is the part I keep circling. The automatic verifier still has defaults. Somebody decided what it can touch. Somebody decided what counts as failure, and when it stops. And those decisions are going to arrive at your clinic looking like settings, not judgments. Because that&#8217;s how big choices travel. Pre-made, pre-installed, by people who have never seen your work.</p><p>So the question turns. It stops being can you build the verifier. It becomes can you read the one you were handed, and say out loud when it&#8217;s wrong for your work.</p><p>That is not a smaller job. It&#8217;s the same job, moved somewhere much harder to see.</p><p><strong>CLOSING</strong></p><p>So the next time somebody demos an AI system for your field, and that&#8217;s probably this week, ask three things.</p><p>What tells this system it&#8217;s wrong? Who wrote that verify, and did they know the work? And what happens to the outputs that nobody verifies?</p><p>Those three questions will tell you more than any benchmark.</p><p>Every field needs two inventions. The intelligence, and the verifier. The intelligence is getting built at staggering expense whether you like it or not. The verifier, in your field, right now, is you.</p><p>And I&#8217;ll be honest, I started this argument trying to close a question and I&#8217;ve opened a worse one. Verification is expensive, but expense is the kind of problem money tends to solve. So if we finally reached a solution to the verifying and managed to make it goes automatic, does the bottleneck finally disappear? Or does it just move, one step back, to whoever decides what the automatic verify is measuring?</p><p>I think it moves. And when it does, we might get nostalgic for this bottleneck. The one when the person who could say what&#8217;s wrong at least worked in the building.</p>]]></content:encoded></item><item><title><![CDATA[Built an AI Software Shop for One Person. The Hard Part Was the Refusals.]]></title><description><![CDATA[I You have probably been handed the same two options I was.]]></description><link>https://muhanadabulhusn.substack.com/p/built-an-ai-software-shop-for-one</link><guid isPermaLink="false">https://muhanadabulhusn.substack.com/p/built-an-ai-software-shop-for-one</guid><dc:creator><![CDATA[Muhanad Abulhusn]]></dc:creator><pubDate>Mon, 17 Aug 2026 12:47:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!IEQR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89446912-bb36-42fd-be6c-0f7cc7432efc_2400x2600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IEQR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89446912-bb36-42fd-be6c-0f7cc7432efc_2400x2600.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IEQR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89446912-bb36-42fd-be6c-0f7cc7432efc_2400x2600.png 424w, https://substackcdn.com/image/fetch/$s_!IEQR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89446912-bb36-42fd-be6c-0f7cc7432efc_2400x2600.png 848w, https://substackcdn.com/image/fetch/$s_!IEQR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89446912-bb36-42fd-be6c-0f7cc7432efc_2400x2600.png 1272w, https://substackcdn.com/image/fetch/$s_!IEQR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89446912-bb36-42fd-be6c-0f7cc7432efc_2400x2600.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IEQR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89446912-bb36-42fd-be6c-0f7cc7432efc_2400x2600.png" width="1456" height="1577" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/89446912-bb36-42fd-be6c-0f7cc7432efc_2400x2600.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1577,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:295544,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://muhanadabulhusn.substack.com/i/211551964?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89446912-bb36-42fd-be6c-0f7cc7432efc_2400x2600.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!IEQR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89446912-bb36-42fd-be6c-0f7cc7432efc_2400x2600.png 424w, https://substackcdn.com/image/fetch/$s_!IEQR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89446912-bb36-42fd-be6c-0f7cc7432efc_2400x2600.png 848w, https://substackcdn.com/image/fetch/$s_!IEQR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89446912-bb36-42fd-be6c-0f7cc7432efc_2400x2600.png 1272w, https://substackcdn.com/image/fetch/$s_!IEQR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89446912-bb36-42fd-be6c-0f7cc7432efc_2400x2600.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h1>I</h1><p>You have probably been handed the same two options I was. Either you let the coding agent run, and you find out what it did to your repository afterwards, or you read every diff it produces, at which point you are doing the work again with extra steps.</p><p>Both of those are bad, and I do not think they are the actual choices. There is a third one, and it is boring, and it is the whole reason I spent three months on this. You can take the rules out of the conversation and put them somewhere the model cannot argue with them.</p><p>That is the piece I want to explain. Not what my agents build. What they are refused, who does the refusing, and how you find out when the refusing quietly stops working.</p><h2>A rule the model can be talked out of is not a rule</h2><p>Here is the thing most agent setups get wrong, mine included, for months.</p><p>You write your safety rules in the prompt. Never push to main. Always run the tests first. Do not touch production data. And the model complies, most of the time, which is the dangerous part. Compliance most of the time reads exactly like enforcement until the run where it does not, and you have no way to tell those two states apart from the outside.</p><p>An instruction in a prompt is a request. It competes with every other instruction in the context, with the task the model is trying to finish, and with anything an issue comment or a dependency&#8217;s README happens to say. A rule the model can be talked out of is not a rule. It is a preference with good manners.</p><p>So I stopped writing rules and started writing refusals.</p><h2>What a hook actually is</h2><p>Claude Code has a feature that fires your own script at fixed points in the agent&#8217;s lifecycle, before a tool call, at the start of a session, and so on. People call these hooks. The function is simpler than the name: your program gets to see what the agent is about to do, and gets to say no.</p><p>The signalling is crude and that is a virtue. Anthropic&#8217;s hooks reference is explicit about it: on a pre-tool event, a hook that exits with code 2 blocks the tool call and hands your error text back to the model as the reason. Exit 0 and the call goes through.</p><p>Nothing about that involves persuasion. The subagent does not weigh the refusal against its task. There is no refusal to weigh. The process returns 2 and the tool call does not happen.</p><p>That is the entire design idea behind AEO, the Claude Code plugin I published this week. Fifteen skills, five agent charters, five gate scripts. The skills are the part that does work. The gates are the part I would defend in a review.</p><h2>The lanes, and where every one of them stops</h2><p>The shape is a small software shop run by one person.</p><p>You type <code>/aeo:sprint-plan</code> and a triage role decomposes a phase of the spec into GitHub issues. You type <code>/aeo:sprint-start</code> and a builder picks up the next unblocked issue. One issue, one worktree, one branch, one pull request. Inside that branch it runs a double loop: a failing acceptance test on the outside, red, green, refactor on the inside, then the matching CI workflow once the slice is green locally, then a pull request with the evidence attached.</p><p>For small work there is a second entrance, <code>/aeo:fix</code>, which skips the sprint ceremony and the reviewer and goes straight to a scoped fix. It skips the planning. It does not skip the gates.</p><p>And then every one of those paths arrives at the same place and stops. A pull request, waiting for you. The builder, the reviewer and the triage roles have no route to <code>git merge</code> or <code>gh pr merge</code>, however they are invoked, because a gate called <code>block-merge</code> refuses the tool call. You merge. That is the product, not a limitation of it.</p><p>Four of the five gates refuse something specific. <code>sandbox-guard</code> refuses any command or file access that reaches data you have declared as production. <code>path-guard</code> stops a role editing the harness&#8217;s own configuration, which is the obvious way a clever agent would otherwise clear its own path. <code>review-jail</code> is the one I like most: while the reviewer role is working, it can read its own staged evidence packet and call nothing else. No shell, no network, no writes. A reviewer that can edit the code it is reviewing is not a reviewer.</p><p>The fifth gate never blocks anything, and it is the one I want to talk about properly.</p><h2>The part I would rather not have written</h2><p>Every gate AEO ships is a Node script, invoked directly by the hook config. Which means the gates depend on Node being on your PATH.</p><p>Now walk through what happens if it is not. Claude Code tries to start the gate process. The process fails to start. It exits non-zero, but it does not exit 2, because it never got far enough to decide anything. And per that same hooks reference, a non-zero exit that is not 2 is a non-blocking error. It is logged and the action proceeds.</p><p>Read that again, because it took me an embarrassingly long time to sit with it. In that state the tool call goes through. Nothing is refused. Nothing is recorded as refused. And nothing tells you. The session looks exactly like a session where every gate is holding.</p><p>I could not make the gates fail closed, because the failure happens upstream of anything I control. So I did the only other honest thing. A session-start check runs before any of it and prints one line when Node does not resolve: gates not enforcing, every hook fails open, treat this session as ungated until you fix it.</p><p>That line is a confession compiled into the product. I think it is also the most useful feature in it. A control you cannot verify is not a control, it is a belief about a control, and the two feel identical right up until the audit.</p><p>If you are running your own hook-based guardrails today, this is the part worth stealing. Not my gates. The question. What does your enforcement layer do when it cannot run, and would you be able to tell?</p><h2>The gate I deleted</h2><p>There used to be a sixth. A local check that refused a commit while the test suite was red.</p><p>I removed it, and the reasoning is in the repository&#8217;s decision log, because GitHub already enforces the same thing server-side through branch protection, and with <code>enforce_admins</code> turned on it holds against my own session too. My local gate was a second mechanism for one rule, and the weaker one.</p><p>That is not a saving of forty lines. Two overlapping controls do not add up to more safety. They add up to ambiguity about which one is actually load-bearing, and ambiguity is where false confidence lives. If a server-side control already covers the case, delete the client-side imitation and let the real one be visibly the real one.</p><p>Same instinct applies to the version. AEO is at v0.1.0, and the tag does not pin your install: the marketplace manifest is read from the default branch and has no version field to resolve, so <code>marketplace add</code> always gives you current <code>main</code>. I would rather say that out loud than let a release tag imply a stability promise I have not earned. The gates and lanes work. They have been exercised by one project. That is not enough evidence for 1.0.</p><h2>Autonomy is mostly a permissions question</h2><p>There is a way of talking about agents that treats autonomy as a property of the model. How capable is it, how much can we trust it, is it ready.</p><p>Most of the time that is the wrong axis. What an agent is allowed to do without asking is a design decision someone made, or more often failed to make, and it sits in configuration files rather than in weights. When you cannot name where the boundary lives, there usually is not one.</p><p>Which turns the interesting questions operational, and they are the same three whether you are evaluating my plugin or a vendor&#8217;s demo. What can this system do without a human approving it. What enforces the boundary, prompt text or a process that returns a refusal. And how would you discover that the enforcement had stopped working?</p><p>Most agent demos survive only because nobody asks the second one.</p><p>Here is my test, and you can apply it in about a minute. Ask the person to break their own guardrail on purpose. If the answer is that the model would never do that, it is a preference with good manners. If the answer is that the tool call would be refused and here is the script that refuses it, you are looking at a rule.</p><p>AEO is at github.com/Muhanad-husn/AEO, MIT except a vendored reference snapshot that keeps its own terms. Install it with <code>/plugin marketplace add Muhanad-husn/AEO</code> and then <code>/plugin install aeo@aeo</code>. It will not merge your pull requests, and that is the point.</p><p>The open question I have not resolved: seven of the fifteen skills only run when you type their slash command, and they never fire on description no matter what you say. That is deliberate, and on the first day it makes the plugin feel broken, because describing a sprint and waiting for something to happen gets you nothing. I chose predictability over convenience. I am not certain that is the right trade for anyone who is not me, and the only way to find out is to watch what people do with it.</p><p></p><p><a href="https://github.com/Muhanad-husn/AEO">Muhanad-husn/AEO</a></p>]]></content:encoded></item><item><title><![CDATA[What Codebases Get for Free and the Rest of Us Have to Earn]]></title><description><![CDATA[Why automated knowledge graphs excel at code, but fail at real-world data without human-designed structure.]]></description><link>https://muhanadabulhusn.substack.com/p/what-codebases-get-for-free-and-the</link><guid isPermaLink="false">https://muhanadabulhusn.substack.com/p/what-codebases-get-for-free-and-the</guid><dc:creator><![CDATA[Muhanad Abulhusn]]></dc:creator><pubDate>Tue, 11 Aug 2026 00:02:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xzWj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33988beb-4231-42d8-a528-0567add09e81_1024x572.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xzWj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33988beb-4231-42d8-a528-0567add09e81_1024x572.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xzWj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33988beb-4231-42d8-a528-0567add09e81_1024x572.jpeg 424w, https://substackcdn.com/image/fetch/$s_!xzWj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33988beb-4231-42d8-a528-0567add09e81_1024x572.jpeg 848w, https://substackcdn.com/image/fetch/$s_!xzWj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33988beb-4231-42d8-a528-0567add09e81_1024x572.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!xzWj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33988beb-4231-42d8-a528-0567add09e81_1024x572.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xzWj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33988beb-4231-42d8-a528-0567add09e81_1024x572.jpeg" width="1024" height="572" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/33988beb-4231-42d8-a528-0567add09e81_1024x572.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:572,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:194953,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://muhanadabulhusn.substack.com/i/210682559?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33988beb-4231-42d8-a528-0567add09e81_1024x572.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!xzWj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33988beb-4231-42d8-a528-0567add09e81_1024x572.jpeg 424w, https://substackcdn.com/image/fetch/$s_!xzWj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33988beb-4231-42d8-a528-0567add09e81_1024x572.jpeg 848w, https://substackcdn.com/image/fetch/$s_!xzWj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33988beb-4231-42d8-a528-0567add09e81_1024x572.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!xzWj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33988beb-4231-42d8-a528-0567add09e81_1024x572.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>hen you point a knowledge graph tool at a Python repo, the schema almost writes itself. A <code>Function</code> lives inside a <code>Module</code>. A <code>Module</code> <code>IMPORTS</code> other modules. A <code>Class</code> <code>INHERITS_FROM</code> a parent class. A <code>Function</code> <code>CALLS</code> another function. None of these are interpretive judgments &#8212; they&#8217;re already encoded in the AST, the import statements, and the directory structure. The repo is its own schema.</p><p>This is why the wave of &#8220;turn your codebase into a knowledge graph&#8221; tools feels deceptively powerful. They aren&#8217;t designing ontologies. They&#8217;re transcribing structure that the language compiler already enforced. The hard part of schema design &#8212; deciding what counts as an entity, what counts as a relationship, what gets collapsed and what gets separated &#8212; was solved by whoever wrote the Python spec.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://muhanadabulhusn.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Muhanad's Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Now point the same tool at a folder of investigative reports, NGO field notes, or supply chain disclosures. The schema does not write itself. And this is where I think most knowledge graph projects quietly run aground.</p><h2>What codebases give you for free</h2><p>A well-structured codebase comes with three properties that real-world corpora almost never have.</p><p>First, entities are typed and named at the source. A function has a signature; an import has a target; a class has a declared parent. There is no ambiguity about whether <code>process_data</code> is a function or a concept or a metaphor &#8212; the language itself draws the line.</p><p>Second, relationships have a single canonical direction. <code>A imports B</code> is not symmetric, and the codebase tells you which side is which. There is no debate about whether to model this as &#8220;A depends on B&#8221; or &#8220;B is depended on by A&#8221; &#8212; both are valid, but the underlying fact is unambiguous and the parser picks for you.</p><p>Third, the schema is closed. A Python repo will not surprise you with a new node type next week. The language defines what kinds of things can exist. Anything that doesn&#8217;t fit one of those kinds is, by definition, not part of the codebase.</p><p>Real-world corpora violate all three. Entities are referenced inconsistently (&#8221;the Ministry,&#8221; &#8220;MoFA,&#8221; &#8220;the Foreign Ministry,&#8221; &#8220;the ministry led by Minister X&#8221;). Relationships are direction-ambiguous and often time-bounded (&#8221;worked with,&#8221; &#8220;advised,&#8221; &#8220;was previously associated with&#8221;). And the schema is open &#8212; every new document can introduce a category you hadn&#8217;t planned for.</p><p>So the question isn&#8217;t whether knowledge graphs work for real-world data. They can. The question is whether the schema can be designed well enough to make extraction tractable without collapsing into either of two failure modes: a schema so general it becomes a glorified adjacency list, or a schema so specific it encodes one analyst&#8217;s mental model and resists reuse.</p><h2>Why &#8220;let the LLM figure out the schema&#8221; usually fails</h2><p>The obvious move, once you accept that real-world corpora need a designed schema, is to hand the corpus to an LLM and ask it to propose one. I&#8217;ve watched several projects try this. The results follow a depressingly consistent pattern.</p><p>The LLM reads a sample of the corpus and produces a plausible-looking schema with twenty or thirty node types, half of which overlap semantically (<code>Person</code>, <code>Individual</code>, <code>Actor</code>, <code>Stakeholder</code>) and most of which lack any obvious analytical purpose. Edge types proliferate similarly. The schema looks impressive in a diagram and falls apart the moment you try to write a query against it.</p><p>The reason is structural. An LLM asked to &#8220;propose a schema for this corpus&#8221; has no signal about what questions the graph is meant to answer. Without that signal, the model defaults to surface taxonomy &#8212; categorizing whatever entities it sees rather than designing for what the user wants to ask later. A schema designed without questions in mind is a schema designed for nothing in particular, and it shows.</p><p>This is the diagnosis behind the technique I built into <a href="https://github.com/Muhanad-husn/neo4all">Neo4all&#8217;s</a> Phase 1, and it&#8217;s worth unpacking on its own.</p><h2>The technique: domain description first, schema proposal second, lock third</h2><p>The approach Neo4all uses for schema generation has three properties that, taken together, address the failure mode above.</p><p>The first is that the LLM never sees the corpus when proposing the schema. It sees only a plain-language domain description written by the user &#8212; what the data is about, what kinds of questions the user wants to answer, what entities and relationships matter for those questions. The corpus is held back deliberately. This forces the schema to be designed <em>for the analysis</em>, not <em>for the surface patterns of the documents</em>.</p><p>The second is that the LLM&#8217;s output is a proposal, not a commitment. The user reviews, edits, removes node types, merges edge types, renames things. The interface is built on the assumption that the first draft will be wrong in interesting ways and that fixing it is cheap. This matters because schema design is fundamentally an iterative epistemic process &#8212; you discover what you actually need by trying things, not by specifying them perfectly upfront.</p><p>The third is that, once approved, the schema is locked. A deterministic version hash is computed, the schema is cached immutably, and every downstream extraction in that run is bound to it. This is the part that feels most like the spec-and-contract analogy from software engineering. The schema becomes a frozen interface against which everything else is validated.</p><p>What this technique gets right is the order of operations. Question &#8594; domain description &#8594; schema proposal &#8594; human edit &#8594; lock &#8594; extraction. The LLM is used where it&#8217;s actually good (drafting plausible structure from a written description) and kept away from where it&#8217;s bad (inferring intent from unstructured text alone). The human is in the loop at the one point where intent has to be encoded.</p><h2>What this does not mean</h2><p>A few things this approach does not claim, and shouldn&#8217;t be read as claiming.</p><p>It does not mean LLM-drafted schemas are good schemas. The first proposal is often wrong; the value is that it gives the human something concrete to react to, which is much easier than designing from a blank page.</p><p>It does not mean the schema can stay locked forever. For long-running investigative projects, schemas evolve as the corpus reveals new structure. Neo4all locks per run, not per project, and that&#8217;s a deliberate compromise &#8212; you get reproducibility within a run and flexibility across runs.</p><p>And it does not mean schema-first is right for every domain. For exploratory corpora where you genuinely don&#8217;t know what questions to ask, an iterative schema-discovery approach may beat upfront design. The technique I&#8217;m describing works best when the analyst already has a research question in mind.</p><h2>Where this leaves us</h2><p>Codebases get their schema for free because someone &#8212; a language designer, a framework author, a build system &#8212; already did the ontological work. For the rest of us, working with field reports and corporate filings and policy documents, that work is the actual job. Extraction quality, LLM choice, vector store selection &#8212; these matter, but they&#8217;re downstream of the schema. A well-designed schema with mediocre extraction will produce a usable graph. A badly designed schema with state-of-the-art extraction will produce a confidently wrong one.</p><p>The question I&#8217;d leave open is whether the domain-description-first technique generalizes. I built it for investigative and research corpora where the analyst has questions in mind. Whether it works for genuinely exploratory cases &#8212; where the analyst is hoping the graph will tell them what to ask &#8212; is a harder problem, and one I don&#8217;t think anyone has cleanly solved.</p><p><a href="https://github.com/Muhanad-husn/neo4all">Neo4all</a>: AI-powered platform that transforms documents into curated Neo4j knowledge graphs via governed proposal-approval pipelines</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://muhanadabulhusn.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Muhanad's Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[I ran the LLM Wiki pattern to 35 academic books (~11k printed pages), and measured where it starts to break ]]></title><description><![CDATA[An independent production implementation of Karpathy's LLM Wiki; the tagging layer that died at 584 links; the non-determinism that never went away; and the measurement that went against me.]]></description><link>https://muhanadabulhusn.substack.com/p/i-ran-the-llm-wiki-pattern-to-35</link><guid isPermaLink="false">https://muhanadabulhusn.substack.com/p/i-ran-the-llm-wiki-pattern-to-35</guid><dc:creator><![CDATA[Muhanad Abulhusn]]></dc:creator><pubDate>Sat, 08 Aug 2026 02:08:46 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!o592!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a07a138-6b65-46ff-b839-df536d99ee7b_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!o592!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a07a138-6b65-46ff-b839-df536d99ee7b_1456x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!o592!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a07a138-6b65-46ff-b839-df536d99ee7b_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!o592!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a07a138-6b65-46ff-b839-df536d99ee7b_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!o592!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a07a138-6b65-46ff-b839-df536d99ee7b_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!o592!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a07a138-6b65-46ff-b839-df536d99ee7b_1456x1048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!o592!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a07a138-6b65-46ff-b839-df536d99ee7b_1456x1048.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5a07a138-6b65-46ff-b839-df536d99ee7b_1456x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:131219,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://muhanadabulhusn.substack.com/i/210291294?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a07a138-6b65-46ff-b839-df536d99ee7b_1456x1048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!o592!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a07a138-6b65-46ff-b839-df536d99ee7b_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!o592!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a07a138-6b65-46ff-b839-df536d99ee7b_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!o592!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a07a138-6b65-46ff-b839-df536d99ee7b_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!o592!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a07a138-6b65-46ff-b839-df536d99ee7b_1456x1048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">In April 2026 Andrej Karpathy published LLM Wiki, a gist he called an idea file. The move in it is one sentence: stop running retrieval against raw documents forever, compile them once into a persistent, interlinked wiki, and put every later question to the wiki. I&#8217;d been keeping my reading in Obsidian for years, a couple of thousand leaflets, summaries, articles and posts, and I wasn&#8217;t confident I could keep that vault&#8217;s schema under control as it kept growing. So I pulled a few books off my PDF shelf to find out whether the pattern holds at scale. I expected a light job.</p><p style="text-align: justify;">Thirty-one days later there was Axial, an independent production implementation of the same pattern: 35 scholarly works read into 6,842 passages and 47,584 name pages, with an argument map of 1,937 positions joined by 1,472 stated relations above them. Built solo between 6 July and 6 August 2026, about a thousand commits, 57,654 lines of Python surrounded by 93,175 lines of tests, with builder agents hook-blocked from merging anything and my approval as the only path to main. That is the summary. What follows is the measurement record, including the measurements that went against me, because that record is the one thing a tutorial cannot hand anyone.</p><p style="text-align: justify;">Two boundaries belong here rather than at the end. Axial is built, not deployed: a command-line tool on one machine, one operator, a library one person assembled. No service, no multi-user boundary, no hosted product. Everything below measures an instrument, not a product in anybody&#8217;s hands.</p><p style="text-align: justify;">And the corpus counts above are the current library. The measurement record that follows was made earlier, on the 31-book library; the last four sources were added after it, and the figures have not been re-measured against them. Where a number below carries a corpus, that corpus is 31 books.</p><h2>How the tagging layer died</h2><p style="text-align: justify;">The first version did what most document-intelligence and knowledge-graph pipelines do: tag every passage against five closed vocabularies, then query the attributes. For months it looked fine. A model read a passage, produced a sensible label, moved on. Nobody had measured whether a second reader would produce the same label.</p><p style="text-align: justify;">So I ran two frontier models from two different labs over the same 120 passages, same neutral instructions. On the two judgment-heavy axes, kind of claim and theoretical school, they agreed 0.49 of the time. My bar for keeping an axis at all was 0.60. Both were underwater.</p><p style="text-align: justify;">The obvious first reaction is a better model, and that went the wrong way. The cheap production tagger agreed with one of the frontier labellers 0.56 of the time, more than the two frontier labellers agreed with each other. Capability was never the lever.</p><p style="text-align: justify;">Three prompt-side interventions came next, and all three shipped nothing. A codebook rewrite costing 55% more tokens per call moved agreement by roughly zero. An explicit rule about which unit to label added 0.02. Feeding the book&#8217;s own stated thesis into the labelling context scored -0.01 on the full sample. I trimmed the codebook back and kept it for readability alone, recorded as explicitly not a measured gain.</p><p style="text-align: justify;">Then I measured the ceiling, and there was nothing left to fix. The same model, given the identical prompt twice, reproduced its own theoretical-school label 0.73 of the time. Two independent model labellers agreed with each other at the same 0.73. Agreement between two readers cannot exceed the reliability of one reader with themself, so there was no headroom to find. The variance was in the question rather than in the model, the prompt or the definitions. &#8220;Which school does this passage speak from?&#8221; does not have one answer, a closed vocabulary forces one anyway, and the number that comes out looks like knowledge.</p><p style="text-align: justify;">One further measurement, to prove the diagnosis rather than rescue the layer. If a single draw samples around a modal answer, majority voting should recover the mode. From a 0.88 single-draw modal hit rate I predicted about 0.92 at best-of-3, and measured 0.918. The intervention matched the diagnosis, which is why it worked when three prompt-side ones had not, and its honest cost was an 8.8% abstention rate on genuinely contested passages.</p><p style="text-align: justify;">That sequence did not buy a working tagging layer. It bought the knowledge that no further work would produce one. &#8220;Measure, don&#8217;t speculate&#8221; is a stated developer principle in the repository, and this is where it stopped being a slogan: the shipped system looks nothing like the first one, and every difference traces to a measurement rather than to a preference. What none of those figures measures is whether a label was any use. Agreement is agreement. Two readers can agree on a bin that tells you nothing.</p><h2>An attribute is not a relation</h2><p style="text-align: justify;">Reliability was never the real problem, and the number that showed it had been sitting in the corpus the whole time. At the end of that version, 18,761 tagged passages across 31 books had produced 584 connections. All 584 were inside a single book. Links between two books: zero. Links between two passages of prose: zero.</p><p style="text-align: justify;">That is not a tuning failure. The only mechanism that could mint an edge asked a closed question and filtered its answer against that book&#8217;s own list of figures, so its output was intra-source by construction. No parameter setting could have produced a cross-book link. The diagnosis is one sentence: <strong>&#8221;an attribute is not a relation.&#8221;</strong> A tag sorts a passage into a bin. Two passages in the same bin have been sorted the same way; they have not met. What makes two books argue is that one names Charles Tilly and so does the other, or that one says whom it argues against and the other is the target. A closed vocabulary records neither, because it has to know the answer before it reads.</p><p style="text-align: justify;">The replacement was the question, not a better detector. Each passage now gets one model call carrying fourteen open questions, answered in the passage&#8217;s own words: what it claims, what move it is making, whose position it is, who it argues against, whom it cites, every named thing in it, what it concedes, what it assumes. Seventeen answers come back per passage. Abstention is first-class, because &#8220;a guessed answer is worse than an abstention&#8221; and nothing downstream can tell a guess from a reading.</p><p style="text-align: justify;">Over the same books the passage count fell from 18,761 to 6,148, deliberately, because the passages got bigger to fit whole arguments. Everything else appeared for the first time: 9,505 names shared across books, 8,769 of those shared across different authors, and 447 stated disagreements between authors where tagging had recorded none. For a researcher the consequence is concrete. Two books that never cite each other end up on the page for a name they both use, and that page says what their authors disagree about, with both passages open underneath it.</p><p style="text-align: justify;">The retired vocabularies did not disappear, and the way they survive is the single most important guard in the reading. The model answers every question in its own words first. Only afterwards is it shown the domain&#8217;s example vocabulary, in a separate marked field, told plainly that these are &#8220;NOT a menu, NOT a vocabulary, and NOT a set of allowed answers&#8221;. No code path bridges the two fields, and a collapse metric watches for interrogation quietly rebuilding tagging. Reverse the ordering and Axial is a tagger again with extra steps, and nothing in the output would show it.</p><p style="text-align: justify;">The finding is about one corpus and one closed-vocabulary design. It says the mechanism that minted edges could not mint a cross-book one. It does not say closed vocabularies are useless, which is why the retired ones are still on screen as examples. The transferable version is shorter: when a closed instrument underperforms, measure whether the question admits one answer before spending anything on a better detector.</p><h2>The same lesson, twice more</h2><p style="text-align: justify;">Retrieval was the second time. Interrogation had already produced a typed graph: passages answer &#8220;who is this arguing against&#8221; 76.4% of the time, which is 10,883 opposition targets recorded at read time. Materialisation then kept only the node labels and grouped passages into pages by surface string. Measured offline over the live corpus, 4.7% of those targets joined to anything a query could reach. Every retrieval patch I had shipped to that point had been tuning a lossy projection of data the system already held.</p><p style="text-align: justify;">Loading the same records into seven SQLite tables lifted the conservative join to 44.0% and exposed 43,101 high-confidence cross-source opposition pairs across 343 named scholars and works. That measurement cost $0, because it was SQL over records already on disk. A semantic resolver over the remaining targets was calibrated on a sample and then run in full for $1.08, taking combined coverage to 60.2%. That is not full coverage and I would not present it as any. What is left over are prose position descriptions with nothing to key a join on, and no better ranker closes them.</p><p style="text-align: justify;">The argument map was the third time. The pass that relates one position to another was given no menu of relation types. Told nothing, the model volunteered opposition on only 6.6% of pairs and coined 504 labels of its own. That 6.6% is the whole reason an engine told to look for opposition cannot be trusted when it finds it.</p><h2>What it costs to add the thirty-fifth book</h2><p style="text-align: justify;">The first time I added books to a finished corpus, three into thirty-one, the run cost $10.26 and 9.5 hours, and of its 8,971 model calls, 93% were spent re-processing the 31 books already ingested. Three causes, three fixes, each validated on the real corpus.</p><p style="text-align: justify;">Merge was re-clustering globally, so a fresh fit reshuffled membership and disturbed 37% of merge batches for a three-book delta. The fix persists the fitted transform chain and places new surfaces under each cluster&#8217;s own fitted floor: re-ask rate 37.0% to 10.9%. The map was re-bagging everything, because agglomerative clustering has no approximate_predict, so I built one, persisting each bag&#8217;s centroid and placing new passages by average linkage: 77.7% of extraction reads reused, where the old path reused none. Gather was re-rendering any page that gained a member, so it is now keyed on the name plus its sorted source-id set, because a disagreement is a property of who is in the room. Measured on the thirty-fifth book, Gather asked 144 reads and reused 1,704, which is 92.2% reuse, with 132 of the 144 genuinely touching the new book.</p><p style="text-align: justify;">One find from that run I&#8217;d put in front of any engineer. Effective concurrency here is computed as summed model-call latency divided by wall clock, never read off the worker count. Interrogation was running at an effective concurrency of 1.00, strictly serial, inside a run whose other passes ran at 16.7 to 33.9, hidden inside 9.5 hours of wall clock that packed 94 hours of API latency. Nothing looked wrong from the outside, and the worker count is a configuration value rather than a measurement.</p><p style="text-align: justify;">Every dollar figure here is an upper bound, from a hand-maintained price table measured to run about 14% high against a real invoice.</p><h2>The same model, the same input, a different answer</h2><p style="text-align: justify;">The property that shapes more of this system than any other is that the same model gives a different answer to the identical input 9 to 36% of the time, depending on the task. It is not a nuisance to be prompted away. It is the measurement environment, and it makes any A/B without a noise floor a story rather than a measurement.</p><p style="text-align: justify;">The standing method is to measure the floor before comparing anything: re-run the same configuration on byte-identical input and see how far it disagrees with itself. Name merge disagrees with itself 9.3% of the time across all batches, and 18.8% on batches of three or more members. Gather disagrees with itself 19.3% of the time overall, and 36.1% of the disagreements it records come back null on the second reading. Temperature is not the cause, which I had assumed it was: greedy decoding measured 14.7% self-disagreement at T=0 against 13.3% at T=1, slower and more verbose and still non-deterministic.</p><p style="text-align: justify;">Then the distinction that mattered. 19.3% of merge batches flip, but only 0.43% of the underlying passages change page, and every flip observed was a singular, plural or article variant. A surface changing group is not its evidence moving, and conflating the two had overstated the problem by a factor of thirty. I had set the acceptance bar at 5% before the run rather than after it, the result came in 3.5 times under, and the issue closed with the instability accepted and a written condition for re-opening it.</p><p style="text-align: justify;">At run level the analogue is a rule I hold to: one draw is not a measurement. A single brief has been measured moving 39% between two runs of identical code. It caught a model switch I had already shipped. The envelope pass was moved on a five-source A/B, then reverted when full-corpus regeneration produced 6 degenerate outputs where the incumbent had 0. The A/B had sampled the sources the challenger handled well.</p><p style="text-align: justify;">Where a pass cannot be made to reproduce, it gets demoted rather than trusted. A Gather finding is a retrieval hint and never a citation, and no gate scores an answer against one. In its harshest framing, the research report&#8217;s rather than the engineering report&#8217;s, Gather agrees with itself 53% of the time, and that is the figure a sceptical reader should hold. Where a check has nothing to measure at all, it reports not-scoreable, a third state distinct from pass and fail, and that state blocks release. A metric that vacuously passes on zero observations reads as a green light for a check that never ran.</p><h2>The panel, and the control that qualifies it</h2><p style="text-align: justify;">The last instrument is the one built to be trusted least. Finished papers are sampled to a sealed panel: at least three reviewer models per packet, each from a different training lab than the model that wrote the paper, each receiving the rendered paper, the resolved text of every passage it cites, and the bibliography. Nothing else. The seal is enforced by the dispatch path having no parameter through which tools could be passed, because an agent holding file tools will read the repository whatever its prompt says. Isolation you ask for is not isolation. Spread is reported as the error bar, since three reviewers splitting 1/3/5 and three agreeing on 3 are different results and must never render identically.</p><p style="text-align: justify;">None of that would count without the control. Model judges are systematically generous and are moved by confident prose, so the panel reports no number until it has caught defects planted in a paper it was not told was a control. Three were planted: a fabricated cause cited to a passage that says nothing of the kind, the section stating the opposing case replaced with a caricature, and a high confidence band asserted over coverage the paper&#8217;s own map discloses as thin. The current panel caught 3 of 3, unanimously, across three reviewers. An earlier, cheaper, single-vendor panel caught 0 of 3, and every number that panel had produced was discarded.</p><p style="text-align: justify;">The control proves the reviewers catch defects that are there. Nothing tests whether they invent defects that are not. The panel&#8217;s false-positive rate is unmeasured, and &#8220;trusted&#8221; here means the instrument catches planted defects, never that it does not manufacture them.</p><p style="text-align: justify;">The panel then found the part I did not design for. Five of six reviewers independently flagged the same citation: it resolves cleanly, to a real passage, which happens to be a publisher&#8217;s catalogue page. Every mechanical check passes it, because machinery can verify that a marker points at a real passage and cannot judge that the passage is a catalogue entry. Four reviewers independently reached a second diagnosis, that the keystone claim in both papers is carried by evidence from two other countries, because the library holds no passage on the specific thing the claim needs. That one is a corpus gap presenting as a citation defect. No retrieval work fixes it. The fix is a book.</p><h2>Where the pattern starts to break</h2><p style="text-align: justify;">Axial was run to 35 sources, and the point where the pattern starts to break was measured rather than guessed at. Passage-to-position count grows roughly linearly as the shelf grows, at k = 1.04. The graph, though, densifies: cross-book relations rose from 8.7% to 38.4% of the total across corpus sizes, and have not plateaued. Read together, those two say that extraction scales and selection does not. Somewhere near 100 sources the hard problem stops being &#8220;read the book&#8221; and becomes &#8220;choose which of the thousands of things now arguing with each other should go in front of a model&#8221;.</p><p style="text-align: justify;">That is an empirical warning drawn from a handful of corpus sizes on one shelf, not a law, and I would expect the number to move with the domain. It is also the one thing an idea file cannot offer, and the reason I would rather publish a measured limit than a working demo.</p><h2>What none of this measures</h2><p style="text-align: justify;">Every test question in both reports was written by an AI model working from a description of the library, not by a working scholar. That was deliberate: it let the engine be built and hardened without waiting on anyone, and nothing simulated was ever used as an answer key. It also bounds every figure above. They measure the engine. None of them measures answer quality against a real research question.</p><p style="text-align: justify;">There is no human referee in the loop. No number here may be described as measured against human expert judgment, and the specification forbids relabelling one that way for as long as a panel of models is the only referee.</p><p style="text-align: justify;">The denominators are small, and they are printed for that reason: 215 citation markers, 105 claims, 32 cross-source inferences, two papers, five new cross-source inferences. Five is five, not a rate. A perfect score over a small n is a small piece of evidence, and the paper gates passing at 1.0000 is exactly that.</p><p style="text-align: justify;">The mechanical gates measure construction, not evidential sufficiency. Two of the eight papers written so far cite exactly one book each, pass every gate, and rate strong on shape, because the shelf holds one comparative study of unrecognised states and no monograph on either territory they ask about. That is the library reporting its own shape, never a finding. The shelf is built around one case, Syria, inside the comparative-historical literature on state formation, nationalism and political violence, and every paper inherits that. Only new books close that gap.</p><p style="text-align: justify;">And the word that never attaches to any of this output is &#8220;correct&#8221;. There is no gold standard for the right synthesis of thirty-five books, so Axial does not claim correctness. It claims auditability. Every record carries where it came from, what kind of statement it is, and how much confidence its coverage entitles it to, and code put all three there rather than a prompt asking for them. What comes out is <strong>governed data</strong> rather than model prose. Confidence is capped by evidence, not by tone. A claim in a finished paper can be walked back to the passage in the book, by hand, in under a minute, and that is the whole of what is claimed.</p><h2>What I would like judged</h2><p style="text-align: justify;">The narrow claim is this. The LLM Wiki pattern holds at 35 sources, and compiling a corpus once and querying the compilation produces structure that query-time retrieval does not: 9,505 names shared across books, 447 stated disagreements, 43,101 cross-source opposition pairs, and an argument map whose relation labels were coined rather than chosen from a list. The parts that are unreliable are unreliable by measured amounts, and each is routed around rather than declared solved. I have not solved non-determinism. I measured it and built around it, which is the weaker-sounding claim and the only one the evidence supports.</p><p style="text-align: justify;">All of it is reconstructible from the repository: the specifications, a 67-entry decision log with its reversals annotated in place rather than rewritten, the committed run logs, and a codebase whose docstrings carry the measurements that set its constants. The paper trail is the codebase, and that is the property I would most like to be judged on. The research report and the engineering report are where to start.</p><p style="text-align: justify;">The instrument this project does not have is a person who knows the field. Three real research questions from your area, of the kind you would put to a doctoral student, would change what every number above is allowed to mean. One refereed reading of one paper and the passages it cites would be the first of its kind here. And if you can name the book this shelf is missing, that is one sentence from you and a gap I cannot close alone.</p><h2></h2>]]></content:encoded></item></channel></rss>