The Third Build Is the Factory
Twelve judgment instruments came off one chassis in a month, and the one that mattered found a fail-open in software Microsoft shipped. The compounding move was not building a good tool. It was noticing the third time I built the same shape and extracting the skeleton then, not the first time and not the tenth.
The compounding move in tooling isn’t building a good tool. It’s noticing the third time you build the same shape, and extracting the skeleton right then. Twelve judgment instruments came off one chassis in a month, and the one that matters found a fail-open defect in software Microsoft shipped.
The most defensible finding my tooling produced this summer is a defect in software Microsoft shipped. A tribunal I built to rule on my own problem, whether thirty-one MCP servers would survive an incoming revision of the protocol spec, refused to rule on the Power BI server from reading its source alone. The case came back UNRULEABLE. So the verdict was re-grounded the only way left: by executing the shipped IL, the compiled bytes that actually run, and watching the failure happen. The verdict moved to BREAKS, and the mechanism was the bad kind, a fail-open, the server proceeding as if a check had passed when the check had changed underneath it. An upstream report was drafted. Two weeks later the spec revision went final, and a delta pass against the final text confirmed both open contingencies: the finding held, and the next major version of the SDK does not rescue it.
I am opening with that finding for a reason that has nothing to do with the bug. The tribunal that caught it went from design document to a full 31-case ruling, 155 verdict rows, in about a day and a half. It could move that fast because it was roughly the ninth instrument of its kind on this machine, and by the ninth, building one no longer means designing anything. It means filling in a skeleton that already knows what a judgment owes you.
The skeleton is the story. And the principle behind it is worth stating plainly, because I got it right mostly by accident and only understood it afterward: the compounding move in tooling is not building a good tool. It is noticing the third time you have built the same shape, and extracting the chassis right then. Not the first time, when you cannot yet tell which parts of the design are essential and which are incidental to the one problem in front of you. Not the tenth, when you have paid the duplication tax seven more times than you needed to. The third.
The shape that kept recurring
Start with what kept getting built, because the pattern is only visible in retrospect and I want you to see it the way I eventually did.
I run a lot of my operation on what I have started calling judgment instruments. The idea has an essay of its own (Sermons vs. Instruments): a rule written in prose binds an agent most of the time, and “most” is the entire problem, so recurring judgments should be embodied in tooling that cannot be talked out of its standards. But “instrument” undersells what these things actually are. Each one is a small court. There is a rubric, versioned like software, that says what the judgment turns on. There is a verdict schema and a validator, so a ruling that skips a required field or cites no evidence is rejected as malformed before anyone reads it. There is a docket assembler that gathers the cases and pins the evidence each verdict must cite. And there is a ledger, so every ruling is durable, auditable, and attributable to the rubric version that produced it.
The first of these I wrote about was the repository tribunal, which put all 182 of my repos on trial and killed 39 of them with citations (Kill With Dignity). What I did not write about was what happened in the three days around it, because the interesting part had not finished happening yet.
In mid-July the same shape got built again. And again. A corpus court, to rule on the lifecycle of 240 entries in my agent memory. A drift court, to rule on documentation that had wandered away from the code it described. A consolidation map, to arbitrate borders between overlapping projects. A regret oracle. A canon court. Each one was a genuinely different judgment about genuinely different material, and each one kept needing the same organs: rubric, schema, validator, calibration, ledger.
Somewhere in there, at the third build, the shape stopped being a coincidence and became a pattern. The commit log records the moment better than my memory does. On July 17 a small repo called instrument-foundry appears with the commit “extract judgment-instrument chassis at n=3.” The same day, the next court to be built ran as the chassis’s proof run, and the friction it hit went back into the foundry as its own commit. By end of day the foundry’s doctrine header had been bumped from n=3 to n=4. The factory was operating the day it was founded.
Why three
The n=3 rule is the part I want to defend, because both failure modes around it are common and I have committed both.
Extract at n=1 and you are not extracting, you are speculating. You have one instance, so you cannot distinguish the load-bearing parts of the design from the parts that exist because of the particular material in front of you. Every abstraction you draw is a guess about a future that has not arrived, and frameworks built this way ossify their author’s first misunderstanding. I have a graveyard of them.
Extract at n=10 and the arithmetic is just bad. You knew the shape by the third build. Every build after that paid full price for organs you could have stamped: a rubric format designed from scratch, a validator written from scratch, a calibration protocol reinvented with slightly different discipline each time, which is worse than no discipline because it looks like discipline. Seven extra payments of a tax you had already diagnosed.
Three is where the evidence changes character. Two instances that share a structure can be coincidence, the same author solving adjacent problems the same way out of habit. A third instance, on different material, with different verdict semantics, needing the same organs anyway, is a pattern making a demand. And at three the cost of being wrong about the extraction is still small: if the chassis turns out to be misdrawn, you have misdrawn it for one consumer, not ten.
There is a quieter reason three works, and it showed up in the foundry’s own history. The chassis was not extracted from a clean design session. It was extracted from three builds’ worth of accumulated friction, and then the fourth build’s friction fed back the same day. A chassis extracted at n=3 is young enough to still be wet. It can absorb what the fourth and fifth builds teach it. A chassis extracted at n=10 has ten consumers with settled expectations, and every lesson after that is a migration.
What the stamping actually looked like
Here is what the marginal cost collapse looks like in the ledger, because “the fourth through twelfth collapse to hours” is exactly the kind of claim this site exists to make checkable.
Over roughly three days in mid-July, off the fresh chassis: the corpus court calibrated on a small docket under rubric v0.1, graduated its rubric to v1.0 with zero operator overrides, and ruled the full 240-entry docket. The drift court judged a six-case calibration docket, graduated at 33 of 33 rows with all six gates accepted, and went live. The regret oracle graduated at 82 of 82 rows, six of six gates, zero overrides. The consolidation map calibrated seven for seven validator-clean, and its calibration caught a real defect on the way in, a degenerate similarity score that had inflated its docket from four cases to seventeen, which the pre-registered expectations flagged before any live ruling depended on it. The repository tribunal calibrated on a 20-verdict sample, took an operator review of 19 agreement out of 20 with the one disagreement recorded as a named override, graduated to v1.1, and ruled all 182 repos to a final ledger of 70 keep, 73 park, 39 kill, zero unresolved. The spec-shift tribunal, the one from the opening, assembled 31 cases against a pinned draft of the spec and ruled all of them.
Each of those is a sentence here and was, at most, a day there. The design-shaped work, deciding what a verdict is, what it must cite, how a rubric earns the right to judge, had been paid for once, in the chassis. What remained per instrument was the part that genuinely differs: the rubric’s content, the docket’s material, the calibration’s known answers.
Two later builds are worth naming because they show the chassis traveling. The transcript autopsy lane, a completely different domain, session transcripts rather than repos or prose, was stood up with the commit “port the judgment chassis onto the Autopsy lane” and ran five blind sittings with registered predictions. And a calibration bench built on the same discipline closed itself at generation six: its generation-seven proposal was tested, refuted, and recorded as refuted, no rule added, lane closed. A factory that can also stop is not a throwaway detail. It is the strongest evidence I have that the gates are real, because a fake gate never says no to its own author.
The ritual is the product
If you take one implementation detail from this, take the graduation protocol, because it is the thing that makes stamping out judges safe rather than reckless. Twelve courts authored by one person in one month should alarm you. It alarmed me. The chassis’s answer is that no rubric touches real material until it has passed a calibration it could have failed.
The ritual, as the chassis enforces it: before an instrument rules on anything that matters, it gets a small calibration docket, real cases plus at least one synthetic plant, a case with a known correct answer seeded in to prove the judge can catch it. Expectations for the docket are pre-registered, written down before the verdicts exist, so the calibration cannot be graded on vibes after the fact. The instrument rules under a v0.x rubric. The operator reviews every calibration verdict. Only if the gates pass does the rubric graduate to v1.0, and any disagreement survives as a named override in the ledger rather than a quiet edit. The repository tribunal’s ledger shows the protocol holding under disagreement: 19 of 20, one override, recorded, rubric amended, graduated at v1.1.
And because a protocol run by its own beneficiary is exactly the kind of self-report this series distrusts, the protocol got pointed at itself. I graded 30 of the corpus court’s calls against my own, blind, on a pre-registered method with the seed and metrics fixed before any data: weighted kappa 0.78, 29 of 30 exact, and the one amendment to the method disclosed in the receipt (Grading the Grader has the whole protocol). That number is not a triumph. It is a measurement of how much a machine judge and I disagree, taken before I let the machine judge things I would not personally re-check.
The pre-registration habit earns its keep in a specific way: it converts calibration from a demo into an experiment. A demo can only succeed. An experiment can fail, and several did. Predictions got registered before sittings and then falsified by them. The consolidation map’s inflated docket got caught. The bench’s generation-seven design got refuted and stayed refuted. Each of those refusals is in a ledger, which means the next agent context that wanders in with the same bright idea finds a receipt instead of a blank page.
The finding that opens the loop
Which brings me back to Power BI, and why it is the exhibit rather than the twelve courts themselves.
Everything else in this essay is, in the end, my system grading my work under my rubrics. I have tried to make that self-referential loop as honest as pre-registration, blind sittings, and operator review can make it, but it remains one author all the way down, and I have written a whole series about why you should discount exactly that arrangement. The spec-shift tribunal is the one place the loop opened. Its material was other people’s software: thirty-one MCP servers, ruled against a pinned draft of an incoming protocol revision, asking one question per server, does this break, with every verdict required to cite the bytes it was grounded on.
The discipline transferred intact. When a first ruling on one cohort was challenged, the adversarial re-check was run and the verdict survived. When a live probe contradicted the stated failure mechanism for another case, the mechanism was corrected in the ledger rather than defended. When the Power BI case could not be ruled from source, the tribunal did not round UNRULEABLE down to fine, it escalated to executing the shipped IL, and the execution produced the finding: a fail-open, the worst polarity a compatibility failure can have, because it fails by proceeding. The upstream report was drafted from the ledger, citations already attached, because the chassis had required them all along. And when the spec went final two weeks later, the delta pass re-ruled the affected rows against the final text instead of letting a draft-based finding stand on momentum. Both contingencies closed. The finding held. The next SDK major does not rescue it.
One external finding is one finding, and I want to be precise about what it certifies. It does not certify that my courts are right about my repos or my prose. It certifies the factory’s claim at its weakest joint: that an instrument stamped from this chassis, pointed at material its author did not write, judged under rules registered in advance, produced a result that someone with no access to me, my machine, or my rubrics could check against shipped bytes. That is the property the whole series keeps demanding of everyone else’s systems. The factory’s ninth product has it.
Limits, counted honestly
The claims above have edges, and this site’s standing rule is to draw them before someone else does.
One author. Every rubric, every schema, every calibration gate in the cluster came from the same mind, and the operator reviews that gate graduation are reviews by the person who wrote the rubric being graduated. The kappa measurement narrows this, one blind protocol with the method fixed in advance, but it does not remove it. An off-family judge over these ledgers is the obvious next deposit, and it has not been made.
One consumer, mostly. The tribunals’ verdicts are consumed almost entirely by me and my agents. The Power BI finding is the exception, not the pattern, and a single external result is an existence proof, not a track record.
The count is closed, not growing. On August 11 the twelve instrument repos were enrolled in the portfolio catalog and the cluster was formally closed out in the decisions log. Twelve is the final number of this run, which is why I trust it enough to write down: it is a result with an ending, not a trajectory I am extrapolating.
And the rule itself is a heuristic with survivorship in it. “Extract at three” summarizes one cluster that went well, plus a graveyard of premature frameworks that argue against one, plus tax receipts that argue against ten. I believe the reasoning about evidence and cost, and I have stated it so it can be attacked. But I did not run the control where I extracted at five, and nobody ever does.
The move, restated
Three days in July, one extraction, and by mid-August: twelve courts, thousands of cited verdicts over repos, memory, prose, and protocol compliance, a graduation ritual that has said no to its own author, a lane that closed itself with its refutations on file, and one fail-open in shipped Microsoft software, reproduced from the compiled bytes, reported upstream, and confirmed against the final spec.
None of that is the point. The point is when the leverage actually arrived. Not with any court, and not with the chassis as an artifact. It arrived in the moment of noticing that the third build was the same build, and stopping to name what repeated. Before that moment, every instrument cost a design. After it, an instrument cost a rubric and a morning, which is cheap enough that judgments I would never have tooled, border disputes between projects, drift in my own documentation, got instruments too.
You almost certainly have a third build sitting in your history right now, wearing three different names. The first one taught you the problem. The second one felt like efficiency, because you remembered the answers. The third one is the factory asking to exist, and the only cost of saying yes is admitting that what you have been building, all along, was the same thing.