Afterword
It would be a strange book about auditing your own controls that exempted itself from the audit, so let me spend these last pages turning the loop on the book itself.
The largest hole is the one I named in the first chapter and never stopped naming: n=1. Every system here is real and every number is real, but one person built the systems, ran the measurements, and scored the results, and that person wanted them to work. I've shown the work so the wanting is at least visible: the bypasses reproduce, the provenance is on disk, the failures aren't inventions. But visible bias is still bias. The honest remedy isn't more confidence from me; it's a second operator. The result I'd most like to see is not a bigger version of any experiment in these pages: it's one reader reproducing a single failure on a harness I've never touched and reporting back whether it held. That's the peer review I'm actually after, and the only kind that would move the n off one.
The second hole is time. Every specific control in this book is a snapshot of a system that was moving while I described it, on a platform that was moving faster. By the time you read this, some of these guards will have been walked around in ways I didn't anticipate, some of the tools will have been renamed, and at least one of the measurements will have been overtaken by a model that didn't exist when I ran it. I've tried to write so that this doesn't matter: to make each chapter about the class of failure and the discipline that catches it, not the particular incident, because the incidents are perishable and the discipline is not. A guard that enumerates its enemies will lose no matter what year it is. A measurement whose provenance you didn't check will lie no matter how good the model. A map of your own system will drift the moment the system moves under it. Those aren't facts about this year; they're facts about operating something you didn't fully build and can't fully see. Hold the discipline loosely enough to swap the details, and you'll get more out of this book than I got out of writing it.
So here's the ask. Read this as a doctrine to test, not a result to cite. Take the move that bothered you most, the guard you think is too paranoid, the eval you think is theater, the memory store you think is overbuilt, and try to break it on your own fleet. If you break it, you've done me a favor and yourself a larger one. If it holds, the n has quietly become two, which is the only way a field this young learns anything worth trusting.
I'll keep operating the fleet, and I'll keep turning the loop back on myself, because the day I stop is the day one of these controls starts going quietly wrong with no one left to catch it. The work is never finished, that was the entire point. But it can be honest, and shown, and handed to the next operator with the receipts attached. The public receipts, Markdown mirrors, corpus JSON, and /llms.txt map are there so another operator can reproduce, argue with, and repair the work instead of merely believing it. That's what I've tried to do here. The rest is yours.