Tolerance panels · the instrument that judged every edit to this post
Green in-lane · amber a little out · red drift. Every panel is a real commit, byte-identical on recompute. Tap any panel to open its shareable receipt.
Geometric Driven Development — 1 measured edit to this post. Recompute any of them yourself, in a clone of this repo: npx thetacog-mcp publish-commit --commit ae776082f
The best thinking I do happens on a footpath, out loud, to a machine that does not get bored. An hour of a man describing an architecture he does not yet understand. What comes back is not code and not a plan — it is a transcript, long and half-wrong, of a thing being learned in real time. Then: "make a spec out of this." The spec is what survives the walk. That part we got right early. What we got wrong for months was everything that happens next — we took the good spec and built it in one enormous conversation, and watched it curdle around hour three, and blamed ourselves for not explaining it well. So here is the method re-plated, including the failure, because the failure is the interesting half.
Every course below is served the same way, and the plating is the argument. First the maître d' presents the dish — flourish, tongue slightly in cheek, the course named as an object before a word of argument reaches the table. Then the inner monologue: the exact sentence the course is built to make you think, written down before you read it. That is not a wish about your reaction. It is a prediction you get to grade. Publishing it in advance is the same trust-inversion as the command that opens the meal, shrunk to the size of one section — showmanship converted into an attackable claim. It is also the opposite of manipulation, because manipulation needs the dark and this is printed on the menu before you taste anything. If a course ends and its sentence did not fire in your head, the course failed and you caught it. Catching it is the meal working anyway.
And a recipe has something a wish does not: a win condition, declared before the first plate. Ours is not your agreement. This meal wins if you leave the table and recompute — run the command, check the number against your own transcript directory, swing at the claim. It fails if you leave merely nodding. Nodding is what a good essay produces. Recomputation is what a working instrument produces. Grade us on that.
A
Loading...
🥂Amuse-Bouche — Why We Believe You Never Have to Take Our Word
The maître d', presenting:The Unshucked Shell — served closed, the command resting on the linen beside it. The kitchen could open this for you and show you a lovely interior; it declines. What the house guarantees is only that nothing was done to it behind the pass — which is the one guarantee a demo that runs on the vendor's laptop can never make you.
Inner monologue it should trigger:"They didn't ask me to believe anything — they handed me the thing and stepped back."
the humble open · the command before the claim · your machine not ours · authority held in reserve
Before any of what follows, run npx thetacog-mcp attest-demo on your own machine. Here is exactly what comes back: a signed placement verdict — in-domain, out-of-domain, or unplaced — for a piece of work, computed with no model anywhere in the path. It runs locally. Nothing leaves your building, which also means we cannot cook it for you. That is the point of serving it first: the proof runs before the belief, so nothing below has to be taken on faith.
There is a second thing you can check, and it costs you thirty seconds. Look in your own agent's transcript directory and find the largest session file. We will put our number on the table first so you can see whether yours rhymes: one working session on this repository, one sitting, one body of work — 41.7 megabytes, 12,939 messages, on the order of ten million tokens of conversation. No context window holds that. No window is going to; the growth is in the work, not in the buffer.
The ingredients on this small plate, and they only work cold and together: the attackable claim you are invited to swing at rather than swallow; the local command that costs you nothing and proves we are not hiding the ball; your own measurement, which is better evidence than ours because you cannot suspect us of choosing it; and the authority we do have, kept visibly in the chair. What you should feel is not this person is impressive. It is wait, I am the one holding the verdict here.
🥂 A → B 🔥
The ladder — five rungs, each one attackable. Reject any one of them and the ladder breaks; tell us which rung it was.
1. Programming was never bottlenecked on typing — it was bottlenecked on learning what the thing actually is, and that is why requirements could never be written up front.
2. Talking to a machine collapses that learning time, which means the bottleneck moved rather than disappeared — and it moved somewhere nobody instrumented.
3. It moved to adjacency. When one context holds several unrelated bodies of thought, retrieval degrades quietly, and the degradation looks exactly like a weak model or an unclear operator. It is neither.
4. A bigger map does not fix this, and neither does a bigger window — because the thing that broke was the address of a symbol, not the amount of knowledge available. Your semantic layer is real and correct for what it decides; this is the one rung it does not reach.
5. Once every action carries a provable address, deviation from spec becomes countable — and countable is the difference between an exposure an enterprise absorbs and one an underwriter will price.
You cannot settle rungs 3 through 5 by thinking about them. Run npx thetacog-mcp attest-demo and swing at the output.
B
Loading...
🔥The Why — Five Centres of Mass on One Plate
The maître d', presenting:Tumbling Terrine, Five Centres — a terrine set with five separate cores, none dominant. It will not sit still on the plate and it will not tell you why; it simply drifts off centre while you are looking at the wine list. Not to be confused with the confidence score, which at least has the courtesy to be wrong in one place.
Inner monologue it should trigger:"That's the thing that happens at hour three, and I've been blaming myself for it."
the one belief · felt before proven · why fast work degrades · relevance is not the scarce resource
Programming was always about the rate at which you learn. You cannot know, at the start of a real project, what you will need to know by the middle — if you could, it would not be a project, it would be transcription. The entire cost of building anything is the cost of finding out what it actually is. Every argument about requirements that ever happened was this argument: the requirement could not be written, because the thing that would have made it writable was on the far side of building it.
So the constraint was never intelligence and never keystrokes. It was learning bandwidth, and for the first time since the field existed that number is not fixed. That is the upside, and it is larger and quieter than the one being marketed. The pitch says the machine writes the code. The truth is better: the machine collapses the time between not knowing and knowing, and knowing was always the bottleneck.
Then the spec is good, the build starts, and it goes wrong anyway. Because it is fast, you move in five directions at once — the data layer, the render path, the guard, the copy, the deploy. Each is a coherent body of thought with its own vocabulary and its own way of being correct. And they are all in one context. What you have built without meaning to is a body with several centres of mass: not one heavy thing but five, each pulling, none dominant. In physics that is not a stable object, it is a system that tumbles. In a context window it is worse, because the window does not report that it is tumbling. It just gets vaguer, and starts answering the question adjacent to the one you asked, and the specific fact from forty minutes ago — the one the whole design turns on — is still in the transcript and no longer in the grip.
Here is the part that stings: the machine loses hold of your repository, and both of you look like fools. You look like someone who cannot explain what he wants. It looks like a system that cannot follow. Neither is true. The transcript overflowed, the centres multiplied, and what fell out was never the intelligence. It was the address.
Every one of those 12,939 messages was relevant when it was written. That is precisely the problem. Relevance is not the scarce resource — adjacency is.
🥂🔥 B → C 🤝
C
Loading...
🤝Connection — The Cache Miss You Have Been Paying Without a Counter
The maître d', presenting:Cache-Line Carpaccio, Sixty-Four Bytes Thin — sliced to the width the hardware actually fetches. Eat it in order and it costs a nanosecond; eat it out of order and it costs seventy-five, and the kitchen will not tell you which one you just did. The bill arrives either way.
Inner monologue it should trigger:"There's no counter for this, which is why I've never been able to point at it."
your own session · spatial locality · the miss nobody counts · misattributed to the wrong party
You have felt this and you have almost certainly filed it wrong. Sixty-four bytes of silicon make a bet every time they load: things stored near each other are related. When the bet is right the next access costs about a nanosecond; when it is wrong it costs about seventy-five. The hardware is not grading your code. It is measuring whether your layout matches your meaning, and it has been doing that since 1985.
A context window makes the same bet one storey up. It assumes what is near in the conversation is near in meaning. In a session with five centres of mass that assumption is violated exactly the way a hash table violates it — related things scattered, adjacent things unrelated. The difference is the accounting. Intel gives you register 0x412e for last-level cache misses: not approximately, exactly, every miss, every time. There is no register 0x412e for a mind losing the thread. So the miss happens silently and gets billed to whoever is nearest — the model, for being weak; or you, for being unclear.
This is the symbol grounding problem wearing this decade's clothes. It did not go away and scale did not solve it. A symbol whose referent has drifted out of reach is not a slightly degraded symbol — it is a token with no address, and everything downstream of it is fluent, confident, and pointing at nothing. Grounding is still king. The book works this through in the section on what actually breaks, and the cache rhyme in the one directly after it.
The ingredients: the locality bet, made by hardware and by transcript alike; the missing counter, which is why this has never been a line item; the misattribution, which costs you confidence in a tool that was working; and the drifted referent, which is the actual failure under all three.
🥂🔥🤝 C → D 🎁
D
Loading...
🎁Contribution — The Spec Your Team Can Hold Without You in the Room
The maître d', presenting:Spec Off the Footpath — an hour of walking, reduced. Served as a single sheet rather than the stockpot it came from, because the stockpot is where the requirements document went to die in every project you have ever staffed.
Inner monologue it should trigger:"Every handoff I've done still routed every real question back to me — that's the thing being named here."
what you hand over · the transcript reduced · orthogonal not sequential · a slice per vocabulary
The walk produces a transcript. The transcript is not the deliverable — it is the stock. The deliverable is the reduction: make a spec out of this. And then the move almost nobody makes, which is the one that pays: cut the spec into areas that do not bleed into one another — not by file, not by size, but by whether they share a vocabulary. Each area becomes a single centre of mass. Each gets its own thread, its own build, its own place to be wrong without contaminating the rest. The tumbling body becomes several stable ones.
You already know what the absence of this costs, because you have handed something off. You wrote the doc, you walked someone through it, and it worked exactly as long as you stayed reachable — every real question came back to you anyway, because the doc had the what and you had the why. That is not a documentation failure. A spec that still needs you in the room was never sliced; it was summarized.
Sliced by vocabulary, it stops needing you. Two people who never speak can hold two slices at once without producing work that has to be reconciled later, and that is the actual constraint on a small team — not headcount, reconciliation. We run it literally: each slice goes to a different terminal with its own spec and its own done-when, written to a shared record that says who owns what. Three slices of this post's follow-on work went out that way while it was being written, and none of them needed a branch or a meeting to exist.
The ingredients: the transcript as stock, not product; reduction as the deliverable; slicing by vocabulary rather than by file, which is the only cut that stops two slices from undoing each other; and a slice that carries its own done-when, which is what lets it leave your hands without leaving your head.
🥂🔥🤝🎁 D → E 🌱
E
Loading...
🌱Growth — The Napkin With Your Position On It
The maître d', presenting:The Twelve-by-Twelve Napkin — one flat square, folded from a shape with far more dimensions than the table has room for. Everything you need is printed on it and nothing else is, which is the courtesy the full rulebook has never once extended to a working engineer.
Inner monologue it should trigger:"So I stop shipping the whole rulebook into every prompt and ship the three rules that actually apply here."
projection not summarization · the coordinate as an address · magnetised rules · the hat you wear here
Slicing stops the bleed. It does not, on its own, keep any single slice on target. The second move is that every prompt gets placed before it runs: the thing you are about to do is projected out of the many dimensions it actually lives in and onto one flat sheet of coordinates — a square in a twelve-by-twelve grid, an actor crossed with a patient. A position, not a summary. The distinction matters: a summary is lossy and arguable, a position is an address and can be recomputed by someone who disagrees with you.
Because it has an address, the rules that live at that address can be magnetised to it — not the whole rulebook, which is noise wearing the costume of thoroughness, but the three or four rules that are load-bearing exactly here, plus the role you are supposed to be occupying while you do it. This post was written under that constraint. So was the commit that carried it.
And this is where the honest limit goes, because a manual that only describes its best day is marketing. We do not run this perfectly. Some days the spec is thin, the placement is wrong, and three commits later I find that two centres of mass have been quietly sharing a context and undoing each other. The method is real; the discipline is uneven. What makes it survivable is that the placement is written down every time — so a day I did it badly is a day I can find. That is the whole difference between a practice and a mood.
The ingredients: projection, not summarization; the coordinate as a recomputable address; magnetised rules instead of the full rulebook; and the written placement, which is what converts a bad day into a locatable one.
🥂🔥🤝🎁🌱 E → F 🌊
F
Loading...
🌊Uncertainty — The Course That Comes Back However Well You Order
The maître d', presenting:The Undecidable Course, Returned Twice — sent back to the kitchen for being wrong, returned identical, sent back again. The kitchen is excellent. The kitchen is not the problem. The class of enterprises currently ordering more context as the remedy will recognise the taste.
Inner monologue it should trigger:"Good enough is a real position — but I've never priced what it actually costs me."
the honest opposing bet · Rice · what good-enough costs · the ceiling nobody puts on a slide
There is a great deal of money going the other way right now and it deserves an honest hearing rather than a sneer. The dominant bet is context: map the enterprise's knowledge completely, connect every entity to every other, give the agent perfect situational awareness, and the failures stop. It is a beautiful idea. It is the idea we would have had. Within its own frame it is not wrong — better context genuinely produces better behaviour, and the people building semantic layers are building real things that solve real retrieval problems.
The difficulty is a quiet assumption underneath: that sufficient knowledge implies sufficient behaviour — that a well-enough-informed agent is a safe one. Whether an arbitrary program will behave well is not a hard question, though. It is a decided one, and the answer has been in since 1953: you cannot determine it in general, at any budget, with any amount of context. More map does not touch it, because the obstacle was never ignorance of the terrain. We argued this badly for a year by swinging at Rice head-on; the better version, and where the gate genuinely is decidable, is in Pydantic at the Door, Rice at the Table.
The objection that arrives next is the good one, and it sounds fatal: there are infinitely many ways for an agent to go wrong, so you cannot count them. True, and it does not matter, because those are two different words doing two different jobs. The ways to deviate are functionally infinite — unbounded in kind, unenumerable in advance. The deviations are finite, because each one happened at a place at a time. An infinite alphabet still writes finite sentences. The fire marshal cannot list the ways a building might burn and prices the risk anyway, because what gets counted is fires, not the space of possible fires. Every attempt to bound the kinds of AI failure — enumerate the harms, taxonomize the risks, write the exhaustive eval suite — is a hand reaching into an infinite set. Nobody has to. Leave the kinds infinite and count the instances: spec, deviation, event, count, frequency, price, and no step in that chain needs to know in advance what the failure will look like. The book works the full argument in Countable, Not Accountable.
So the honest question is not is good enough acceptable. Often it is. The question is what good enough costs, and it has a name: an enterprise running agents without a defensible record is self-insured, and self-insurance is not paid in premiums. It is paid in a deployment ceiling — you may deploy exactly as many agents as you can afford to eat the losses on. Nobody writes that ceiling on a slide, which is why it is the most expensive line item in the building.
The ingredients: the opposing bet, stated at full strength; the assumption underneath it; the 1953 result that no budget moves; and the unwritten ceiling that self-insurance actually buys.
🥂🔥🤝🎁🌱🌊 F → G 🔒
G
Loading...
🔒Certainty — The Second Pour, Identical to the Sediment
The maître d', presenting:Sediment-Free Second Pour — the same bottle, poured again in front of a guest who came here hoping to catch us. Identical to the sediment or the house is wrong, publicly, at its own table. Telemetry is welcome to try this; telemetry is a photograph of the exhaust.
Inner monologue it should trigger:"Reproducible by someone who wants it to be wrong is a different category from a dashboard."
determinism as a property · the adversarial rerun · exhaust versus event · what cannot be sandbagged
Assume the crash. Rice says to. What remains available is not was this good — that is undecidable and always will be — but where did this land, which is a fact. The record is computed as a deterministic function of the work itself: same input, same reading, byte for byte, no model anywhere in the path. Open it, then open it again six months later from a different machine, and get the identical value.
That property is what a hostile party needs, and it is the one thing a software-layer imitation cannot supply. When this market realises that context was never the bottleneck, the fast follow will be liability receipts assembled from telemetry — token counts, latencies, keyword filters, trace spans. That is behaviourism: it measures the exhaust after the fact and infers the event. An inferred record can be sandbagged, because whatever produced the inference can be tuned to produce a friendlier one. A record that is the event has nothing between it and the thing it records to forge. The book makes the general form of this argument in the section on refusing to erase.
The ingredients: determinism as a property, not a promise; the adversarial rerun as the only test that counts; exhaust versus event, which is the whole difference between telemetry and a receipt; and nothing left in between to forge.
🥂🔥🤝🎁🌱🌊🔒 G → H 👑
H
Loading...
👑Significance — The Radio the Tugboat Did Not Carry
The maître d', presenting:The Carving Knife and the Missing Radio — carved at your seat, from a 1932 cut. Two barges went down in a storm; the tugboats had no radios; radios were cheap and available and not yet customary. The court declined to accept "nobody else had one either." The dish is served cold on purpose.
Inner monologue it should trigger:"I've been in that meeting, and we couldn't answer the question — that's the thing this fixes, not the insurance part."
who moves first · the standard of care · negligence per se · the officer who signs
You have sat in the meeting this is about, and it had nothing to do with insurance. Something an agent did got escalated, and someone senior asked the only question that mattered — what was it actually operating on when it did that? — and the honest answer was a scroll through logs and a reconstruction. Not a fact. A reconstruction, assembled by the same team being asked. Nobody called that a liability problem. It just cost you a week and left the question unanswered.
That is the whole thing, and it is why you do not have to wait for anyone. The version of this that reaches you first is not a carrier mandate — it is your own counsel or risk officer noticing that a deterministic check exists, costs nearly nothing, and is now on the record as available. There is a 1932 case behind why that noticing is enough: a fleet of tugboats lost two barges in a storm without radios, radios were cheap and not yet customary, and the court declined to accept nobody else carried one either. A cheap available safeguard you did not use is your problem whether or not your competitors also skipped it. Which means the deadlock everyone describes — carriers waiting for a standard, enterprises waiting for cover — is not actually yours to wait out. It breaks from inside the building. The longer version is in Incidents, Countable.
And the practical version, which is smaller than it sounds: the next time that meeting happens, you answer the question in one sentence instead of a week — because the record was emitted when the work ran, not reconstructed afterwards by the people being asked about it.
The ingredients: the cold-start problem named honestly; the 1932 holding that removes the excuse; the internal forcing function, which fires before any carrier does; and the seat you occupy once you have moved.
🥂🔥🤝🎁🌱🌊🔒👑 H → I 🐉
I
Loading...
🐉The Turn — We Are Not Selling You the Physics
The maître d', presenting:Dragon's Breath, Flambéed at the Table — the match we have been holding since the first course. It is a brief flame and it is not the meal; anyone who came for the flame has misread the menu, and we would rather say so now than take their money.
Inner monologue it should trigger:"They just told me the technology isn't the product — which is the first honest thing I've heard in this category."
the declarative constraint · what is actually being sold · the standard others adopt · precedence over persuasion
Now the authority, held back until it is the last thing rather than the first. Everything above about cache lines, addresses, and grounding is the moat, not the product. If we spent your attention explaining the physics, we would deserve to be filed under interesting research and forgotten by Tuesday.
What is actually being sold is a flight recorder that makes the agentic economy insurable — a trigger that a carrier can write against and an adjuster can adjudicate. The physics exists for exactly one reason: to make that record legally defensible when someone with an incentive to break it tries. A parametric policy pays on a measured parameter, not on a proven chain of causation; hurricane cover pays on wind speed and nobody demands a study correlating wind to damage first. The measured boundary is the trigger. That is the entire product, and the rest is how it survives a hostile reading.
The reframe is not that we are right and the semantic-layer builders are wrong. It is precedence: once a deterministic placement record is cheap and available, it becomes the artifact every other layer in the stack has to be able to emit — the map, the eval harness, the guardrail, all of them. Not because we won an argument. Because the first claims dispute will ask for it, and after that it is customary.
There were no cars on the road before there was a way to write the policy. The engine was never the blocker. The social contract was — and the contract needed a countable event before it could exist at all.
🥂🔥🤝🎁🌱🌊🔒👑🐉 I → J ☕
J
Loading...
☕Digestif — The Ingredients, Not the Conclusion
The maître d', presenting:L'Addition, With the Walk Attached — the bill, and stapled to it the recipe, the sources, and an invitation to cook it worse than we did and tell us where it broke. The house has been wrong before and prefers to be told at the table.
Inner monologue it should trigger:"One command, thirty seconds, and I find out whether their number is anything like mine."
evidence last · sources not verdicts · the to-do · your move
Here is what is on the record, handed over as ingredients rather than as a conclusion you are supposed to draw. Rice's theorem (1953): every non-trivial semantic property of a program is undecidable — the load-bearing constraint under all of this, and the one worth attacking first if you want to break the argument. The T.J. Hooper (2d Cir. 1932): an industry-wide failure to adopt an available precaution does not excuse the omission. The cache hierarchy, shipping since 1985, which has been reporting spatial-versus-semantic locality at nanosecond resolution the entire time — Intel's 0x412e counts the misses whether or not anyone reads it. And your own transcript directory, which is the piece of evidence we cannot have selected for you.
The to-do, and it is two lines long. If you check one thing, check the cheap one. Your largest session transcript, one line, no signup:
ls -lS ~/.claude/projects/*/*.jsonl | head -1
Ours came back 41.7 MB. If yours is in the same order of magnitude, the argument in course B is describing something you have already paid for. If it is nowhere close, we would genuinely like to know — that is a result either way, and it is the one number in this post we could not have chosen for you. Then, when you have thirty seconds more, npx thetacog-mcp attest-demo.
And the win condition, declared before the first plate. Ten courses, ten sentences, each printed above the course that was built to produce it. Scroll back and count the ones that actually fired. Not whether you agreed — whether the sentence was already in your head before you read the paragraph under it. Four or more and the format worked on you. Fewer and it did not, and you now hold a specific list of which courses failed, which is worth more to us than a nod and is the reason the predictions were printed in the first place.