Tolerance panels · the instrument that judged every edit to this post
Green in-lane · amber a little out · red drift. Every panel is a real commit, byte-identical on recompute. Tap any panel to open its shareable receipt.
Geometric Driven Development — 7 measured edits to this post. Recompute any of them yourself, in a clone of this repo: npx thetacog-mcp publish-commit --commit 06e8fa729
Not the part everyone is arguing about. Nobody can decide whether an AI's answer was correct — for non-trivial semantic properties, no general procedure exists, and pretending otherwise is how this conversation has stalled for two years. But that was never the question an underwriter needed answered. They needed to know where the thing landed relative to what it was allowed to do, cheaply, months later, to somebody who does not trust you. That one is decidable. We made it cost a fraction of a cent and shipped it open.
EVERYONE ALREADY HAS THIS WHAT MAKES IT INSURABLE
───────────────────────── ───────────────────────
Prompt Prompt
↓ ↓
LLM LLM
↓ ↓
Output Output
↓
WHERE did it land? ← decidable · no model
↓ · anyone recomputes it
in-lane / out of lane / unplaced
↓
signed receipt
↓
┌───────────────────┴───────────────────┐
↓ ↓
deployer signs insurer signs
"that was my agent, "I accept this as trigger basis
that was its lane" — from these oracles"
└───────────────────┬───────────────────┘
↓
closed circle
↓
countable → priceable → insurable
That right-hand column is the whole meal, and the two signatures at the bottom of it went live today. The left column is what your vendors ship. The gap between them is why the capital that wants to fund autonomous deployment currently cannot.
Every course below is plated the same way, because the plating is the argument. The maître d' presents the dish — pure flourish, the idea named as an object before a word of argument reaches the table. Then the inner monologue: the exact sentence the course is built to make you think, written down before the course is served. That is not a wish about your reaction, it is a prediction you get to grade. Then the ingredients. Publishing the predicted sentence in advance is trust-inversion at the scale of one section — showmanship converted into an attackable claim. It is also the opposite of manipulation, which needs the dark; this is printed on the menu before you taste anything.
The win condition, declared before the first plate: not your agreement. The meal wins if you leave the table and recompute — run the command, watch a verdict get signed, then verify it with a key we do not control. It fails if you leave merely nodding.
A
Loading...
🥂Amuse-Bouche — The Clause We Cannot Take Back
The maître d', presenting:Le Contrat Irrévocable — a licence clause, plated alone, that removes a power from the house.
Inner monologue it should trigger:"Wait — they made it impossible to charge for the thing that proves their own product works?"
the grant · the fork test · the number and its interval · what remains licensed
Do you worry about $1.2B in AI liability?
If the property is trivial, software can check it — and why are you paying to check trivial properties? If it isn’t trivial, Rice’s theorem says nobody can. So we fixed the math.
press enter
Who did this make you think of? We’d love to know.
Here is the clause, added to the licence today and shipped to the public repository, verbatim: "Verification is unconditional, perpetual, and irrevocable. Any person may read the specification, run the math, fork the runtime, and recompute or independently verify any receipt, endorsement, or circle document — including receipts issued by ThetaDriven Inc. or by any licensee — forever, without permission, without notice, and without fee. No term of Part B, and no commercial agreement, may condition, meter, or withdraw this right."
That is Part A-1, and until this morning it was a sentence in a marketing document, which is to say it was not a commitment at all. A promise about verification made by the party being verified is worth exactly nothing, and the fix was not to say it more loudly.
Alongside it, a test that passes right now: a competing oracle — a forked runtime, a signing key we have never held, no involvement from us at any point — produces a verdict, and an insurer who names that key in their own accepted list closes the loop completely. circle returns CLOSED. The assertion is nine lines long and it is the load-bearing one in the file, because its inverse also passes: a receipt we sealed fails unless the insurer independently named us. Our own signature is not self-authorising. There is no constant in that file containing a blessed key, and a separate test fails the build if one appears.
The calibration behind all of it: 15.4% breach frequency, 95% interval 10.9–21.3%, sealed and pre-registered before it was read. That is a real number, honestly obtained, on a corpus we chose — the run against a repository the instrument has never seen has not happened yet, and when it does, whatever comes back gets published.
What stays licensed is narrow and worth stating plainly: building or monetising a financial instrument on one of these attestations. Running a fork is free. Verifying is free forever. Underwriting on the output is the licensed act, and it attaches to the act rather than to whose key sealed the artifact.
If "free forever" reads as a catch, here is the arithmetic that makes it not one. We were never going to be paid for verification, because an instrument nobody can independently check has no commercial value to begin with — the checking is what makes the number admissible, and a number that isn't admissible can't be underwritten by anyone at any price. So the clause gives away revenue we could not have collected and buys the one property the instrument cannot function without. The catch, if you want to call it that, runs the other way: it means we can never quietly become the gatekeeper later, which is a door we closed on ourselves on purpose and cannot reopen.
npx thetacog-mcp prove
A verifier that can be switched off is not a verifier — it is a vendor with a good reputation, and reputations are the thing insurance exists because you cannot price.
🥂 A → B 🪜
The ladder. Six rungs. Reject any one of them and the argument breaks — and if you do, that rejection is the most useful sentence you could send back, because it tells us exactly where this fails for someone who actually holds the risk.
Rung 1. Your existing controls are real and correct for what they decide. A policy engine that blocks an unauthorised call did its job; a code review that catches a bad migration caught it. Nothing below replaces either.
Rung 2. There is one question none of them answer after the fact, cheaply, to a hostile third party: where did this action land relative to what it was authorised to do? Not "was it allowed at the time" — "can you establish, months later, to someone who does not trust you, where it landed."
Rung 3. Today that question is answered by people at consulting rates, months late, producing a judgement another expert can contest. That is rigour, and it is expensive rigour.
Rung 4. An opinion cannot be pooled. You cannot build a loss distribution out of professional disagreement — so the exposure cannot be priced, so the policy is not written, so the deployment does not get funded.
Rung 5. Therefore the binding constraint on autonomous deployment is not the technology and not appetite. It is the unit cost of establishing what happened.
Rung 6. A determination that costs a fraction of a cent, arrives before the next decision, and returns the same answer to a stranger is a different kind of object than a cheaper audit. It is a statistic rather than a judgement.
You cannot settle rungs 5 and 6 by thinking about them. Run the command.
B
Loading...
💧Why We Believe — The Dam Is Not Made of Doubt
The maître d', presenting:L'Eau Retenue — water at a wall, served at pressure, with the wall explained.
Inner monologue it should trigger:"They're not sitting out because they don't believe. They're sitting out because nobody can tell them the number."
where the money stops · why refusal looks like scepticism · price as the real property
Watch where capital actually goes. It floods the model layer, the chips, the data centres — every part of the stack whose risk fits on a page an investment committee already knows how to read. It stops at the edge of deployment. Not at the edge of belief: at the edge of deployment, which is a different border entirely.
The book puts it more carefully than I can here:
"The same institution that will fund a training run will not underwrite the agent that run produces, and the people declining are not sceptics. They are enthusiasts with a fiduciary problem." — Tesseract Physics, Ch. 12
An underwriter who cannot bound an exposure does not write a smaller policy. They write none, because a number they cannot defend is worse for them than a deal they did not do. From outside, every one of those refusals reads as doubt about the technology. From inside it is smaller and far more mechanical: nobody can cheaply establish what happened.
Which relocates the interesting property. Rigour is available at any cost — the entire audit industry is rigour at high cost. What has not existed is a determination that is cheap enough to be routine. Cheapness is not a convenience feature here. Cheapness is what converts a judgement into a statistic, and only a statistic can be pooled.
🥂💧 B → C 🤝
C
Loading...
🤝Connection — You Have Already Priced Something Nobody Could See
The maître d', presenting:Le Pluviomètre — a rain gauge, on a white plate, saying nothing whatsoever about the harvest.
Inner monologue it should trigger:"This is the same move as parametric weather cover — and I already know that instrument works."
the instrument you already price · the question refused · why refusal is the credibility
Weather insurance does not ask whether the harvest was good. It asks whether rainfall fell below a stated threshold. The gauge says nothing about crop yield, farmer competence, or market prices — and that narrowness is precisely why the instrument settles in days instead of litigating for years.
We refuse the same question, in the same shape. We do not measure whether an agent's output was correct, safe, or good. For non-trivial semantic properties, no general procedure decides that — a fact with a name and a date attached to it, and not one we intend to argue around. We measure where an action landed relative to the scope it was declared authorised for. In-domain, out of domain, or honestly unplaceable.
That third verdict matters more than it looks. On a live run of the placement demo against twenty real commits, sixteen came back unplaced — nearer to noise than to any lane, so the instrument declined to guess. An instrument that reports its own blind spots is the only kind whose readings are worth pricing, and a tool that always returns an answer is a tool that is sometimes lying to you.
🥂💧🤝 C → D 🎁
D
Loading...
🎁Contribution — The Six Questions We Cannot Answer From Here
The maître d', presenting:La Clef Inachevée — a key, deliberately unfinished at the bit, served beside the lock.
Inner monologue it should trigger:"They're not asking me to buy. They're asking me to specify — and I actually know the answer to number five."
what we shipped open · what only a risk team knows · why we would rather be corrected in public
The primitive is live and the endorsement schema is deliberately at version 1, because the last mile is not derivable from first principles. It is knowledge that lives inside people who clear claims for a living. The questions are published in full in ATTESTATION-ASK.md in the repository; a one-line answer to any of them is genuinely useful:
Trigger sufficiency — a closed circle currently asserts this payload, against this declared scope, produced this verdict, and both counterparties signed it. Is that enough to fire a parametric trigger, or is a field missing?
Resolution — scope is declared as a set of authorised cells. Is that granularity meaningful in policy language, or does it need to be coarser or finer?
Aggregation — receipts pin to a single Merkle root. Is a periodic root the right reporting unit, or do you need every receipt?
Time — endorsement timestamps are self-asserted and supersession is by latest. Acceptable for a trigger, or does it need an external anchor?
Schema fit — what shape drops into a Solvency II or NAIC path without a translation layer?
Revocation — there is none; a wrong endorsement is corrected by a superseding one. Right, or does a claims process need explicit revocation with reason codes?
If you do this for a living and something above is wrong, saying so costs you a sentence and saves us a year. And if it is not your field — the person it belongs to is probably one forward away, and being the person who sent it is a real contribution to them.
🥂💧🤝🎁 D → E 🌱
E
Loading...
🌱Growth — Two Signatures Turn a Reading Into an Obligation
The maître d', presenting:Le Cercle Fermé — a ring closed with two seals, served on the record it binds.
Inner monologue it should trigger:"A measurement nobody has agreed to act on is just telemetry. This is the part that makes it a contract."
measurement vs obligation · the two roles · who names the oracle
Until today the chain produced a verdict and stopped. A spec published, a payload submitted, a verdict sealed, a stranger able to recompute it. All true, all useless commercially, because nothing said what the measurement obligated anyone to do.
Two signatures close it. The deployer signs: that was my agent, that was the lane I declared, I accept this record of it. The insurer signs: I accept this as trigger basis, under these terms, from these oracles. Both signatures bind to the receipt hash and the verdict, so an endorsement cannot be lifted onto a different receipt or survive an edited verdict — both are tests, both pass, and both were written before the feature was called done.
The fourth field on the insurer's signature is the one that took the longest to get right: accepted_oracles. They name whose word they take. We never adjudicate it. An insurer who names nobody is refused by the schema itself — an endorsement that obligates nothing is a schema error, not a permissive default.
🥂💧🤝🎁🌱 E → F 🌫️
F
Loading...
🌫️Uncertainty — The Four Things We Will Not Claim
The maître d', presenting:L'Assiette Vide — an empty plate, presented deliberately, with its contents named.
Inner monologue it should trigger:"They just told me the weakest parts before I found them. Which means I can probably trust the rest."
consequence · corpus · what we measure at the metal · the missing series
We do not measure consequence. Whether landing outside a declared scope causes a loss of any size is not something we have measured, so we do not assert it. That correlation is the modelling work of whoever performs it, and we would be lying if we pretended to have done it.
The calibration is on our own corpus. 15.4%, sealed and pre-registered — and chosen by us. The unseen-repository run is next, and its result gets published whichever way it goes.
We measure execution latency, unprivileged. Raw hardware event registers sit behind a privileged interface we deliberately do not require, because demanding root puts an instrument outside the sandboxes it needs to live in. A published post of ours once implied we read those registers directly. It was wrong, it is corrected across five surfaces, and a test now fails the build if the claim reappears anywhere.
We are pre-series. The ninety-day longitudinal record has a start date that has not been pinned. Until it is, each receipt is a datapoint and not a trend, and counting them as though they were one would be the easiest available dishonesty.
🥂💧🤝🎁🌱🌫️ F → G 🔒
G
Loading...
🔒Certainty — What Runs, Right Now, On Your Machine
The maître d', presenting:Les Deux Empreintes — two digests from two processes, plated side by side, identical.
Inner monologue it should trigger:"I can check every single one of these in under a minute without asking them for anything."
determinism · no model in the verdict · the guards that hold it
The same probe, run twice in two isolated processes, returns a byte-identical digest. That is the whole claim, and it is checkable in about a minute with no network call.
No model appears anywhere in the verdict path. Not as a tiebreak, not as a fallback, not as a re-ranker. This is a code fact rather than a positioning choice, and it is the difference between an instrument and a very confident opinion.
The guards shipped alongside, because a fix without a fence is a fix that gets undone: an endorsement cannot be lifted onto another receipt; an edited verdict breaks the seal; an insurer must name an oracle; a hardcoded blessed key fails the build; a competing oracle the insurer trusts settles exactly as well as ours does. Nine assertions on the loop, six on the public claims, seven holding the licence and the terms of service to the same boundary.
And the bug, which is the only reason to believe any of the above. This morning countersigning a verdict appended a field to the sealed receipt — and appending a field to a sealed receipt silently invalidates the signature over it, because the canonical body covers every key outside the signature envelope. The receipt still looked fine. It verified as false. Caught on the first real run, fixed structurally so the sealed artifact is never touched again, and a test now asserts byte-identity before and after countersigning. Build, break, catch, fence — that sequence is exactly what we are asking a risk market to be able to do to us, and the six questions above are where we are asking you to do it.
🥂💧🤝🎁🌱🌫️🔒 G → H 👑
H
Loading...
👑Significance — The Standard Gets Authored By Whoever Shows Up
The maître d', presenting:La Plume Partagée — one pen, laid across two place settings.
Inner monologue it should trigger:"If this becomes the format, I'd rather have shaped it than inherited it."
who writes the standard · why we cannot write it alone · the position that is actually on offer
Every settlement standard that ended up universal was authored by a small number of people who happened to be in the room early, and then inherited by everyone else as a constraint. The people in the room were rarely the cleverest available. They were the ones who answered when the format was still soft.
We cannot write this one alone, and the six questions above are not modesty about that — they are the parts where our answer would be a guess and yours would be knowledge. What is on offer is not a discount or a pilot. It is the pen, while the schema is still at version 1.
🥂💧🤝🎁🌱🌫️🔒👑 H → I 🐉
I
Loading...
🐉The Pivot — Precedence, Not Persuasion
The maître d', presenting:Le Tarmac — a section of road surface, served cold, without apology.
Inner monologue it should trigger:"The seatbelt framing was wrong. This is the road."
the correction to our own framing · what actually moved the speed limit · the constraint stated once
I described this work as protection for a long time and the word was doing me damage. A seatbelt is bought against fear of the crash, and everyone in the room prices it against their private estimate of the crash — which is a negotiation you lose slowly.
But cars do not travel at speed because belts were invented. They travel at speed because someone built a road where that speed became an ordinary thing to do, and then a policy could be written about it, and then ordinary people drove fast without deciding to be brave. As the book has it: "Nobody thanks the tarmac, and nobody drives at that speed without it."
The constraint, stated once and without decoration: the cost of establishing what happened sets the speed limit of the entire autonomous economy. Whoever makes that cost negligible does not win an argument about safety. They set the surface everyone else drives on.
🥂💧🤝🎁🌱🌫️🔒👑🐉 I → J 📚
J
Loading...
📚Digestif — The Evidence, as Ingredients Not Conclusions
The maître d', presenting:La Bibliothèque — the sources, uncut, with nothing built on top of them.
Inner monologue it should trigger:"They handed me the sources instead of the conclusion. I'll draw my own."
primary sources · the book · the sibling post · grading the win condition
The code, forkable now.github.com/wiber/thetacog-mcp — the countersign implementation, the licence including Part A-1, and ATTESTATION-ASK.md with all six questions in full.
The book chapters that argue it at length.The Dam Is Not Disinterest on why the capital is frozen; What Unlocks When the Signal Turns On on the stack that becomes possible — premium pricing, reinsurance tranches, securitisation, capital treatment — every layer of which already exists in other domains and is waiting for the measurement to become portable.
The sibling post.The Map the Size of the Empire — why grounding is a coordinate lookup rather than a moving van, which is the mechanism that makes the placement cheap enough to matter here.
Rice's theorem (1953) for the undecidability result the narrow claim is built to sidestep rather than defeat. Parametric insurance literature for the trigger design this borrows wholesale. T.J. Hooper (1932) for the doctrine on adopting an available safeguard, which is a live question rather than our conclusion to draw for you.
Now grade the win condition. Count how many of the ten predicted sentences actually fired in your head — you are the only person who can compute that number, which is the point. Then run the command, and if you have the standing to answer one of the six questions, answer it.
npx thetacog-mcp prove
Do you worry about $1.2B in AI liability?
If the property is trivial, software can check it — and why are you paying to check trivial properties? If it isn’t trivial, Rice’s theorem says nobody can. So we fixed the math.
press enter
Who did this make you think of? We’d love to know.