You're using a chatbot with the object permanence of a goldfish to govern an autonomous agent that holds your balance sheet. There is now a decidable, hardware-signed alternative. It's free. We do something narrower, and far stronger — and you never have to trust us to believe it. That's the entire point.
↑ a verdict being walked across the 144×144 lattice. it does this the same way every single time. your eval could never.
Show me the radio ↓Because for the first time in computing history, semantics run on the chip — north of 6 million times a second, on silicon, with no model in the loop. Not all semantics: the decidable kind — the part where meaning has an address (where-is-what), so resolving it halts at chip speed instead of regressing forever. The spec isn't a string we checksum — it's compiled into the same vocabulary the work speaks (named coordinates on a curated lexicon), so "did the work drift from what was asked" collapses to a distance between two positions on one lattice, not a verdict a model renders. The work product of one function lands at an exact coordinate on a 144×144 lattice, and whether it drifted from the coordinate the spec authorized is a recomputable, signed physical event — byte-for-byte identical on your machine and ours — not a second chatbot's opinion. "Does this output sit where the spec said it should?" used to be answerable only by another fallible model. Now it's measured in hardware. Semantic-on-chip wasn't a better tool. It did not exist — and the instant it does, every eval that flips between runs and signs nothing is, by law, below the standard of care.


"Wait — isn't 'does it mean what the spec asked' the exact non-trivial semantic property Rice's theorem says is undecidable? You can't wield Rice to bury the eval and then sell the thing Rice forbids."
↓ not a metaphor — every commit tolerance panel below is a real commit from this repo, each a decidable semantic placement that recomputes byte-for-byte. that's the six-million-a-second, made visible.
Yesterday, not knowing was an excuse.
You just read this. Welcome to the information hazard.
we're genuinely a little sorry. the measurement didn't make your stack uninsurable — it made that legible. that's the only thing that changed, and it doesn't change back.
It asks a second chatbot if the first one did okay.
Same input, three runs: PASS · FAIL · PASS. That's not a control framework. It's a magic 8-ball with a system prompt — and Brenda from Risk Advisory bills $400/hr to shake it for you. Stochastic governance, dressed up in a slide deck with your logo on it. And you're the one who signs it.
npx thetacog-mcp prove-rice sweeps a payload across the lane boundary: the LLM judge flips on 3 of 6 payloads (no two runs agree near the edge); the on-chip Oracle is 5/5 byte-identical at every step.scripts/pmu/prove-rice.mjs — run it, watch it happen.


Three real commit receipts from this repo. Same instrument, three readings — and every one is byte-identical if you recompute it. That's the opposite of a magic 8-ball you pay $400/hr to shake. Click the drift reading to open commit 60d4a3c6b's full receipt and recompute it on your own machine — no model in the loop.
Rice's theorem: software cannot decide a non-trivial property of software without infinite recursion. Who checks the checker? Another checker. Forever. It's turtles, and you're paying GPU rent on every turtle. You can't fine-tune your way out of a theorem — you change layers, to one that isn't software grading its own homework.
We don't beat Rice. Nobody beats Rice.
We do the one thing the theorem never forbade: move the question off software entirely — onto silicon, where a stranger recomputes the same answer to the same bytes. The theorem still stands, untouched — it's just standing on the incumbent's neck now, not ours. An eval is software judging software's behavior over every possible input (the undecidable thing). Ours is a physical layer underneath both, asking only "does this finished work sit where this spec authorized?" — a finite comparison of two fixed artifacts, not an infinite regress. Same theorem. Opposite side of it.
prove-rice.mjs.tests/pmu-simulator/competence-walk-is-real.test.mjs.Ask a model what a word means and it hands you more words. Ask what those mean — more words. To check a checker you need a checker; to ground a symbol you need another symbol already grounded. Software has no floor. That is Rice's staircase seen from below: not "who checks the checker" but "who defines the definer" — and the honest answer, in language, is no one, ever.
A flat tool — cosine, embedding, an LLM judge — pretends to end the regress and only defers it: it grounds your text in another pile of ungrounded reference text and calls the nearest one a match. The staircase is still there. It just moved into whatever dictionary the tool was trained on.
We end it by refusing to define with words at all. On the ShortLex lattice a concept's meaning is its address — where it sits is what it is. And the definer-of-definer chain becomes a thing you can watch run: stand on a lit coordinate, walk its row, jump to the column that row lights, walk that row, recurse — each step resolving a meaning by the meaning it is defined in terms of. The staircase is real; the difference is this one has a bottom. The lattice is finite and acyclic, so the walk halts — and it halts on silicon, north of six million times a second. The regress with no floor in language has a floor in physics, and that floor is the placement. docs/architecture/ballistic-walk-the-definer-of-definer.md
Each independent walk asks one small question — did the real work land where the intent did, more than chance would explain? — and answers with a number: σ, the separation of true alignment from noise. One walk is a whisper. But the walks are independent, so their evidence adds, and the sum is a divergent series — it climbs with no architectural ceiling (the patent's own phrase). Add walks, sharpen the lanes, and the certainty does not approach a limit; it keeps going — past 10σ, past 100, toward the 600σ a physicist would call decided.
Now the honest half that makes it real instead of magic: a divergent series buys unbounded precision, never unbounded coverage. We become almost violently certain about the lanes we have carved — and we know nothing about the territory we have not. Perfect meaning-is-position (ρ→1) stays out of reach; we close the gap inside the covered lanes, never over all of meaning. Infinite sharpness on a finite map. That is not the hedge — it is the reason the map is worth insuring: a claim with an edge you can see is the only kind an underwriter can price. scripts/pmu/walk-significance.mjs
Two tugboats lost their barges because they carried no radio to hear the storm warning. The whole industry carried no radios — so they pleaded custom. Judge Learned Hand was unmoved: "a whole calling may have unduly lagged in the adoption of new and available devices… there are precautions so imperative that even their universal disregard will not excuse their omission."
An LLM eval is the tugboat with no radio. The radio is now on the shelf. And you've seen the shelf.
gzip bridge → 144×144 lattice → ballistic PMU walk → ed25519 receipt

INTENT (what you commanded) · REALITY (what the agent did) · Δ (exactly where they disagree). this exact triptych rides on every commit we make. it's not a render — it's the receipt.
A spec is published. Work is submitted, signed. The gate runs on the chip and signs a verdict binding what was asked, what was delivered, and who delivered it. Either side recomputes it. No blockchain. No clearinghouse. No second model asked to bless the first.
.thetacog/pmu/target/release/pmu-onchip — the same daemon digest is stamped into every receipt.npx thetacog-mcp attest-demo — Node A signs the spec, Node B signs the work, the chip gates it, an underwriter prices it, and an LLM judge (any onboard CLI on your PATH — Claude Code, Cursor, codex, ollama, …) renders a verdict that signs nothing a stranger can recompute. Measured 2026-06-19: a small model (llama3.2:1b) flips on the same spec; a large one (current Claude) holds — the flip is a class of error that rides on model capability, and the chip removes the whole class. (Or drive each pillar by hand: attest publish-reef / submit / gate / verify · scripts/pmu/attest-demo.mjs.)A bullseye is one shot you hit or miss. This is the other thing: every one of the 144×144 cells is a competence coordinate, lit by where a commit's reality actually fired against where its intent said it should. You don't score a point — you get a picture of where you are as what you meant. Green fired in-lane · amber a few zones out · red too far to absorb. The magenta crosshair is your one located pixel.


both rode on real commits in this repo — the filename is the sha. left: the work fired where the intent declared it. right: the red gathering in the corner is reality firing where the intent was weak — drift you can see before it ships, not in a post-mortem after.
This is what arrives on every commit we make — not a dashboard someone curated, an artifact the chip emitted. And when a commit is tight and nothing fires out of lane, the field is silent — near-black, just the crosshair. An instrument that stays quiet when nothing happened is the only one you trust when it finally lights up.
The strength isn't one heroic claim — it's ten independent ones. Tap any to see exactly how the Rust runner proves it. We'll wait.
prove-rice.mjs STRONGchip-cloud-lattice-golden.test.mjs:64 STRONGforge-test.test.mjs:51 STRONG*npx thetacog-mcp attest verify STRONGprove-rice.mjs STRONGdata/pmu/reef/reef-144.json STRONG.thetacog/sigma-panel-trend.ndjson MEASURED.thetacog/reef-interventions.ndjson MEASUREDattest publish-reef STRONGscripts/pmu/attest.mjs STRONG* T1 holds under host-key pinning; the hardware-attested key (Secure Enclave / tape-out) is what makes pinning unforgeable. We state the bound out loud.
"We deployed a secondary AI to monitor the primary AI.
Then we asked a third one to summarize their argument."
Your agent just made a $10M unrecoverable error. Plaintiff's counsel is holding a physical memory receipt; you are holding a probabilistic text output — and the only other thing on the defense table is a goldfish bowl. There is no second sentence.
Every competitor's deck hides its bounds. We lead with ours — because the fence is what makes the signal underwritable.
Everything above is the instrument. This is the ground it stands on — the one thing we want you to walk away believing. For seventy years the field has tried to get a program to certify that another program is good, safe, aligned — and kept failing, because it assumed the failure was a tooling gap next year's bigger model would close. It is not a tooling gap. It is a theorem. We took a question philosophy proved you cannot answer, and instead of answering it, we turned it into one a machine can: semantic meaning became a spatial coordinate — and a coordinate has a price.
Rice (1953), downstream of Turing and Church (1936), drew a wall around the whole effort: every non-trivial semantic property of a program is undecidable. The part the field skipped is why. Undecidable does not mean hard — it means there is no computable line a machine could draw between the cases that satisfy the property and the cases that don't. Ask ten careful people where "helpful" ends and "obsequious" begins; you get ten lines, because there is no fact of the matter to decide against. At the speed of an autonomous agent, a meaning you cannot compute is the same as a meaning that isn't there. You cannot grade your way over that wall — and the deck above shows the incumbent walking into it at $400/hr.
We did not solve Rice's theorem. Nobody solves Rice's theorem. We changed the question. Computable measure theory gives the hinge: a property is decidable exactly when its characteristic function is computable — when an actual procedure returns yes or no for every case — and a measurement is just the evaluation of such a function. So the whole game becomes: find the property of an agent's work that is both load-bearing and decidable, and price that, instead of forever litigating a word that was never going to resolve.
There is such a property, and it was hiding in plain sight: did the agent stay in the lane it was hired for? Not "is the output good" (undecidable, the interior) but "where did the work land relative to where we hired it to land" (decidable, a geometry placed around the interior). The whole field looked inside for understanding. We put a rigid lattice outside and read off a coordinate. Meaning stops being a property you judge and becomes a position you measure — the same number on your machine as on ours. We carved a decidable slice out of an undecidable whole, and we price the perimeter, never the interior. Call it the priced undecidable perimeter: a rigid boundary drawn around a thing no one can read, where the boundary itself is the readable, checkable, sellable object.
If your first instinct is that's a reframing, not a solution — you are right, and that is the entire point. Solving is forbidden by the theorem; reframing onto a decidable property is the only move that was ever on the table. The question is whether the reframing is real or merely rhetoric, and the test is brutally concrete: does it emit a number a stranger recomputes byte-for-byte, with no model in the loop? It does. Section 3 above is a real commit from this repo landing at an exact coordinate — 35% outside its authorized lane against a 25% tolerance — and you can re-run the gate on your own hardware and get the identical verdict and the identical σ. A reframing you can recompute is a measurement; a reframing you have to trust is a slogan. This is the first kind — that is the only reason it is worth your time, and the only reason it is worth an underwriter's.
Physics has a smallest distance. We have a smallest lie.
— Tesseract Physics · The Razor's Edge — below it, placement is noise; above it, placement is a fact.
Three people already made this exact move, one field over, and proved it works. We are adjacent to them, not claiming them — the honesty is the load-bearing part.
Black, Scholes & Merton (1973) did not crack open the company to decide whether its stock was "really" worth more. They left the stock sealed and priced the volatility around it — and turned the unpriceable into a traded instrument worth trillions. Our move is mechanically identical, just one field over: we do not open the model to decide whether its reasoning is sound. We leave the model sealed and price the drift around it. Seal the asset, price the perimeter. Black-Scholes didn't touch the stock; we don't touch the weights.
Harnad (1990) proved why the coordinate isn't a crude shortcut. His symbol grounding problem: a symbol cannot be grounded by more symbols — it grounds only when placed in a structure outside the symbol system. So a lattice coordinate is not a thinner substitute for meaning; placement is what meaning is made of, and the coordinate is the only version of it a machine can hold without begging the question with yet more ungrounded symbols. Every LLM-judges-LLM eval fails precisely here — it grounds your text in another pile of ungrounded text.
Knight (1921) is where it pays. He drew the line between risk — measurable, insurable — and uncertainty — unmeasurable, excluded from every policy ever written. AI competence has lived entirely on the uncertainty side; that is why no one will write the coverage your board needs. The decidable slice drags one load-bearing question across that line into risk: a countable event, a frequency, a premium. That is not a metaphor about insurance — it is the literal precondition for an insurance contract to exist.
An idea is worth what it is worth at the point of highest pain — so here is the concrete one. Your board wants to hand an autonomous agent root on production: rotate this cert, scale that fleet, patch this node, drain that pool. This is the Agentic-SRE wave, and you are the person who has to sign off on it. It is the beachhead because the blast radius is largest here — which means the decidable boundary is worth the most here, and the cost of not having it is the most literal.
A bad doctor costs thousands. A competent plumber in the wrong lane costs millions.
A bad doctor makes a wrong call — one patient, one room, a bounded harm a malpractice system already prices. A plumber who is competent but working in the wrong lane — who opens the gas main when he was hired for the sink — costs the whole building, not the one fixture. The catastrophe is not incompetence. It is competence aimed off-target. An action executed perfectly, in the wrong lane, is invisible to every quality check you own — a quality eval sees a clean, well-formed action and passes it, while the agent has just modified the IAM policy it was never scoped to touch. Off-target-done-well is the failure a score literally cannot see; a negation is just a minus sign.
Now put the plumber at machine speed with credentials to the whole fleet — that is an autonomous SRE agent. The reason your board wants it deployed and you are the bottleneck is that one agent touching the wrong system is an extinction-level compliance event, and you cannot get comfortable on an undecidable promise that it will behave. So stop trying. Hold it to the decidable thing instead: a cryptographically sealed, recomputable receipt that it stayed in the lane you assigned. The lane constraint is not a cage on the agent — the constraint is the trust signal. A narrow, signed, recomputable boundary is exactly what lets you say go. You stop being the bottleneck and become the enabler — and the same receipt, read by an underwriter, is the line they can finally write, because a decidable breach is a countable event and a countable event is a premium. Decidability Is Meaning · Competence Is a Shape, Not a Score.
So — are you out of your pixel? For an agent with root on your infrastructure, that is not a turn of phrase. It is the only safety question that has an answer.
Ancestors cited, not claimed: Rice (1953) · Turing & Church (1936) · Harnad (1990) · Black–Scholes–Merton (1973) · Knight (1921) · computable measure theory. The floor is filed: Patent US 19/637,714. Quotes deep-linked to Tesseract Physics — Fire Together, Ground Together and the dispatches above.
Read the section above as the physics and this one as the consequence. You are the officer who signs off on handing an autonomous agent something that matters — root on production, a payments rail, the balance sheet. The law has always said you owe a duty to oversee a risk like that. What it could never give you was a control that discharges the duty, because the question the duty turns on — is the agent's work sound? — is the undecidable interior from the section above. You cannot install a monitoring system for a property no machine can compute. So for ninety years the duty was real and unmeetable, and everyone quietly agreed not to mention it. That just ended — and the way it ended lands on you, by name, not on the company.
Since In re Caremark (Del. Ch. 1996), a board has owed a duty to make a good-faith effort to install and monitor a system that surfaces mission-critical risk. For two decades it was the hardest claim in corporate law to win. Then Delaware gave it teeth: Marchand v. Barnhill (2019) revived it and set the duty at its maximum for a "mission-critical" risk; In re Boeing (2021) applied it to the 737 MAX and settled for $237.5 million, the largest duty-of-oversight settlement in Delaware history; and In re McDonald's (2023) extended the duty, for the first time, from directors to officers — the people who actually deploy the AI. The reason it lands on you and not the corporate veil: oversight liability sounds in bad faith, a breach of loyalty — precisely the category that charter exculpation (DGCL §102(b)(7)) is built not to cover, that indemnification (§145) needs good faith to reach, and that D&O conduct exclusions are written to carve out. It is one of the few exposures your usual backstops are designed not to absorb.
The question a Caremark court asks is not was the decision wrong. It is show me the system that was watching.
— Tesseract Physics · The Budget Is the Proof — for a deployed agent, the honest industry answer today is that no such system exists.
Your last line of defense is the oldest one — everyone in our industry deploys agents this way; there is no accepted way to monitor them. It has been dead since 1932. In The T.J. Hooper, 60 F.2d 737 (2d Cir. 1932), two tugs lost their barges in a storm they would have dodged with a weather radio almost no tug then carried. It was not industry custom to carry one. Judge Learned Hand held the owners negligent anyway: "a whole calling may have unduly lagged in the adoption of new and available devices… there are precautions so imperative that even their universal disregard will not excuse their omission." Custom is no defense. Hand's test is precise — liability attaches when a precaution is available, reasonable, and proportionate to the risk, not for skipping any new gadget.
Now apply that test to the agent you are about to deploy. The section above made the in-lane question decidable: not the undecidable "is the work good," but the recomputable "did it stay in the lane you authorized" — a coordinate, signed with an ed25519 key so it cannot be forged or back-dated, that the chip recomputes six million times a second, byte-for-byte identical on your hardware and a stranger's. That is a precaution that is now available at reasonable cost. The instant it exists, Hooper inverts: "no one else in our industry monitors their agents either" stops being your shield and becomes your admission. The two doctrines fuse — Hooper supplies the negligence standard, Caremark supplies the fiduciary one, and both now read the same way. Deploying an agent on mission-critical work without a recomputable in-lane monitor is the modern tug with no radio.
This is not a future risk to plan for. ISO/Verisk filed standard endorsements CG 40 47 (broad generative-AI exclusion, Coverage A and B) and CG 40 48 (Coverage B), attaching at U.S. commercial general-liability renewals from January 1, 2026; W. R. Berkley has filed an "absolute" AI exclusion reaching D&O, E&O, and Fiduciary. They trigger on loss "arising out of" AI — among the broadest causal phrases in policy drafting. Translation: the "silent AI" your tower used to cover by default is being converted into an explicit exclusion. The personal exposure in move 01 is arriving on the same renewal cycle as the deletion of the coverage that used to absorb it. The endorsements are already filed.
The control that discharges this duty has to do exactly one thing a court can rely on: produce, for every consequential action an agent takes, a signed record an independent party can recompute and check byte-for-byte — establishing whether the action stayed inside the lane its operator authorized. Not a confidence score an AI grader can flip from "pass" to "fail" on the same input near a boundary; a court cannot lean on a number that is not stable. A fixed result that re-verifies identically every time, signed so it cannot be forged, checkable by anyone offline without trusting us. It is the audit trail that lets a director say, truthfully and with proof in hand, "we would have caught it."
What this page does not claim.
We do not claim to prove an agent's work is correct — that is the undecidable interior, and anyone selling it is selling a solved halting problem. We claim the narrower, decidable thing: whether the action stayed in its authorized lane, signed and recomputable. The cryptographic receipt layer is proven and verifies independently today; the physical-drift signal beneath it is stated, by our own ledger, as a partial result on a single architecture — not yet the bankable claim. The legal authorities below are cited, not paraphrased, so you can check every one. Honesty about the edge is the load-bearing part: it is what separates a control a court can use from a vibe with a logo.
You have now seen the radio on the shelf — a recomputable in-lane receipt you can generate for yourself, in one command, on your own hardware, trusting no one, including us. The instant you have seen it, "no one told me I could" stops being available — to you, or to the court that later asks why it stayed off.
The doctrine is developed in full in Tesseract Physics — Fire Together, Ground Together, The Budget Is the Proof. The instrument that discharges the duty is filed: Patent US 19/637,714. "The Hooper Standard" names the 1932 legal doctrine — not a person. The physics is why we are here, not the patent.
Way 1 — run it yourself. Install nothing; it runs straight from npm.
Way 2 — let the AI you already trust be the witness.
Your agent reads the install output into its own context. It tries to confirm the placement by reasoning — and catches itself, because Rice applies to it too (it's a chatbot; it can't grade this any more than the eval-chatbots could). So it does the only honest thing left: it recomputes the receipt (prove-rice --check → exit 0), then turns to you and explains, in its own words, that you're standing in a T.J. Hooper moment — the radio is on the shelf, and you've now seen it. The defense's own expert just took the stand for the plaintiff. The goldfish filed an amicus brief against the goldfish bowl. We didn't argue it into agreeing — the math did, on your machine, in the tool you already use. The skeptic you brought becomes the witness.
Run it on your own outputs. If your AI governance can't survive a stranger — or your own agent — re-running the math on their laptop, it was never governance. It was a vibe with a logo.
Get the package → Read it again ↑so… are you out of your pixel? (it's okay. yesterday, everyone was.)