Tolerance panels · the instrument that judged every edit to this post
Green in-lane · amber a little out · red drift. Every panel is a real commit, byte-identical on recompute. Tap any panel to open its shareable receipt.
Geometric Driven Development — 4 measured edits to this post. Recompute any of them yourself, in a clone of this repo: npx thetacog-mcp publish-commit --commit 84a2ac34a
There is a clip going around — Christopher Carter, CEO of Approyo, a man who runs enterprise SAP infrastructure for a living — where he compresses the entire AI-governance conversation into six words borrowed from Peter Drucker: you can't govern what you can't count. He's pointing at shadow AI: the average enterprise knowingly runs dozens of agents, unknowingly runs dozens more, and a security chief guessing at the footprint doesn't have a strategy, he has a hope. He's right — and the sentence has a trapdoor he doesn't open. Run the discovery audit he's calling for, and run it perfectly: your register firms up from "dozens, roughly" to an exact integer, the slide improves, the room applauds — and your governance has not moved an inch, because a census of agents says nothing about what any one of them did since lunch. The unit of risk was never the agent. It is the action — and an action becomes countable the moment it lands somewhere: inside the lane it was chartered for, or outside it. Sum those landings and you get say-vs-do drift, a number per agent. That number is the physics accountability stands on. Everything below is the same claim from ten angles.
One habit of the house, printed up front: every section opens with the exact sentence it is built to make you think — a prediction you get to grade, published before you read, because manipulation needs the dark and this is on the menu. The win condition, declared now: this piece wins if you leave and recompute — count your own register, run the command at the bottom — and fails if you leave merely nodding.
A
Loading...
🧮Amuse-Bouche — Why We Believe the Count You Have Is Not a Count
The maître d', presenting:The Uncounted Cover — a dining room set for twelve; the till says forty ate. The house that billed from memory served its last winter on a spreadsheet nobody dared open.
Inner monologue it should trigger:"My agent inventory is a guess, and everything downstream of a guess is also a guess."
census vs. count · the noun error · behavior as the unit · the rung under the governance word
THE COUNT YOU HAVE THE COUNT GOVERNANCE NEEDS
────────────────── ──────────────────────────
agents on the register agents on the register
plus the shadow ones ↓
(an estimate, quietly) every action: WHERE did it land?
↓
in-lane / out-of-lane / unplaced
↓
signed receipt, one per action
↓
say-vs-do drift — a NUMBER, per agent
↓
countable → priceable → insurable → accountable
Do you worry about $1.2B in AI liability?
If the property is trivial, software can check it — and why are you paying to check trivial properties? If it isn’t trivial, Rice’s theorem says nobody can. So we fixed the math.
press enter
Who did this make you think of? We’d love to know.
The left column is the census — where the whole market currently lives, where the discovery audit runs, and where the applause happens. It has an error bar nobody says out loud, and even at zero error it describes a fleet, not a yesterday. The right column is what a count of actions looks like, walked top to bottom: each output gets a coordinate — where it landed against the agent's stated lane; each coordinate gets a signed receipt; and the receipts sum into say-vs-do drift — the standing distance between what the agent is chartered to do and what it does. That number is the physics. Every governance word you want to say later — audit, gate, cap, insure, fire — stands on it or floats.
Drucker's line has a silent second clause. You can't govern what you can't count — and you can't count what has no unit. "Agent" is a census unit. "Action, placed against its lane" is a count unit. The whole argument is which noun you pick.
🧮 A → B 🔦
THE LADDER — SIX RUNGS, EACH ONE REJECTABLE
The whole argument, as internal monologues you must climb. Reject any one of them and the ladder breaks — and you'll know exactly which rung to tell us about. Accept all six and you were co-opted by the argument, not by us.
"My register undercounts, and I've known it longer than I've said it." — your discovery audit is real and worth running; this rung only asks you to admit its error bar.
"An inventory is a census. A census has never governed anything." — your asset register is correct for what it decides: budget, licensing, attack surface. Behavior is the one thing it doesn't reach.
"An action, placed against a stated lane, is a countable event." — in-lane, out-of-lane, or unplaced: three buckets, a denominator, arithmetic.
"A count the agent produces about itself is testimony, not measurement." — your eval suite is real and catches real regressions; it is still the model grading the model.
"A count that re-runs identically on my machine is the only count an auditor accepts." — reproducibility isn't a feature, it's what "count" means.
"Priceable follows countable, and accountable follows priceable." — the ladder's top rung is the governance word everyone was trying to say first.
You can't settle rungs 3 through 5 by thinking about them — they're claims about a running system. The system is a one-minute local install: npx thetacog-mcp attest-demo runs on your own machine, reads nothing of yours, and places one sample action against a lane so you can watch what "counted" means before anyone asks you to believe it. The rest of the meal digests the rungs one at a time.
B
Loading...
🔦The Why — The Second Count
The maître d', presenting:Consommé of the Second Count — clarified until you can read your own register through it. The kitchen that counted covers while the stove burned took the insurance adjuster's word for what the fire cost.
Inner monologue it should trigger:"I've been counting the wrong noun — agents, when the unit of risk is actions."
the belief before the mechanism · why inventory feels like progress · the drift that accrues while you audit
Here's why the census feels like governance while not being it: discovery is visible work. You run the audit, the number firms up, the slide improves — twelve knowns become seventy-one knowns and the room applauds. But watch what the number does overnight: nothing. It sits there while every one of the seventy-one takes hundreds of actions, each one either inside its charter or outside it, and none of those events land anywhere. The register improved; the ledger still doesn't exist. This is the belief under everything that follows: an agent is not a risk, an action is a risk, and a security program that counts the former while the latter accrues is measuring its own effort, not its exposure. The drift between what an agent was told and what it does — the say-vs-do gap — is accruing whether or not anyone looks, and an accruing quantity has never once paused for an audit cycle. You don't need the mechanism yet. You need to feel the difference between a number that describes your fleet and a number that describes yesterday.
🧮🔦 B → C 🤝
C
Loading...
🤝Connection: The Question With Your Name Pre-Filled
The maître d', presenting:Carpaccio of the Shadow Agent — thin slices of a tool your own building shipped, plated on a mirror. The garnish was added by the marketing department, who did not inform the kitchen.
Inner monologue it should trigger:"Half of these are running in my building right now, and I could not name them all under oath."
your register, not a vendor's demo · the meeting already on the calendar · the error bar nobody speaks
This is your table, not ours. Somewhere on your calendar — this quarter, next at the latest — is the meeting where someone senior asks "how many AI agents do we run, and what are they authorized to do?" and the honest answer has an asterisk you will not say aloud. Carter's shadow-AI point is your lived Tuesday: the sanctioned copilots you counted, the browser extension legal installed, the workflow bot a product team wired to production data because the API key was right there. The seat that has to answer is yours — not the vendor's, not the intern's — and the question is already scheduled whether or not this post exists. What connection means here: we are not describing a hypothetical enterprise, we are describing your register's error bar, and you can check that claim against your own asset list tonight without touching anything of ours. If your count is exact and your charters are current, this meal is dessert. If it isn't — and the room Carter is talking to says it isn't — then the rest of these plates are about your Tuesday.
🧮🔦🤝 C → D 🎁
D
Loading...
🎁Contribution: The Number You Get to Bring Upstairs
The maître d', presenting:Terrine of Say-versus-Do — pressed in visible layers: the charter on top, the behavior beneath, the gap you can see from across the table. Built to survive the elevator to the board floor, where the adjective-laden slide never once arrived intact.
Inner monologue it should trigger:"I could walk into the next audit holding a number per agent instead of an adjective per slide."
what you hand upward · a number that survives forwarding · the drift column in the deck
What the count gives you is the rarest artifact in an AI meeting: something to hand to the people above you that doesn't need you standing next to it. Today your board deck says "monitored," "governed," "low risk" — adjectives, each one a small loan against your credibility, each one defensible only as long as you're in the room. A say-vs-do number per agent is a different object: this agent, this month, this many actions, this fraction out-of-lane. Forwarding it stakes nothing, because the receiver doesn't have to trust you — the receipts recompute. The ingredients on this plate: the drift column your register is missing; the denominator — actions placed, not incidents remembered — that turns "how safe are we" into arithmetic; and the forwarding asymmetry: an adjective defended is credibility spent, a count forwarded is credibility banked. Your auditor, your CFO, your underwriter all already know how to carry a number. None of them knows how to carry "we take governance very seriously."
🧮🔦🤝🎁 D → E 🌱
E
Loading...
🌱Growth: The Count Widens the Lane
The maître d', presenting:Sorbet of the Widening Lane — served to reopen the palate. The queue outside the approval committee has its own microclimate; several promising dishes froze solid waiting in it.
Inner monologue it should trigger:"If drift is a number I watch, I can say yes faster than I say no today."
governance as throughput · the gate that scales · saying yes with instruments
The standard objection to governance is that it's a brake, and for once the objection is right — assumption-based governance can only brake. When you can't see behavior, the only safe answer to "can this agent also handle X?" is a committee, and the committee's only safe answer is slow. Watch what the count does to that: an agent with ninety days of receipts and a flat drift line is an agent you can widen — more autonomy, more surface, a bigger lane — and the instrument carries the justification, not your neck. The pieces here: the flat line as a permission slip — history replacing hypotheticals in the approval meeting; graduated autonomy — lanes that widen on evidence and narrow on drift, automatically, in both directions; and the reversal you get for free — when something does go out-of-lane, you know the same day, not at the quarterly review, which is precisely what makes the wider lane defensible. Teams that can't count behavior are stuck choosing between velocity and control. The count is how you stop choosing.
🧮🔦🤝🎁🌱 E → F 🌫️
F
Loading...
🌫️Uncertainty: The Count the Model Can't Give You
The maître d', presenting:Soufflé of Self-Certification — risen entirely on the testimony of its own steam. Magnificent at the pass; the house stopped serving it after one too many collapsed in the elevator, mid-ride, on the way to the people who sign things.
Inner monologue it should trigger:"I can't ask the model to count itself — the count has to come from outside the thing being counted."
testimony vs. measurement · the 1951 boundary · what we refuse to claim
Here is where the trap lives, and it's worth naming precisely because the vendors circling this space mostly don't. The obvious way to count agent behavior is to ask the agent — self-evals, model-graded rubrics, confidence scores. All of it is testimony: the system under examination producing the examination. And the boundary here isn't a vendor opinion, it's Rice's theorem, 1951 — every non-trivial claim about what a program will do is undecidable from inside computation, which is why software auditing software inherits the exact uncertainty it was hired to remove. The boundary is not ours and predates every vendor in this market: the theorem was proved in 1951, seventy-five years before anyone filed a patent near this space — including us, in 2026. We found the boundary; we didn't draw it, and you can check the dates without trusting either claim. And note what the boundary leaves standing, because this is what we refuse to claim: WHERE an action landed against a stated lane is decidable, recomputable, a count. WHETHER the agent is correct, safe, aligned — undecidable, by anyone, at any price. A vendor claiming the second number is selling you the soufflé.
🧮🔦🤝🎁🌱🌫️ F → G 📏
G
Loading...
📏Certainty: The Count That Re-Runs
The maître d', presenting:Canelé Counted Twice — baked, then baked again from the same recipe in front of you: identical to the crumb. The demo that only demos once is a magic trick, and the house does not employ magicians.
Inner monologue it should trigger:"Same input, same coordinate, on my own machine — that part I can check without trusting anyone."
reproducibility as the definition · your silicon, not our cloud · the verdict with no model inside
Everything the last plate took away, this one gives back — on the narrow ground where it can actually be defended. The placement verdict — this action, this lane, in or out — is computed with no model in the loop: a deterministic walk over a fixed 144-coordinate lattice, the same input producing the same coordinate every run, which is what separates a count from an estimate wearing a number. And it runs locally, via npx, on your own silicon — your prompts and outputs never leave the building, which your security team will verify in the source rather than in our marketing, as they should. The ingredients: determinism as the definition of countable — if two runs disagree, you never had a count; the audit property — an auditor doesn't accept your number, they re-derive it, which is only possible because there's nothing stochastic to re-derive around; and the honest scope — this certainty covers placement, full stop. It's a smaller certainty than the market is being promised elsewhere. It has the advantage of being checkable, twice.
🧮🔦🤝🎁🌱🌫️📏 G → H 👑
H
Loading...
👑Significance: One Column Away
The maître d', presenting:The Ledger Brought to Your Seat — your own asset register, returned to you with one ruled column added in pencil. The kitchen claims no credit for the book you were already keeping.
Inner monologue it should trigger:"The register I already keep is one column away from being a governance instrument."
what you already own · the column, not the platform · who signs the first behavioral report
Strip our product off this page entirely and the move survives, which is how you know it's yours: take the agent inventory you already maintain and add one column — "behavior last counted." Not "last reviewed," not "risk rating" — the date someone last had a count of that agent's actions against its charter, from any source: your logs, your gateway, your evals with their testimony-limits noted, anything. For most registers the column starts entirely blank, and the blank column is the report — it converts "we should improve AI governance" from a sentiment into a checklist with a denominator, and it costs fifteen minutes. What follows from a filled column is the part worth wanting: the first security lead in your industry whose AI section is a table of counts instead of a paragraph of adjectives doesn't just pass the audit — they get quoted in it, because auditors reuse whatever finally gave them a denominator. That artifact is yours to author with or without us. The pencil column is the whole ask.
🧮🔦🤝🎁🌱🌫️📏👑 H → I 🐉
I
Loading...
🐉The Pivot: Count-First Becomes the Bar
The maître d', presenting:Dragon's Breath, Counted — flambéed at the table with the match the house has been holding, visibly, since the first plate. The flame is measured before it is admired.
Inner monologue it should trigger:"Count-first is going to be the bar my board deck is judged against, whether or not I like it."
precedence, not persuasion · the tugboat and the radio · the standard others must adopt
Now the register shifts, because it can afford to. In 1932 a court held two tugboat operators negligent for sailing without weather radios — not because radios were required, but because radios had become cheap and available, and The T.J. Hooper established that a whole industry's custom stops being a defense the day the instrument exists. That is the mechanism bearing down on the census-only security program: the moment behavioral counting is a fifteen-minute local install, "we counted our agents annually" stops being diligence and starts being the custom that didn't survive contact with the instrument. The regulatory pincer is already set — SOX makes officers certify controls they cannot currently see into, and the EU AI Act's staged obligations keep raising what "adequate oversight" means — and neither framework accepts the system's own testimony about itself, for the reason the soufflé collapsed three plates ago. So here is the judo of it: this standard doesn't need us to win. Any enterprise that starts publishing say-vs-do counts sets the bar for its peers, because the first count in a market redefines what guessing looks like. You can be the one who sets it or the one measured against it. The dragon was always on the menu; we just served the count first.
🧮🔦🤝🎁🌱🌫️📏👑🐉 I → J 🧾
J
Loading...
🧾Digestif: L'Addition, Itemized
The maître d', presenting:L'Addition, Itemized — the bill arrives as a count, never a vibe. The house that presented totals without line items is remembered fondly, by its creditors.
Inner monologue it should trigger:"Every number in this bill is one I can recompute."
evidence last, as ingredients · the misattributed quote · sources on the table · the to-do that closes the loop
Ingredients, not conclusions — check each against the source. The clip itself:Carter's shadow-AI segment, six words at the top, the guess-vs-strategy point fourteen seconds in. A finding about the quote: the Drucker Institute has itself noted that the famous measurement line is likely misattributed — Drucker probably never said it. Sit with the irony: the most-quoted sentence about counting fails a provenance check, which is precisely the argument for receipts over quotations — a signed count doesn't care who said it. The regulatory ingredients: SOX officer certification and the EU AI Act's phased obligations, both of which refuse self-testimony — the insured cannot certify its own innocence. The house's own ledger: the market logic of a countable incident is in Incidents Countable: The Line That Opens a Market, the instrument-precedent in Telematics for Semantics, and the reason actuaries have been waiting for a denominator is in the book at The Actuarial Blindspot, with the deeper claim — that an unaddressed action is unaccountable in principle, not just in practice — at Erasure Takes the Address.
The to-do repeats the opening move, because the loop closes or it wasn't a loop: run npx thetacog-mcp attest-demo on your own machine and watch one action get placed — in-lane, out, or unplaced, with a receipt. Then the win condition, as declared before the first plate: Count how many of the ten predicted sentences fired in your head. Ten predictions were published in advance; you are the only one holding the tally — and the tally matters for the same reason the whole post does. We just spent ten courses arguing that a system should be graded on counted behavior, not on its own testimony about itself. This post is a system. Your count of fired predictions is its say-vs-do number — the one measurement of this argument that its authors are structurally unable to produce. Send it, and you've run the audit we've been describing, on us.
Do you worry about $1.2B in AI liability?
If the property is trivial, software can check it — and why are you paying to check trivial properties? If it isn’t trivial, Rice’s theorem says nobody can. So we fixed the math.
press enter
Who did this make you think of? We’d love to know.