a working page · fresh off our call · updated as we learn

What’s next — and what we’re agreeing on

This is a live page, not an email. The top is the short action list — three things, kept deliberately tight. Below it, new this week, is the deeper track: everything you’d want in your pocket the moment this comes up inside your building — the exact lines, the frame the industry already understands, and the numbers if anyone technical leans in. Scroll as deep as the moment needs; none of it is homework. The spine of all of it, in one sentence: without a hardware receipt, an agent’s liability is unbounded by construction — no record exists that could cap it. The receipt is what makes a loss finite: in-lane, priced, arguable. And the shape it takes is one your industry has run for three hundred years: the spec is the policy wording. Insurance has always been named perils — specified behaviour, deviation detectable, covered; unspecified behaviour, excluded, not by a lawyer’s trick but because we never claimed to measure it. Coverage resolution equals spec resolution, which is why the honest first peril is the one you named at lunch.

last updated 2026-08-06
▶ Listen on the walk — the whole page, ~4 min
Today
Your sentence is now the product’s boundary — and the one thing for today is to run it yourself, once.

At lunch you put it in your own words before we ever did: undecidability, plus detecting if a plumber starts doing brain surgery, is what makes for insurable AI. That is now the fence, in writing. Scope escape is the named first peril — the doctor doing plumbing, detectable today, uninsured today. Everything we cannot see is excluded by construction, and we say so on the pricing page rather than waiting to be asked. You found the fence six weeks before it was drawn.

Before we talk, spend ten minutes making it yours. Open Claude Code in a brand-new empty folder and run the line below. That’s the whole ask — nothing to prepare, nothing to send back.

Why this and not a doc: the moment you’ve run it on your own machine, you stop being someone forwarding a thing you were told about and become the one person in your building who has actually seen it work. That’s the position no memo can hand you, and it’s yours the second the receipt prints.

npx -y thetacog-mcp@latest attest-demo
📱 On a phone, no laptop? You can’t run the command yet — but you can watch the exact same proof in your browser: open thetadriven.com/d and thetadriven.com/pmu-simulator/demo right there. The difference matters and it’s the whole point: in the browser you watch it; on a computer you run it yourself and recompute it byte-for-byte. When you’re back at any machine, the line above is the real thing.
Then ask it: “Walk me through whether this is true, and ask the hardest skeptical questions you can.” Bring me the one question it couldn’t shake — that’s the whole agenda for our call.
The “in-house” trap — where this stands, and the opening it gives you

What we’re looking for hasn’t changed: the person inside a firm who hears “we’ve got AI risk handled internally” and knows why that’s the opening line, not the closing one. That person is you.

What the field taught us
The most capable firms don’t reject this — they say “we’re building it ourselves.” With proprietary data and client distribution, they treat AI liability as an ordinary software problem their own engineers can own. Read it correctly: it’s a concession the capability is necessary.
And it walks into the same wall every time. The math: software can’t certify the behavior of other software from inside the same failure domain. The regulation: separation of duties under Sarbanes-Oxley and the EU AI Act means a firm can’t self-certify its own AI on a monitor built from the same stack.
So the attestation that actually caps liability and satisfies a court has to come from below the software layer — a neutral primitive the firm doesn’t control. Built entirely in-house, that spend produces an uninsurable liability. The one piece an internal team can’t hand itself is the hardware-grounded proof — and that’s exactly the piece already here when they reach the wall.
New this week, and it changes how you read a “no” from the top: the senior leader who declines is not the obstacle — he is the alarm system working correctly. A practice leader who defers to his trusted domain experts is doing the executive job as it is actually designed: sense the risk, delegate the work, refuse to personally absorb what the seat was never built to absorb. If reopening the question requires him to overrule his own team, you are asking him for courage — and a well-run seat never supplies courage, by design. The play is never a better argument. It is cover: one demand he can make that needs no bravery and no math — a placement receipt as a condition of sign-off — so his experts keep building, his alarm gets a referent, and the question reopens itself as his own diligence rather than anyone’s pitch.

The pocket playbook — three patterns, run on sight

You don’t have to hold the theory in your head under pressure. Each of these is a pattern you’ll hear in a room. When you hear the trigger, you already know what’s really happening — and you have the one line that interrupts it. Read them once; they’ll be there when the moment is.

The “in-house” trap
When you hear“Our internal team is already building an AI oversight tool.”
You already knowThey’re building software to check software — walking into the innocence trap. They think it’s a capability problem. It’s a separation-of-duties problem.
You runDon’t argue with their engineers. Nod, then plant the flag: “That makes sense for operational analytics. How’s the firm handling the SOX / EU AI Act separation-of-duties requirement when the monitor shares a failure domain with the agent it’s watching?” Then step back and let it sit.
The closed-file trap
When you hear“Leadership already passed on this — the file is closed.”
You already knowHe didn’t reject the idea; he refused unmeasured risk while trusting his own experts — which is exactly what a good executive does. Asking him to reverse his team is asking for courage. Cover reopens files; courage keeps them closed.
You runNever re-argue the original pitch, and never ask anyone to overrule their experts. Route a question the leader can hand to his own risk chain at renewal: “Does our policy respond if a court finds we operated AI agents with no reporting system capable of assigning responsibility?” His own broker answers it, the reopening arrives as his diligence, not your persistence — and you’re on record as the one who handed over the umbrella before the storm.
The horizon trap
When you hear“We’ll just wait for the next generation of models — they won’t have these problems.”
You already knowThey think reliability is a bug that gets fixed. It’s undecidability — the horizon just moves. I don’t debate it; the math does.
You runOpen the laptop: “Don’t take my word for it — let’s run this on your preferred model right now.” Run the 90-second proof and show the wobble: “The model always wobbles. The hardware receipt never does.”
The budget trap
When you hear“We don’t have budget to onboard another software vendor right now.”
You already knowThey’re filing this as an IT expense. It’s an insurance clearing instrument that unlocks revenue they’re currently frozen out of.
You runRe-file it for them: “This isn’t a software subscription — it’s a per-agent license that lets you place the business your carriers are actively excluding. It’s the telematics dongle, not another SaaS seat.”

Action items

1. Show the one friend you trust — this is the fun one, and there’s no deadline on itnow
  • You said you might walk a friend you trust through this. Do it whenever the weekend allows — and here’s the move that makes it click: they open Claude Code in an empty folder, drop the prompt in, and ask it anything they’d have asked me. You don’t have to defend it or be in the room — the proof answers for itself, on their machine.
  • The showstopper for anyone technical: ask the same question of a weaker model and watch it flip-flop — then a stronger one, still no stable answer. That wobble isn’t a bug a bigger model fixes; it’s the proof that "a better model will settle it" never arrives — the horizon just moves. The one thing that doesn’t wobble is the signed receipt. That ninety seconds is the whole thesis, and it’s the part your friend won’t forget.
Monday: what your friend poked at, and where it landed.
2. Run it once yourself first — 20 minutes, one empty foldernow
  • Open Claude Code in a brand-new empty folder. Run: npx -y thetacog-mcp@latest attest-demo · hooper · prove-rice.
  • Then ask it: "Walk me through whether this is true, and ask the hardest skeptical questions you can." Do this before Saturday and the friend demo runs itself — you’ll already know where the sharp edges are.
Bring me the one question it couldn’t shake.
3. The 15-minute run-together — any slot you name, still opennow
  • You offered a Monday and it’s still on the table, whichever Monday works. 15 minutes, live (FaceTime / phone), while you paste the commands and I fix anything on the spot. No prep, no exhaustion tax — just the laptop and a new folder. Call anytime; I’m always around.
Name the slot — I’ll hold it.
4. Keep an ear to the ground — your one insider playnow
  • If anyone in the Emerging-Tech or cyber-practice orbit mentions hardware-verification models or the "Trust Debt" math, you say one line: "I actually know the architect who built that standard." Then stop — the line works because it isn’t a pitch.
  • Back pocket: thetadriven.com/d — if anyone asks what it is, just point them there.
Whether you’ve heard anything land, from either desk.
5. The internal owner — already found, closeddone
  • The material went straight to the two people who own this inside the firm — the Emerging-Technologies lead and the US cyber-practice lead — each by name, hand-delivered one-page brief and cover letter. Nothing for you to push, carry, or explain.
What we need from you — and what we don’t
You never have to produce a work product — no deck, no deliverable, no homework. That is not your job here.
The one thing worth more than any of that: tell us if you hear the material land on either desk. That’s the signal that turns this from a maybe into a real conversation — and you’re the only one positioned to hear it first.

When it comes up — the whole talking track

You never have to open this conversation. But the moment someone opens it near you, this is everything worth having in your pocket — staged from the one sentence you must remember down to the numbers a technical skeptic would ask for. Read it once; it will be there when you need it.

Line one — the only sentence you must remember
"I actually know the architect who built that standard." Say it and stop. Whoever raised it will carry the conversation from there, and every section below is just backup for wherever they take it. The worst move is to follow the line with a pitch — the line works because it isn’t one.
If you want to OPEN it rather than wait — the one question, and why it is a question
Since 28 July the front page of thetadriven.com asks one thing: "Do you worry about $1.2B in AI liability? We fixed the math." That is the whole opener, and it works as a question for a reason worth understanding before you ever use it. It is a filter, not a claim about anyone’s book — we are not telling your client they are carrying $1.2B, and you should never say that either. A round billion reads as rhetoric and everyone scrolls past; an oddly specific number reads as somebody’s actual exposure, so the only person who stops is the one who already has a figure of that shape in their head. If they flinch, they self-identified and you did not have to guess. If they shrug, they were never the reader and you spent one sentence finding out. Either answer is a win, which is why it beats a pitch: a pitch has to be right, a filter only has to be asked. The second half is the part that makes them move — if the semantic property is trivial, software can already check it, and why is anyone paying to check trivial properties; if it is not trivial, Rice’s theorem says nobody can check it, ever. That pincer has no third door, and the way out of it is the only thing we sell: where the agent’s action landed is decidable, signed, and re-runnable on their machine.
If they ask what it is — the telematics frame
"It’s telematics for AI agents — the OBD2 dongle. It doesn’t judge whether the work was good; it counts whether the agent stayed in the lane it was hired for, and signs a receipt either way." This is the frame that lands, because it’s the business your industry already runs: the dongle on a car never knows WHY the driver braked — a child in the street, a hundred times a day, entirely possible — and it doesn’t care. It prices the pattern. A book of drivers who look risky pays more; a book that drives clean pays less. Insurance has never needed ground truth. It needs a correlated, countable signal — and until now, AI had none.
If they push — why "green" doesn’t mean "good," and why that’s the point
Be the first to say it, because it’s the strength, not the hole: a green receipt does not certify the work was good. Whether work is "good" is formally undecidable — no instrument will ever certify it, ours included, and anyone who claims otherwise is selling something. What the receipt certifies is decidable: the agent stayed inside the domain it was hired for. Green is the seatbelt, not the guarantee. And when reds concentrate in exactly the lane an agent was working — red in the legal column while it edits a legal document — that is "you should have been paying attention," made countable.
The grand-slam line — "you should have checked"
Random spot-checks of AI work are zero-leverage — checking 1% at random catches 1%. The receipt makes oversight targetable: put human eyes on exactly the events that left the lane. From there the logic writes itself, and it’s 1932 logic your industry already knows (the tugboat with no radio): once a flagging instrument exists and is cheap, a flagged-and-unchecked event is legible negligence, and a flagged-and-checked-and-fine event is a countable, priceable act of oversight. One more piece and the loop closes: an AI checking an AI is never countable — the checker fails the same way the checked does. Only a receipt of human attention closes an insurable loop. That single fact is why this is an instrument business and not a better-model business.
The exclusion arithmetic — why the firm should care this quarter
From January 2026 the standard forms started carving AI out: an ISO generative-AI exclusion attaching to renewals, and "absolute AI" exclusions spreading through D&O, E&O, EPLI. Every one of those endorsements is a client your desks currently cannot fully place — a placement that walks away or self-insures. The signal is the thing that moves a client from excluded back to underwritable: telemetry a carrier can price against instead of a risk it can only refuse. The book of business locked behind those exclusions is the market, and right now nobody is placing it.
The license-inventory play — why a broker, specifically, wins this
The license side has one structural fact in it, and it is a market observation rather than a rule anyone imposes: an attestation only one party can read is worth nothing to the counterparty, so the deployer running the agents and the carrier pricing against the signal each tend to hold a license to the same patent — and the broker is the only seat at the table that touches both sides of every single deal. Nothing is conditioned on that; the same recomputable unit is sold to both sides at one price, and a deployer’s license is fully effective on its own. So be precise about what a broker actually earns here, because it is not a margin on the paper: the license is $20 per agent-year, sold for use, re-registered to whoever ends up running the agent for a flat administrative fee, with no resale spread available to anyone. What a broker earns is the placement — the exclusions took a book of business off your desks, and a signal a carrier can price against is what moves those accounts from excluded back to underwritable. Being early to a standard is a relationship position, not an inventory position. Full mechanics are public at thetadriven.com/pricing.
If anyone technical leans in — the numbers
All of it recomputes on a laptop in about twenty minutes, no access to anything of yours: 0.90 separation on a sealed, blind, cross-domain held-out; 10 of 10 off-domain deliverables rejected; 0 of 7 adversarial forgeries accepted; chip and cloud produce the same answer byte-for-byte (diff 0); one walk on the 144-node lattice runs in ~14 ms with no model in the loop; receipts are ed25519-signed coordinates and one-way hashes — never the client’s work product. And the honest fence, stated first because it’s what a good skeptic finds anyway: the ~13% breach rate is measured on our own lived ledger — a realized loss ratio, not a sealed blind premium — and mass-scale actuarial tuning is exactly what the first deployments exist to fund. Perturbation tests separate at many sigma; an insurer doesn’t need perfection, it needs better-than-random, stable, and countable.
Where this stands inside your building — so you never overreach
Both desks already have the material, by name, on paper. Nothing is waiting on you, and nothing should come FROM you — your entire position is the ear and line one. If either desk engages, point them at thetadriven.com/d and step back: at that moment you’ve already done the one thing nobody else could, which was being there when it surfaced. The bridge position is won by being early and light, never by pushing.

What else to research

You already named the first two. Here’s the rest of the trail — take it and run with it; this is yours to use in any room.

Rice’s theorem (1953) — you have this
Why no system can decide another system’s meaning from inside the same failure domain — the math under "undecidable."
Caremark + Judge Learned Hand / The T.J. Hooper (1932) — you have this
The duty of oversight, and the rule that once an available precaution exists, not using it is the negligence.
The symbol-grounding problem
Why an AI’s words aren’t welded to meaning — the deeper reason verification can’t be left to another model.
Silent-AI exclusions + ISO’s 2026 generative-AI exclusion
The concrete market move: coverage pulled from D&O / E&O exactly as exposure rises.
Affirmative AI products — Munich Re aiSelf, Armilla/Lloyd’s, Mosaic
What’s actually being sold today — and the white space: they price model drift, not domain drift.
"Second-order red flags" — Caremark-for-AI commentary
Oxford / Columbia / Akin Gump: model drift as a board-level duty. The legal hook — honest fence: not yet a court holding.
Usage-based insurance / telematics — Progressive Snapshot and the UBI book
The direct precedent: an industry that already prices a correlated behavioral signal without knowing why any single brake was pressed. The AI receipt is the same instrument pointed at agents.
Broker economics of a two-sided license
Why the seat that touches deployer AND carrier captures a standard: the fax machine, the Visa network, the DTCC — every clearing position started as whoever sat between both sides first.

Roles, goals & assumptions

Here is what you’d actually be signing up for — the role, the goals, and the assumptions behind them, stated plainly. None of it is fixed; it’s written down so you can push back on the exact parts that don’t match your reality. Tell me where I have it wrong.

why we believe this is true
We believe the strongest proof we have is also the simplest: a capable agent working in the WRONG lane — a brain surgeon doing the plumbing — is something we can catch, sign, and let a stranger recompute, beyond a shadow of a doubt. That one signed record turns a risk the market is actively EXCLUDING into one your clients can actually carry — and you’re positioned to be the person who saw it first.
Connection — where you already stand
You move between desks and can start a real conversation in any of them, and you’d rather hand a client a real answer — coverage that actually holds when they claim on it — than a quiet carve-out they discover too late. The role this assumes for you is the connector — the one inside who can route a thing to the people who decide.
Contribution — what this lets you give
This isn’t about what the instrument gives us. It’s what it lets YOU give: your clients a way to keep coverage that’s being written out from under them, and your firm a first-mover instrument instead of another memo. The contribution is yours to make.
Growth — the new capability
You become someone who can stand up a recomputable proof on a stranger’s laptop and explain undecidability to a skeptic in one sentence — a capability almost no one in your building has. That’s a goal worth naming out loud.
Uncertainty — the open variables (argue with these)
The honest open question: is the wedge specifically D&O oversight liability, or is it broader — E&O / professional liability, or a paid discovery engagement? We’ve been assuming D&O is the issue; your book is what confirms or corrects that. Also open: who owns this internally, and what "first step" actually fits your process. These are the things we should settle together.
Certainty — what does not depend on any of that
Regardless of how the above lands: the lane-and-drift verdict is decidable, signed, and reproduces byte-for-byte on any machine. That part you can verify yourself today — no trust, no agreement required.
Significance — who you become
The one inside who turned an excluded, undecidable risk into a priced, monitorable one — early, and on the record. That’s the role on offer if the goals below hold up.
A few things we take as given
  • The material already reached the two owners directly — we need you now as the listener, not the router. ← confirm?
  • The first dollar is a discovery engagement, not a product license. ← confirm?
  • The exposure that matters most to your clients is AI oversight liability (D&O). ← or is it really E&O / professional liability?
  • You can raise this as a discovery, not an accusation, with no career risk. ← confirm the framing that protects you.
  • The right first internal conversation is the cyber desk ↔ the D&O / management-liability group. ← which one first?

Evidence

The gap is being actively created
ISO introduced a generative-AI exclusion attaching to commercial general-liability renewals from January 2026 (ISO forms underpin ~82% of US P&C). "Absolute AI" exclusions (e.g. Berkley, Hamilton) are spreading into D&O, E&O, EPLI and fiduciary lines — one expressly excludes "a board’s failure to oversee AI."
Fenwick "The End of Silent AI"; Zelle Law; IndependentAgent (Mar 2026)
The legal hook
Multiple firms/academics argue the Caremark duty of oversight extends to AI: model drift and monitoring anomalies are "second-order red flags" a board must escalate. Honest fence: this is firm/academic commentary — no court has yet held it; the duty extends, liability is not easily triggered.
Akin Gump; Oxford Law Blog; Columbia Blue Sky (2026)
The white space
The affirmative AI products that exist (Munich Re aiSelf, Mosaic-aiSure ~$15M, Armilla/Lloyd’s ~$25M) price MODEL / KPI drift — not domain / competence drift. Nobody yet prices the "plumber doing brain surgery" event. The market is tiny and early ("≈five products worldwide") — first-mover, not crowded.
Munich Re; Mosaic/Business Insurance (Feb 2026); Armilla/Lloyd’s
The math and the standard
Why an AI can’t underwrite an AI: Rice’s theorem (1953) — no system can decide another’s semantic properties from inside the same failure domain. Why "a device now exists" matters: The T.J. Hooper (1932) — once an available precaution exists, not using it is the negligence.
Rice 1953; The T.J. Hooper 1932 · see thetadriven.com/pixel
The telematics precedent
Usage-based auto insurance built a multi-billion-dollar book on exactly this move: price the correlated signal (hard braking, mileage, hours) without ever knowing the ground truth of any single event. The dongle doesn’t know why the driver braked, and the actuary doesn’t need it to. The drift receipt is the same instrument class — countable behavioral telemetry — pointed at agents instead of drivers, which is why an insurance audience doesn’t need the thesis explained so much as recognized.
Progressive Snapshot / UBI industry practice since 2011
The license mechanics (public)
Per-agent annual license under US 19/637,714 — application pending, not a granted patent. The unit is metered like a lease: 365 days or 10,000 attestations, whichever comes first, sold for use rather than for resale. The unit price does not increase for a licensee for the duration of their agreement. Licenses are transferable to whoever will actually run the agent — re-registered through ThetaDriven for a flat administrative fee, keeping the caps they were minted with, and never above the unit price plus that fee. The same recomputable unit is sold to the deployer and to the carrier, with no condition that either accept the other. The full terms are public.
thetadriven.com/pricing
The unit under the license — what an "agent-year" actually counts
The denominator is not wall-clock time; it is receipts. A doctor performs a countable number of attestable actions a year; an agent performs the same kind of countable actions, just faster — a complicated decision with many insurable points burns several receipts, a simple one burns one. Pricing per receipt-capacity rather than per calendar year is what lets the same license structure hold as agents speed up.
the 144-lattice receipt model · thetadriven.com/pixel
The lane-and-drift verdict is decidable and recomputable today; the price is advisory until it’s calibrated on real claim history. That fence is deliberate.