The Benchmark · public adoption board
The claim, stated so you can attack it: an AI attestation only means something if the number can be recomputed by the party who doesn't trust it — including when the number prices its own administrator badly, which that one does. Everything else on this board is built the same way round: 3,154 receipts sealed in 30 days, 2 independent forks of the meter, 1 star, 0 licensed constituents, every one of them read from the public fork graph of github.com/wiber/thetacog-mcp, countersigned registers, and an append-only ledger that replays shape-identical. No self-reporting anywhere in the loop, ours least of all — the first constituent of an honest index is the org that runs it.
You don't have to trust this
npx -y thetacog-mcp@latest attest-demoThat command runs the meter locally — your machine, your data, a recomputable receipt back. What it returns is the same object every metric below counts. Prefer a browser? Verify a signed receipt at /verify-receipt — signature check plus a full recompute, no install.
The index · accumulating
constituents 1 priced / 1 registered
readings pooled 200
vega, readings-weighted 7.49
dispersion n/a
INDEX OF ONE — the administrator is currently the only priced constituent. The pooled number is our own, said plainly rather than dressed as a market.
Live: /.well-known/semantic-vega-index.json · the constituent registry
The pooled vega is the readings-weighted mean of published constituent vegas. Not a standard deviation over a merged series: we hold prints, never anyone else's rows, so a merged-series number would be one we are not entitled to compute. Dispersion across constituents is reported once there are at least two of them and is null with its reason until then — a lone constituent has no dispersion, and printing zero would read as perfect agreement when it means there is nobody to disagree with.
The administrator's own print is read from disk so the index is never a deploy stale about itself. Every other constituent is fetched from its own domain, because an index that reads its constituents out of its own database is not an index. Unreachable constituents are reported as unreachable rather than dropped, and never carried forward from an earlier run — a number that survives its own source going dark is the beginning of a fiction.
Publishing the aggregate before there is an aggregate to be proud of is the point. An index of one, labelled an index of one, with a join path that costs a pull request, is a stronger position than a polished figure nobody else can enter — and it is the honest version of what comes next, which is the T-ladder below.
Depth: fork the instrument, not the vendor's verdict · countable precedes accountable.
The self-metric · published, not reported
A receipt is one reading: this commit, this lane, this deviation. A semantic vega is the dispersion of your readings over time — how far delivery wanders around declared intent, run after run. That second number is the one a parametric warranty is written against, because two repositories with the same average can have completely different tails, and the tail is the entire product. Volatility is what an underwriter buys and sells; here it is measured over your own history rather than assumed from a sector table.
thetadrivencoach · curve v0-advisory · advisory, not a filed rate
semantic vega 7.49
series 200 rows · 6a5efbc7d4747fb0
breach at the gate's kill line 19.5%
excess kurtosis 0.1
advisory premium 214 GPU-credit-hours / agent-month (in kind, never fiat indemnity)
SENSOR-UNCALIBRATED: the miss classes this gate catches in this domain have not been measured yet, so the premium above is a shape, not a rate. Stated here rather than omitted. Live JSON: /.well-known/semantic-vega.json
That reading does not flatter us. A vega above ten with a fat tail is a repository that moves in bursts, and the advisory curve prices it accordingly — ours loads at more than twice the base. Publishing the number that prices us worst is not humility, it is the only order of operations that leaves the index worth anything: the administrator's reading goes up first, uncorrected, or the board is a scoreboard its own scorekeeper is exempt from.
Publish yours — two commands, your machine, your tape
npx -y thetacog-mcp@latest vega-backtestnpx -y thetacog-mcp@latest advisory-premium --series attest-out/backtest-series.ndjsonThe first replays the meter over your existing git log — every commit becomes one row, thin commits refuse rather than get faked into the statistics — so a repository that installed the meter this morning still has a day-one vega computed from its own history. The second prices that series on the published curve. Neither call leaves your machine, and both are recomputable: the same log produces the same print, on anyone's silicon.
Then publish it at the same address we do — /.well-known/semantic-vega.json — and the index builds itself. One small file per repository, at a path that is identical everywhere, means an index of publishers is assembled by fetching rather than by asking permission, and any stranger can assemble the same one. That is the map of maps: this board is not the territory, it is an index over prints that each belong to whoever produced them. The meter that computes yours is MIT and stays MIT, with no field-of-use restriction, ever — which is the load-bearing part. A metric you cannot compute without the administrator's cooperation is a metric the administrator can quietly bend.
Three properties make a metric healthy, and all three are structural rather than promised. It is recomputable by the party who doesn't trust it, which removes the need for anyone's good faith. The administrator's own number is on the board first, which removes the exemption. And gaming it costs more than fixing it: drift is measured in the same lattice the remediation bounty pays out of, so engineering a flattering reading means doing the work that would have produced an honest one, with extra steps.
Depth: fork the instrument, not the vendor's verdict · countable precedes accountable.
The loss side · schedule v0-advisory-proposed
Two sentences if that is all you need: the pool never pays surge pricing — rates are pre-agreed per severity band, in the paper, before any incident. And hours release only when a recomputed test passes on the public tape — nobody bills on the honour system.
| band | eng hrs | compute hrs | ceiling | on our tape |
|---|---|---|---|---|
| B1 DIAGNOSTIC | 2 | 8 | 4 | 36 |
| B2 BREACH | 8 | 40 | 12 | 18 |
| B3 PERSISTENT | 24 | 120 | 32 | 21 |
75 events over 200 readings = 720 engineer-hours and 3,528 compute-hours, against a rule-bounded ceiling of 1,032 engineer-hours. Measured on the same rows the premium is priced from — not modelled.
An hour is boundable and a cash indemnity is not. Compute hours have a market price and engineering hours have a rate card, so the maximum draw per event is known before the event happens — which is what lets a reserve be sized, and what keeps the repair form intact rather than sliding into open-ended indemnity. That is the whole reason the payout is in kind, and it is the same reason the collateral is compute rather than cash.
Bands select on the receipt's own published fields, never on an adjuster's read: the lens diagnostic line, the gate's kill threshold, and persistence — a run of 3 or more consecutive breaching readings, because a single breach and a breach that will not clear are different failures and only one of them needs a deep bundle. Anyone holding the tape can recompute which band an event falls in; the marketplace clears WHO is dispatched, never what an hour costs mid-incident.
Schedule v0-advisory-proposed — every constant published, none of them filed. Superseded by the T5 hardening documents — thresholds, bundles and ceilings fixed there, per pool, in the paper. Depth: the payout arrives as labour · we don't insure the work, we insure the worker · you never insure the catastrophe, you insure the crossing.
A benchmark earns the right to be referenced the way an index does: inclusion criteria first, constituents second. The ladder starts at T0 because individuals count: the fork graph is partitioned by owner type, so one person carrying the meter gets a row of its own rather than being filtered out on the way to an institutional headline. The licensed counts further up start at zero and are tracked in public beside the live baselines that ground them — the sequencing is the product. Nothing downstream may reference this oracle until the hardening documents (T5) exist in public, then a pool (T6a), then riders (T6b), then anything further (T6c). A board that can't show its discipline early can't be trusted at a thousand constituents.
The loop · every step measured on this board
The rows below are not a roadmap of intentions. They are one loop, cut at the seven places where it can be counted by someone who does not work here — which is why every milestone on this board is a number a stranger can recompute rather than a status we award ourselves. Read in order, the steps answer the question the whole ladder exists to answer: what has to be true, and measurably true, before anyone can write a warranty against an AI agent's deviation from what it was asked to do.
2 independent forks · 104 private cloners in 14 days · 1,585 package fetches in 30 days
Reach is not marketing here, it is the verification layer: every stranger holding the meter is someone who can recompute any claim on this page, including the ones about them and including the administrator’s own. An oracle nobody outside the building can run is an opinion with a domain name.
0 qualifying · 0 org-owned forks in the graph
A fork with CI enabled is the first moment somebody else’s reality enters the tape instead of ours. Their pipeline runs the meter over their code, the receipts commit to their repository, and what becomes public is a placement rather than a diff — which is why this step needs no procurement conversation and no trust.
0 merged upstream
A merged pull request is how a fixer claims territory in the open: coverage published before any incident exists, in a register nobody can backdate. Mapping territory is claiming it, so the meter deepens fastest exactly where the people who will later be dispatched already work.
0 agent-years under licence
The hinge, and the row most easily misread as a price list. You cannot write a warranty against a phenomenon — nor an option, a treaty, or anything derivative-shaped. You write against a UNIT and a TERM, and the agent-year is both: one agent, 365 days or 10,000 sealed attestations, whichever comes first. Until agent-years are under licence there is no notional for a premium to attach to, no exposure period for a claim to run against, and no denominator to divide losses by, which is why the rate card is published while the count is still zero. The other half is the market this unit exists for: an agent acts in milliseconds and does not wait for a human to sign off, so responsibility has to attach at that speed or it does not attach at all. A unit that needs review before it counts cannot cover a market that clears before the review starts.
0 from licensed deployments · 3,154 sealed on our own tape in 30 days · our own vega 7.49
What gets priced was never the level of one receipt. It is the dispersion of readings around declared intent over time — how often the reading breaches, how wide it scatters, how heavy the tail is when it does. Receipts per month is the frequency half, the semantic vega above is the dispersion half, and both come off a tape that replays shape-identical.
3 of 3 published
Three public documents, published to be attacked: the lit-mass floor, the portfolio distributional audit, and the false-negative protocol — the last being the testable procedure for when the trigger fails to fire on a real breach. Nothing downstream may reference the print until all three exist. This is the gate the whole loop sits behind, and the only step entirely inside our own hands.
0 pools · 0 riders · 0 further
A pool references the print, a rider follows a performed quarter, anything further follows a demonstrated record. Then the loop closes instead of ending: the fee funds standing capacity, a trigger dispatches to whoever holds the highest-confidence pixel over the failure region, and that dispatch is billable work for the contributors who mapped the territory back at step 3. Paid work deepens coverage, deeper coverage improves the meter, a better meter reaches further — and wider reach means more strangers able to verify the print the paper settles against.
Why the trigger being parametric is the whole unlock
Cover on a novel exposure is rationed by the scarcest input in specialty insurance: expert judgment. Somebody has to adjudicate each loss, that somebody cannot be hired quickly, and that — not appetite — is why agent liability today is quoted narrowly or declined outright. A trigger that settles against a recomputable print removes the adjudication step, and with it the ceiling judgment imposes. What still bounds the paper is the reserve, and a reserve held in kind — compute and remediation hours — is an input you can buy, lease out while it sits unclaimed, and then spend in exactly the currency a payout consumes. The constraint stops being how many experts you can find and becomes how much capacity you can hold, which is a constraint that answers to money instead of to hiring.
From the deployer's side that arithmetic is about speed, not safety. Coverage that must be adjudicated arrives at the pace of a risk committee; coverage that settles against a reading arrives at the pace you ship. The question stops being whether you are permitted to put the agent into production this quarter and becomes whether this run landed in lane — which is the question you were going to ask anyway. Competence cover in kind is what lets an organisation deploy faster rather than merely more carefully.
It is also why the object being measured was never “AI risk”. Hand an executive a wand that answers one question about their organisation and almost nobody asks whether a model will misbehave. They ask is my organisation doing what I intended? That question has been unanswerable for as long as organisations have existed, because it had no unit. Step 4 gives it one, the print gives it an answer, and the paper is what you get when the answer is countable enough to settle against.
Naming the limit is the honest half of the claim. Removing adjudication removes a throughput ceiling; it does not remove exposure. A pool can only promise the hours it actually holds, which is why the loss side on this page is denominated in engineer-hours and compute-hours with a stated ceiling per event rather than in an open-ended sum. Nor does the print judge quality: it decides WHERE work landed against declared intent, and whether an arbitrary system is free of defects stays undecidable — the claim perimeter further down this page says so in the same words we would use under oath.
The second bound is calibration. A gate that cannot see a domain's vocabulary is stably blind, and stability alone would price beautifully while measuring nothing. That is what the battery exists to catch before money attaches, and why the reading above carries its own calibration state rather than presenting a premium as a filed rate.
The countdown, in measured units
What stands between this board and the first AI-risk warranty is not time, it is four countable things: 0 hardening documents still to publish (T5 stands at 3 of 3), then the first sponsored pool at T6a (0 live), then one full continuous quarter of that pool performing, then the first rider at T6b (0 live). No date is offered here, because a date would be the one number on this page nobody could recompute — and every other number is chosen precisely so you never have to take ours for it.
stars: 1 (0 excluding the repo owner) · watchers: 0 · first independent fork 2026-07-10, most recent 2026-07-24 — read straight off the GitHub fork graph
What countsA fork of the MIT meter repo (github.com/wiber/thetacog-mcp) owned by a GitHub User rather than an Organization. One person deciding the meter is worth carrying is a real adoption event and it gets its own row — the fork graph is partitioned by owner type, never filtered down to institutions with the rest thrown away. Every fork appears in exactly one of T0 or T1, and T0 + T1 equals the public fork count anyone can read off GitHub.
Recompute itThe public fork graph — `gh api repos/wiber/thetacog-mcp/forks`, partitioned on owner.type. Two API calls and you have the same numbers, including the star count with the repo owner subtracted out so our own star is disclosed rather than quietly counted.
This is you if you run agents on your own machine and would rather hold a receipt you computed than a dashboard someone else controls.
Most aligned: Independent AI engineers & consultants · Solo operators running agent fleets · Researchers reproducing attestation claims · Anyone who forks before they trust
Forward to the engineer who reads the source before the README.
Forward this metric →1,585 npm installs in the last 30 days · 224 clones total · 32 unique repo visitors — none of them appear in the fork graph
What countsAnyone who took the meter without leaving a public trace: a git clone, or an npm install of the package. A fork is the only adoption GitHub draws as a graph, and it is by far the rarest kind — most people who try a tool clone it, run it, and are never counted anywhere. They are adopters and this row says so. Two aggregate APIs make them countable without making them identifiable: repo traffic (a rolling 14-day window, readable only with push access, so this number is generated from a laptop and never in CI) and the public npm download counter. Neither can name a person, which is the point — we can thank you in aggregate and could not deanonymize you if we wanted to.
Recompute itGitHub repo traffic (/traffic/clones, /traffic/views — 14-day rolling, push access required) and api.npmjs.org/downloads/point/last-month/thetacog-mcp, which anyone can curl.
This is you if you already ran the meter privately and told nobody. Nothing is owed and nothing is asked — but if it produced a receipt that surprised you, that is the one thing worth sending back.
Most aligned: Engineers evaluating quietly before telling anyone · Teams whose policy forbids public forks · npm installers who never opened the repo
No forward needed — this row is a thank-you.
Forward this metric →org-owned forks in the graph today: 0 of 2 public forks — the rest are counted in T0, not discarded
What countsAn Organization-owned fork of the MIT meter repo (github.com/wiber/thetacog-mcp) at institutional scale (public-company scale, ≥$100M fund AUM, or a frontier-run lab) whose CI runs the attestation suite. The fork is the org's own attestation surface: signed receipts — verdict, cell, σ, the encircled panel — commit to the forker's own repo. Their source never leaves their building; the union of fork receipts is the public actuarial tape.
Recompute itPublic GitHub fork graph + the receipts committed in each fork — no self-reporting anywhere in the loop.
This is you if you own platform, model-risk, or AI-governance at an org that size — and your CI could run one extra suite this quarter.
Most aligned: Foundation-model platform teams · Hyperscaler AI infrastructure groups · Bank / insurer model-risk functions · Open-source program offices
Forward to whoever owns your model-risk CI.
Forward this metric →What countsA pull request from a qualifying fork, merged into github.com/wiber/thetacog-mcp main. Public merge history is the register.
Recompute itGitHub merge history — recomputable by anyone with a browser.
This is you if your engineers already patched their fork and the patch is good enough to send home.
Most aligned: Staff engineers inside T1 forks · OSPO maintain-upstream policies
Forward to the engineer who complains about carrying patches.
Forward this metric →What countsExecuted licenses at the published rate card: $20/agent-year, where one agent-year = 365 days or 10,000 sealed attestations, whichever comes first. Settlement-style referencing is licensed separately, priced in basis points on limit/notional per Schedule A. Counted org-by-org from the countersigned register. Read the row as the UNIT step of the loop above rather than as a price list: an agent-year is the unit and the exposure period every referencing structure is arithmetic over — warranty, option, treaty alike — so a count of zero here is precisely the reason no paper exists yet, and the rate is published before the count moves so nobody has to ask what it will be.
Recompute itThe signed, hash-chained license tape — each license mints an ed25519 identity and chains to the previous entry, so the register itself is recomputable.
This is you if you run 100+ agents in production and one named person answers for what they do.
Most aligned: Heads of AI governance / CIO office · Tech-E&O buyers and their brokers · Procurement at agent-fleet operators
Forward to whoever signs for your agents’ errors.
Forward this metric →engine baseline, tracked from our own tape: 3,154 receipts sealed in the trailing 30 days across 2,502 commits — we are constituent #1 of our own index
What countsAttestation receipts sealed on the canonical ledger by licensed deployments, sustained month over month.
Recompute itThe tape itself — deterministic, LLM-free, re-runnable by any licensee.
This is you if your fleet ships daily and every Monday someone asks what the agents actually did.
Most aligned: SRE / platform owners of agent fleets · Internal audit at AI-heavy operators
Forward to the person who compiles the Monday agent report by hand.
Forward this metric →What countsThree named public documents, each versioned: (1) the lit-mass floor — the minimum measured (non-UNMEASURED) surface before any receipt is eligible to feed a downstream product; (2) the portfolio distributional audit — the receipt series’ statistical properties characterized across a portfolio of companies, published, not spot-checked; (3) the false-negative protocol — the documented, testable procedure for when the trigger fails to fire on a real breach. Nothing downstream references the oracle before all three exist.
Recompute itThe documents themselves, versioned and public — their existence is the metric. Open each one below, sha256 the file you are served, and compare the first sixteen hex characters to the hash printed beside it. No repo access, no account, no trust in us required.
v1 · sha256 9951ca17b474092a
v1 · sha256 d616f486d6a190a3
v1 · sha256 e316a75fa22fb745
This is you if you are an actuary, cat-modeler, or model-validation quant who would enjoy tearing these three documents apart before anyone relies on them.
Most aligned: Actuaries & catastrophe modelers · Model-validation quants · Parametric-trigger skeptics (especially)
Forward to the sharpest skeptic you know — that is the job posting.
Forward this metric →What countsThree sub-registers, in mandatory order. T6a remediation pool (0 live): the first third-party pool referencing the index — eligible only after all three T5 documents publish. T6b SLA rider (0 live): eligible only after a T6a pool has performed a full continuous quarter. T6c further paper (0 live): anything further derivative-shaped — eligible only after T6b has a demonstrated record. Their paper, their regulatory perimeter, our settlement-reference license. Zero here is a promise kept, not a gap.
Recompute itExecuted settlement-reference licenses, counted from the register.
This is you if you structure parametric paper — MGA, ILS desk, reinsurance innovation team — and want a trigger you can recompute instead of trust.
Most aligned: MGAs & program administrators · ILS structurers · Reinsurance innovation teams · Parametric-insurance desks
Forward to the desk that got burned by a trigger it could not audit.
Forward this metric →The gate — T5
A parametric warranty pays out on a rule a stranger can rerun, so the rule has to be written down and survive a hostile read before any pool references it. Three engineering specifications carry the whole claim. Nothing downstream — no pool, no rider — references this oracle until all three are public. Their completion status is tracked on the board above; what each has to establish is here.
What counts as inside the spec you gave the agent — the minimum measured surface a commit must light before any rate is quotable at all. Below the floor the reading is honestly direction-only, never a number dressed up as coverage. It also names the failure mode out loud: drift is measured against a specification, so a vague spec makes the number worth little — “decidably drifted, but from what?” The aperture can infer intent from semantic momentum when a spec is thin, but a warranty is only ever as sound as the spec it is written against. This is the document that stops a green instrument from meaning “we didn't look” — or “we never said what looking meant.”
And the signal is not a model's self-report or a log's say-so. Positions are addresses in a fixed lattice, and a semantic boundary crossing is a physical one: a dependent load's latency across a 64-byte cache line is a cost a stranger can re-time on their own machine, from userspace, no privileged counters required. That is what makes “recomputable” mean something harder than “the same script ran twice.”
2 · The portfolio distributional audit →
The statistical properties of the receipt series across a portfolio of companies and commits — so a breach rate is measured against a distribution, not asserted. The object being counted is Trust Debt: the delta between what an agent was specified to do and what it ungroundedly did, permanently on the record. The recompute is the point: every reading replays shape-identical on a stranger's machine, scored against random and impostor realities so the number means something.
3 · The false-negative protocol →
The documented, testable procedure for the case that actually matters: the trigger fails to fire on a real breach. The hard case is a semantic boundary, not a crash — you specified a plumber and it started acting like an electrician, and the whole question is where plumbing ends and wiring begins. A fail-closed gate is supposed to halt the out-of-role action before it reaches the ledger; a warranty is only as sound as its misses are bounded, and the payout condition lives or dies on how well this boundary is caught.
Swing at any of the three — we would rather you find the falsifier than not. Show a lit-mass floor that passes a spec nobody actually wrote; a breach rate that does not move when you swap in an impostor reality (σ is measured against random and impostor realities, never asserted — if it does not move, the instrument is decorative); or a semantic boundary the fail-closed gate lets through. And note the standard the whole rack is held to: decidable is a higher bar than interpretable — an objective, discrete, recomputable state, not a qualitative read. Tagging output with a model adds metadata; it does not make the record decidable, and a warranty cannot trigger on metadata.
Why an unusual channel is putting these in front of you at all is a separate argument, made plainly and without naming anyone: the pamphlet, the disclosed playbook, and the advisory invoice it points people to.
Upstream of the paper
The T6 rungs above are financial structures only in the sense an index future is: paper that settles against an official print. Read as pure structure, the whole stack is one line:
sealed print → parametric trigger → warranty paper → in-kind payout (GPU-hours + remediation hours) → compute reserve
Three structural facts carry it. The trigger is parametric: the sealed benchmark reading — public, recomputable by the party who doesn't trust it — never an adjuster's judgment. The payout is in-kind: remediation capacity — GPU-hours and engineering hours, dispatched through an accredited remediation network, never open-ended cash indemnity. The repair form is what keeps the paper in warranty territory rather than insurance territory, a determination each sponsor makes with its own counsel; the accreditation of who may perform the repair stays upstream, with the mark. The collateral is compute: a reserve denominated in GPU-hours is a hard, liquid, priceable asset — infrastructure, not code. Nothing in the paper requires the code explained; it requires the print licensed.
We call that paper a parametric warranty, and the precision is not cosmetic. Insurance indemnifies a judged loss in cash. A warranty promises a repair — and a parametric one promises it against a published reading instead of an adjuster's opinion. Which lane a particular structure lands in is a determination its sponsor makes with its own counsel, state by state; what is stated here is which lane this is built for: repair form, accredited network, capacity in kind, never open-ended cash indemnity. The other half of that sentence travels with it always — a note that funds the reserve is a security inside its sponsor's own perimeter. Nothing in this stack is unregulated. It is differently regulated, on purpose.
What a trigger looks like from the inside is worth saying plainly, because it is the part that sounds like marketing right up until you watch it happen. Nobody rings a claims line. The print is the notice: the reading crosses the band written into the paper, and dispatch goes out against the coverage map to whoever holds the highest-confidence pixel over that failure region. They are not a vendor you selected, and not a bench you were paying to sit idle. They are accredited under the mark, rated by coverage they published before your incident existed, and they arrive already knowing which region of your system moved — because the lattice that fired the trigger is the same one they hold territory in. The first thing they hand you is not an invoice. It is the receipt that dispatched them.
Depth: you insure the crossing, not the catastrophe · what a countable downside is worth · why the barbell's precondition has to be manufactured.
Two questions are open on any pool that would reference this oracle, and they are the sponsor's to answer, not ours:
Sequencing is unchanged by anyone's enthusiasm: hardening documents (T5) public first, then a pool (T6a), then riders (T6b). The benchmark is administered upstream; everything downstream is created independently by third-party sponsors, under their own counsel, in their own regulatory perimeter.
Why a benchmark at all
The object measured here is not a property of AI. It is deviation from declared intent — did we get what we asked for — the oldest unpriced cost in any principal–agent relationship. Agentic AI is the wedge because there the deviation is currently uninsurable; the lattice never cared whether the agent is a model, a vendor, or a team. Four roles close the loop, and none of them requires trusting another:
The flywheel is the point: every payout event is simultaneously billable work for a fixer, a confidence event for a deployer, and a new constituent on this board. Premium turned over into remediation capacity is not a cost — it is the sector being built, in public, one pixel at a time. The inversion is deliberate: as agentic systems scale, the premium on grounded human judgment rises with them — the network is paid to be exactly what the models are not.
The bargain
Read them together and the design goal is visible: no participant is exposed to another participant's opinion. The trigger is a reading, not an adjuster. The dispatch is a map, not a phone tree. The rate is a band, not a negotiation. The failure mode that remains is the honest one — not doing the thing you said you would do — and that one is supposed to be visible, because it is the entire object being measured.
The meter, run on its owner
Everything above asks you to recompute our numbers. So the honest place to start is the numbers we got wrong — 3 of them, inside our own operations, in the 2 days before this page was last built. None was caught by being careful. Each was caught by an oracle the claimant did not control.
2026-08-16
How many receipts this repository holds
declared 3,425 on one deck, 3,460 on the other
observed 3,485
the tell — Two surfaces, one fact, two numbers — neither recomputable
verify: git show cb5269fee^:src/app/deck/actuary/page.tsx | grep RECEIPT_DIRS
2026-08-16
Whether thetacog-mcp 2.52.0 was published
declared 2.52.0
observed 2.51.0
the tell — The script already checked, printed the answer, and discarded it
verify: npm view thetacog-mcp version --prefer-online
2026-08-15
Whether delegated work was being executed
declared A queue of pending asks
observed 15 consecutive failed fires over 27h, every ask re-queued
the tell — A failure that re-queues its own work is indistinguishable from a backlog
verify: grep '"status":"failed"' .thetacog/delegation-fires/fires.ndjson | wc -l
One mechanism produced all three, and it is not carelessness. In each case the claim and the observation were stored as a single number instead of two. Collapse them and the disagreement has nothing left to disagree with — which is why the fix is never “be more careful” and always a field that holds both. That is the entire content of an attestation, and the reason it has to be recomputable by the party who does not trust it: we are that party for our own numbers, and we still needed the instrument to catch these.
The claim perimeter
WHERE a state landed is provable, decidable, and LLM-free — recompute it.
WHETHER an arbitrary system is bug-free is UNDECIDABLE (Rice) — we never claim it.
We state the boundary exactly so the claims inside it stay unpuncturable. And the perimeter in the other direction: we license the decidable attestation oracle and the accreditation standard. We never issue, underwrite, structure, or market financial instruments — every referencing structure is created independently by third-party sponsors under their own counsel, in their own regulatory perimeter.
Each metric above carries a fit-filter. If one describes you, the fastest path is the command above — run it, then write us with the receipt it returned. If none describes you but one describes someone you know, use the forward button on that row: every metric is written to survive being sent one hop without explanation.
elias@thetadriven.com · the meter code is MIT; the administration — sealed ledger, certification mark, settlement reference — is what stays owned.
Do you worry about $1.2B in AI liability?
If the property is trivial, software can check it — and why are you paying to check trivial properties? If it isn’t trivial, Rice’s theorem says nobody can. So we fixed the math.
a number we can call — or whatever you would actually ask
Who did this make you think of? We’d love to know.