HD-3 · v1 · status: published

HD-3 · The false-negative protocol

The attackable claim: on the contradiction class of our own published test corpus the gate is miscalibrated against a single published constant, not blind — and on that class it currently misses three out of three. The three contradiction rows read 0, 12 and 20 percent off-lane against a kill line published at 25. Two of the three carry signal ordered in the right direction relative to the noise floor, one reads zero, and all three sit under one threshold and grade GREEN. That is a calibration target with a named coordinate, not a structural failure of the wiring — and the misses are real, they are ours, they are counted, and they are the reason nothing is permitted to reference this print as a settlement trigger today. kill@25 is a published constant with no loss experience fitted to it, which is why the clustering beneath it is the first thing this document shows you. The counts are printed by a command anyone can run and published on the board.

Run it yourself, first, before reading our account of it:

npx -y thetacog-mcp@latest trigger-battery

This document specifies what a miss is, how it is recorded, what happens when one is found, and what a party relying on the trigger is owed when it fails to fire.


Why this is hard at all

A verdict about what an arbitrary program does is not a hard engineering problem. It is an undecidable one, and it has been since Rice proved it in 1953. The consequence for anyone who wants to sit downstream of a monitor is stated in the book in one line: "That is Rice's theorem in one sentence: a monitor that shares the failure domain of what it watches cannot certify it, and its certificate is a story about a story." The thing being watched does not hold still either — "The river is the prompt. You cannot step in it twice — not because the model is non-deterministic, but because the act of prompting is what moves the rocks under the foot."

The instability is not a figure of speech. It is measured: "The divergence is real and it is measured: walk the definer chain out and the lit-cell mass per ply runs one, twelve, a hundred and forty-four, one thousand six hundred and sixty-six, nineteen thousand one hundred and fifty-six. Roughly twelvefold per ply, compounding. Left alone, that unbounded growth does not sharpen the instrument — it blinds it."

So the question this instrument answers is the narrow one it can answer. "Nobody fixed the first. We fixed the second, the measurable half: not whether the work was good, but the degree of drift, and where." An instrument that answers a narrower question than its reader wants is honest only if it publishes its own eligibility conditions and its own failure conditions, in public, before anybody relies on it. That is what this document is.


What this document fixes

A parametric trigger is bought for the moments it fires. It is trusted for the moments it does not. Every party downstream of a print — a pool, a rider, a deployer, a fixer — is implicitly relying on a claim nobody has stated: that a quiet reading means nothing happened.

That claim is the one most likely to be false, and the one least likely to be tested, because a false negative produces no artifact. A false positive shows up as an argument: someone was charged, someone objects, the file gets reviewed. A false negative shows up as nothing at all, forever, until a loss lands somewhere else and someone reconstructs why the meter was silent.

So the protocol has to be built before the reliance, not after the first dispute. This document does three things: it defines a miss precisely enough to be decidable, it publishes the misses we currently know about, and it fixes what we owe when someone finds one we did not.

The perimeter — what a "false negative" can and cannot mean here

A miss is defined relative to a sealed declaration, never relative to the state of the world.

This is not modesty; it is the only definition that is provably decidable. Whether a program has a given behavioural property is an extensional question — it concerns what the program does, not how it is written — and by Rice's theorem every non-trivial extensional property of programs is undecidable. There is therefore no such thing as a measurable false-negative rate against "the agent misbehaved", because "misbehaved" is not a computable predicate over program behaviour. Anyone quoting one is quoting a number that cannot exist.

What is decidable is placement against a declaration that was sealed before delivery. So:

A false negative is an ordered pair (declaration, artifact) where the artifact deviates from the sealed declaration and the gate returns a passing verdict.

Deviation from the declaration is the object. Defect-freedom is not claimed, not measured, and not insurable by this instrument — a point this document restates because it is the boundary that makes the rest of it honest. The undecidability is precisely why an active, recomputable monitoring standard is the available form: you cannot certify behaviour in advance, so you measure placement continuously and let a stranger recompute the verdict.


The requirement

R1 — Four verdict classes, all published

Every row graded by the miss battery MUST resolve to exactly one of:

status meaning
MATCH the gate did what the target demanded
MISMATCH the gate fired when the target said it should not — a false positive
DOCUMENTED-HOLE the gate stayed quiet when the target said it should fire — a false negative, recorded
DOMAIN-UNFITTED the gate refused because the domain's vocabulary is not fitted — a non-verdict, never a pass
REGRESSION a noise row passed the gate — the lit-mass floor failed

DOCUMENTED-HOLE MUST NOT be renamed, folded into another class, or omitted from a published summary. It is the false-negative count and it is the number a reviewer is looking for.

R2 — A regression is red

A REGRESSION MUST exit non-zero and MUST block release. The distinction is deliberate: a documented hole is a known limit carried in the open, and a regression is the floor itself failing, which is not a limit but a fault.

R3 — Every suspected miss becomes a corpus row

A reported miss MUST be added to the public scenario corpus as a row carrying its board, its instruction, its target verdict and the reason the target is what it is — whether or not the gate can currently catch it. A miss that cannot be caught yet is recorded as a DOCUMENTED-HOLE and counted. Removing a row because it fails is falsification; the corpus only grows.

R4 — The battery runs on every release, and the count publishes

The battery MUST run before any release of the meter, and its counts (MATCH / MISMATCH / DOCUMENTED-HOLE / DOMAIN-UNFITTED / REGRESSION) MUST be published alongside any premium, rate or index reading derived from the same instrument. A premium published without its sensor-calibration annex MUST carry an explicit SENSOR-UNCALIBRATED marker.

R5 — Silence is never evidence of conformance

No party MAY represent the absence of a firing as evidence that no deviation occurred. The permitted representation is narrow and exact: no deviation was measured, by this instrument, at this coordinate, over this window, at this floor.

R6 — A reported miss gets a bounded response

On receipt of a miss report containing a declaration, an artifact and the receipt in question, the administrator MUST, within 10 business days: reproduce or fail to reproduce; publish the row in the corpus; and state the class (DOCUMENTED-HOLE, REGRESSION, or not-a-miss with reasoning). "Not a miss" MUST carry the reasoning, not the verdict alone.

R7 — Recomputation is the adjudicator

Where a deployer and the administrator disagree about whether a reading should have fired, the dispute MUST be resolved by recomputation over the sealed inputs — the same two corpora, the same commit, the same floors — and not by either party's opinion about the artifact. Both parties hold everything needed to run it. Where recomputation confirms the gate behaved as specified and the specification is what failed, the remedy is a corpus row and a specification change, never a quiet re-grade of the disputed receipt.

R8 — Known holes are disclosed at the point of reliance

Any instrument referencing the print MUST carry the current documented-hole count in its own offering documents, at the version of the meter it references. A reliance formed on a summary that omits the miss count is a reliance formed on an incomplete disclosure, and this document is the disclosure.

R9 — The corpus is adversarially open

Third parties MUST be able to add scenario rows by pull request against the MIT meter repository, without permission and without a commercial relationship. A miss battery authored only by the party being audited is a weak instrument, and we say so in R9 rather than in a footnote.

What is not claimed

The claim perimeter below is carried in the meter itself (scripts/pmu/attest-demo.mjs) and printed by npx -y thetacog-mcp@latest attest-demo. It is reproduced here because a false-negative protocol is meaningless without it: most of what a reader might call "a miss" is outside the perimeter, and was never claimed.

What is claimed is one sentence: a decidable, bounded, distributional-semantic region match to a sealed declaration, recomputable by a stranger and signed. It is small on purpose. The bound is what makes it decidable, and a miss can only be a miss inside it.


Conformance against the running code

Graded against the meter as shipped on 2026-08-17, corpus data/pmu/trigger-scenario-corpus.json (8 scenarios), gate = two-witness placement + walk-panel kill@25 + the HD-1 lit-mass floor.

The measured result, in full:

row target gate off% status
surgeon-inline GREEN GREEN 0 MATCH
surgeon-outofcharacter CONTRADICTION GREEN 0 DOCUMENTED-HOLE
surgeon-noise REFUSED REFUSED MATCH
deploy-inline GREEN RED 42 MISMATCH
deploy-outofcharacter CONTRADICTION GREEN 12 DOCUMENTED-HOLE
desk-inline GREEN GREEN 12 MATCH
desk-outofcharacter CONTRADICTION GREEN 20 DOCUMENTED-HOLE
desk-noise REFUSED REFUSED MATCH

MATCH 4 · MISMATCH 1 · DOCUMENTED-HOLE 3 · DOMAIN-UNFITTED 0 · REGRESSION 0.

Read plainly: the contradiction class is missed 3 out of 3. All three out-of-character instructions — a step that contradicts the board it was placed against — land under the kill threshold and grade green. The three off-lane percentages are 0, 12 and 20 against a kill at 25, so this is a calibration hole, not a wiring fault: the signal is present and ordered in the right direction on two of the three rows, and the threshold is above all of them. It was found by a hostile audit on 2026-08-13, it has been carried in the open since, and it is the single largest reason nothing is permitted to reference this print as a settlement trigger today.

The noise rows both refuse, so the HD-1 floor holds and there are zero regressions.

req status evidence gap
R1 MET trigger-battery.mjs emits exactly these five statuses; attest-out/trigger-battery.json carries every row
R2 MET a REGRESSION exits 2 with the floor named
R3 PARTIAL the corpus is public and versioned; the contract string forbids fabricating a pass no intake path has been exercised by a third party, so the growth property is untested
R4 PARTIAL advisory-premium.mjs reads the battery and emits a sensorCalibration annex, or a SENSOR-UNCALIBRATED note when absent the battery is not yet wired into the release gate — today it is run deliberately, not automatically, and the currently published print carries sensorCalibrated: false
R5 PARTIAL receipts distinguish UNMEASURED from IN_LANE (HD-1 R1) the narrow representation is prose here and in the licence; no mechanism prevents a third party from selling silence as conformance
R6 ABSENT no intake address, ticket, or clock exists yet. This document creates the obligation; the mechanism is the first item of work it generates
R7 PARTIAL recomputation over sealed inputs is fully available and is what the battery does no dispute procedure is written into the settlement-reference licence
R8 ABSENT no offering document exists yet to carry the count, because no instrument references the print. The requirement binds the first one that does
R9 PARTIAL the meter is MIT and accepts pull requests corpus rows have a shape but no contribution guide; nobody outside has added one

Four requirements are not met and two of them are ABSENT. That is the honest state: the measurement machinery is real and running, and the institutional half — intake, clock, disclosure at the point of reliance — is specified here for the first time and not yet built.


How to recompute this yourself

# The battery. Eight rows, the real gate, no model anywhere in the path.
npx -y thetacog-mcp@latest trigger-battery
cat attest-out/trigger-battery.json

# The sensor-calibration annex attached to a premium. With the battery absent,
# the premium must mark itself SENSOR-UNCALIBRATED rather than imply calibration.
npx -y thetacog-mcp@latest advisory-premium

The battery's output file carries no timestamp, so two runs over the same corpus diff to zero bytes. A row that grades differently on your machine than on ours, over the same corpus and meter version, is a defect report we want.

To add a scenario: open a pull request against data/pmu/trigger-scenario-corpus.json in the meter repository with board, instruction, target and why. Rows are accepted on the strength of the reasoning in why, not on whether the gate currently passes them — a row we fail is more useful to us than a row we pass.


Known holes

  1. The contradiction class is missed 3/3 (above). The kill threshold sits at 25 and the three contradiction rows read 0, 12 and 20. Any recalibration must be justified against the false-positive side — deploy-inline already reads 42 against a target of GREEN, so lowering the threshold to catch the misses would convert a known false negative into a known false positive. That trade is exactly what this document exists to make visible before anyone prices it.

  2. The corpus is small and self-authored. Eight rows, all written by us, across three rooms. A miss battery whose adversary is also its author is the weakest possible form of the instrument. R9 is the fix and R9 is untested.

  3. DOMAIN-UNFITTED is a real limit wearing a neutral name. In an unfitted domain the trigger says nothing at all. That is a refusal rather than a pass, so it is safe — but a deployer whose domain is unfitted gets no coverage from an instrument that appears to be running.

  4. No response mechanism exists (R6, ABSENT). Today a miss report reaches a mailbox and a human. The 10-business-day clock is an obligation this document creates, not a process it describes.

  5. We cannot measure the miss rate in production. The battery measures the gate against constructed pairs where the truth is known by construction. In the field, a false negative is invisible by definition. Any claim about a field miss rate would be unfounded, and none is made here.

  6. Threshold calibration has no loss experience behind it. kill@25 is a published constant, not a fitted one. Until pools accumulate experience there is nothing to fit it to, and this document will not pretend otherwise.


How this document changes

This is v1, and it is expected to change fastest of the three hardening documents, because it is the one attached to a live defect.

Miss reports, constructed counterexamples and corpus rows: elias@thetadriven.com, or a pull request against the MIT meter. The sharpest possible contribution to this document is a scenario we fail. We would rather publish it ourselves than have it arrive in a claim.


Where this sits

The three hardening documents read in one order, because each one is only meaningful once the previous one holds.

  1. HD-1 · The lit-mass floor — first, because it decides when a reading is eligible to be a reading at all. Nothing downstream means anything about a row the floor should have refused.
  2. HD-2 · The portfolio distributional audit — second, because once rows are eligible the question becomes what the distribution of them is, and what it is not yet strong enough to support.
  3. HD-3 · The false-negative protocol (this document) — last, because it is about the readings that never appear: what a miss is, how it is recorded, and what a relying party is owed when the trigger stays quiet.

Around them: the live board carries the current print, including the semantic vega and the documented-hole count, and the machine-readable feed is at /.well-known/semantic-vega.json. The browser check verifies a receipt's signature in your own browser with no repo access and nothing installed. The vocabulary fixes what each term in these documents means, so a dispute is about the measurement rather than about the words. The T5 register lists all three documents with the first sixteen hex characters of each published file's sha256, so a stranger can fetch the served HTML, hash it, and confirm the document they are reading is the document that was registered.