The attackable claim: the index currently has one constituent, that constituent is us, and cross-sectional dispersion is therefore published as null rather than as a number. An index of one is not a portfolio. We are saying so in the title paragraph of the document whose job it would be to hide it, because the alternative — publishing a dispersion figure computed over a sample of one and letting a reader assume otherwise — is the exact failure this document exists to prevent.
Fetch the print before reading our description of it:
curl -s https://thetadriven.com/.well-known/semantic-vega.json
curl -s https://thetadriven.com/hardening/index.json
This document specifies what statistical properties the receipt series must have characterised before any instrument references it, what those properties currently measure, and what is structurally missing.
A verdict about arbitrary program behaviour is not provably decidable. That is Rice, 1953: “a monitor that shares the failure domain of what it watches cannot certify it, and its certificate is a story about a story.” Nor does the watched thing hold still: “the river is the prompt. You cannot step in it twice … because the act of prompting is what moves the rocks under the foot.”
And the drift compounds, measurably:
“… walk the definer chain out and the lit-cell mass per ply runs one, twelve, a hundred and forty-four, one thousand six hundred and sixty-six, nineteen thousand one hundred and fifty-six. Roughly twelvefold per ply, compounding. Left alone, that unbounded growth does not sharpen the instrument -- it blinds it.”
That is lit mass, and the balance it runs up is Trust Debt. This instrument does not attempt the undecidable question — it takes the sidestep: “Nobody fixed the first. We fixed the second, the measurable half: not whether the work was good, but the degree of drift, and where.” An instrument that narrow is worth exactly what its stated eligibility and failure conditions are worth. A statistic with no domain, no floor and no miss class is not a measurement but a numeric decoration.
A parametric trigger is a distribution wearing a threshold. Everything an underwriter needs to know lives in the shape of the series, not in any single reading: how wide it is, how fat its tails are, how often it crosses the line, whether crossings cluster, and whether the answer is stable across the population being covered rather than an artefact of one repository's habits.
The failure this prevents is subtler than a wrong number. It is the spot-check masquerading as a characterisation: one flattering series, a handful of statistics, and an implied claim that the population behaves this way. A pool priced off that assumption is not underpriced or overpriced in a way anyone can see; it is priced off an unstated and untested homogeneity assumption, and it will find out at the tail.
So this document does two things. It fixes the method — which statistics, over which series, at which floors, published where — and it publishes the current reading with its sample size and its biases stated in the same breath.
Every reading MUST be derived from a named, hash-identified series of receipts. The receipt MUST carry the series hash, the row count, and the path of the series it read, so that a stranger can confirm they are auditing the same data.
A published print — the statistic set that constitutes a repository's Semantic Vega reading — MUST carry, at minimum:
| statistic | definition |
|---|---|
vega |
sample standard deviation of sigmaDrift over the series |
breachPctAtKill |
share of rows with offPct above the gate's own kill threshold (25) |
breachPctAt15 |
share of rows above the diagnostic line (15), for cross-reference |
kurtosisExcess |
excess kurtosis of the drift series — the tail-shape signal |
rows |
the number of readings the statistics were computed over |
A print carrying the moments without the row count is non-conformant: the sample size is part of the statistic, not metadata about it.
Fewer than 30 rows ⇒ REFUSED, with the reason attached and no premium emitted. A thin series is never priced, for the same reason a thin panel is never graded (HD-1 R5).
Where fewer than two constituents publish a print, cross-sectional dispersion MUST be published as null, never as 0. A zero implies agreement between constituents that were never compared.
The word "portfolio", and any statistic described as characterising a population, MUST NOT be used of an index with fewer than five independent constituents, where independent means: distinct owning organisations, no shared authorship of the measured artifacts, and no constituent contributing more than 50% of total readings. Below that threshold the index publishes as what it is — a single-constituent reading, or a small sample — and says so on the surface a reader lands on.
Constituency is opt-in: a repository publishes a print and asks to be indexed, and can leave by deleting a file we do not control. That property is deliberate and it is also a self-selection bias, since repositories with unflattering series can decline to publish or can withdraw. The bias MUST be disclosed wherever the index reading is published, and constituent departures MUST be recorded in the index history rather than silently reducing the count.
The series is not i.i.d. Readings are ordered in time, generated by a changing codebase under a changing declaration, and drift readings cluster. Any statistic computed as if the rows were independent draws — the standard deviation and the excess kurtosis both are — MUST be published with that caveat attached, and persistence MUST be characterised separately (below).
The remediation schedule selects its deepest band on a run of consecutive breaching readings (persistRun = 3). The distribution of breach run lengths MUST therefore be published alongside the breach rate, because one breach rate composed of isolated crossings and the same breach rate composed of long runs price differently and draw differently on the reserve.
Where a correction is known but not yet applied — the scope-gate correction for off-lane-as-scope-breadth is the current example — the published reading MUST name itself as raw and state the direction of the correction. A reading that will move must not be presented as one that will not.
Any premium derived from these statistics MUST carry its curve version and MUST NOT be represented as an actuarially filed rate. Curve v0 weights are provisional; the inputs are measured, the weights are not fitted to loss experience, and there is no loss experience yet to fit them to.
The perimeter below is not written for this document; it is the claim perimeter the meter prints for itself, in scripts/pmu/attest-demo.mjs, on every run of npx -y thetacog-mcp@latest attest-demo. It is reproduced here because a carrier-side reader needs the negative space before the statistics, not after them.
What is claimed: a decidable, bounded, distributional-semantic region match to the spec, reproducible and signed. That is small on purpose — the bound is what makes it decidable, and it is the object a state-proof seals and a parametric trigger reads. Every distributional statement in the next section inherits this perimeter; none of it survives being quoted outside it.
Graded against the meter as shipped on 2026-08-17, reading corrected 2026-08-18 after the defect in hole 10. The reading below is the print published at /.well-known/semantic-vega.json on that date; the live values are always at that URL and will differ from these as the tape grows.
The current reading, in full:
| field | value |
|---|---|
| state | ADVISORY |
| rows | 200 on the tape, 167 priced — 33 refused by the panel and excluded, never counted as zero |
| series source | data/pmu/measure-history.ndjson |
| series sha | recomputed per publish — see the print |
vega (stddev of sigmaDrift) |
8.48 |
| breach at kill (offPct > 25) | 40.1% — was published as 29.0% until 2026-08-18; see hole 10 |
| breach at diagnostic (offPct > 15) | 60.5% |
| excess kurtosis | −0.88 |
| curve | v0-advisory |
| advisory premium | 265 GPU-credit-hours per agent-month |
| sensor calibrated | true — 4 MATCH · 3 DOCUMENTED-HOLE · 1 MISMATCH · 0 REGRESSION (HD-3) |
| loss side over this series | B1 34 · B2 19 · B3 48 events over 167 priced readings · 1,372 engineer-hours · ceiling 1,900 · 822 engineer-hours per 100 readings (was 527 while the refused rows were priced as clean) |
| constituents in the index | 1 |
| cross-sectional dispersion | null |
Two of those rows are the whole story. The sensor-calibration annex is attached — and what it attaches is three documented false negatives, so the premium is a number computed over a sensor with a known and published miss class, which is a different and better thing than a number computed over a sensor nobody measured. And one constituent means every distributional statement here is a statement about a single repository's habits, ours.
These figures move with the tape. The values above were read on 2026-08-17; the live ones are always at the URL, and a reader who finds different numbers there has found the document working as intended, not a discrepancy.
That paragraph covers a value moving. It does not cover a conclusion inverting, and one has. Checked against /.well-known/semantic-vega.json on 2026-08-19:
Excess kurtosis published above as −0.88 currently reads +4.76. That is not the same statistic at a different value; it is a change in kind. −0.88 is platykurtic — thin tails, extreme readings rarer than a normal draw, the comfortable direction. +4.76 is fat-tailed: extreme readings more frequent than a normal draw, which prices worse, widens the interval around every other figure in the table, and makes hole 5 heavier rather than lighter. The document's posture on this statistic — reported because R2 requires it, not because we lean on it — was written against the thin-tailed reading and is now pointing the wrong way. The unflattering number is the live one.
sensorCalibrated currently reads false. The table above records true with three documented false negatives attached. False means the calibration annex is not attached to the live print at all: the premium there is computed over a sensor whose miss class is not published beside it, which is the condition hole 9 names and exactly the thing this document says a reader must not be asked to assume away. Until it reads true again with its annex, the live rate should be read as uncalibrated, and the sensor-calibration row above should not be carried forward to it.
The other live values, for completeness: breach at kill 7% against 40.1% here; breach at the diagnostic line 18.5% against 60.5%; vega 8.74 against 8.48; advisory rate 225 credit-hours per agent-month against 265; loss side 23 B1 · 10 B2 · 4 B3 over 200 readings against 34 · 19 · 48 over 167 priced, which puts B3 back to the smallest band and undoes the R8 finding below as a live statement. The live print carries neither pricedRows nor unmeasuredRowsExcluded, so the hole-10 fix cannot be confirmed present in it from the print alone — which is itself a finding, and it is recorded here rather than smoothed over.
None of this is retracted arithmetic; the 2026-08-17 reading is what it was, and it is left standing so the movement is visible. What is retracted is any use of this table as a current characterisation. Fetch the URL first, then read the table as the dated snapshot it is.
| req | status | evidence | gap |
|---|---|---|---|
| R1 | MET | the print and the advisory receipt both carry series.sha, series.rows, series.source, and now pricedRows + unmeasuredRowsExcluded |
— |
| R2 | MET | advisory-premium.mjs computes and publishes all five fields |
— |
| R3 | MET | CURVE.minRows = 30; below it the receipt is REFUSED with the reason and no premium |
— |
| R4 | MET | vega-index-history.ndjson carries "vegaDispersion": null at constituents: 1 |
— |
| R5 | PARTIAL | the constituent count is published on every row of the index history and on the board | the five-constituent threshold and the 50%-concentration cap are fixed here for the first time; no code enforces them, and the word "portfolio" appears in this document's own title against a sample of one — deliberately, as the disclosure |
| R6 | PARTIAL | the registry states the join and leave mechanics in its own _this_is field |
departures are not yet recorded in the history; a constituent that leaves would reduce the count with no trace |
| R7 | PARTIAL | the caveat is stated in this document | the print itself carries no non-stationarity note beside the moments |
| R8 | PARTIAL | runs are detected and banded over the real series: advisory-premium.mjs marks every reading inside a run of ≥ persistRun breaching readings as B3 (never double-counted as B2) and publishes the loss side — 34 B1, 19 B2, 48 B3 events over 167 priced readings. B3 is now the LARGEST band, which it was not while 33 refused rows sat in the series breaking up runs |
the distribution is not published: number of distinct runs, longest run, and the length histogram are all uncomputed, so the B3 count is known and its shape is not |
| R9 | MET | the print carries: "raw series — the scope-gate correction is NOT yet applied; the sealed rate will be lower" | — |
| R10 | MET | every receipt carries curve.version and the ADVISORY label; the file's own header says it is published to be criticised and superseded |
— |
Five requirements fall short; none is fully ABSENT. The one that matters most commercially is R8: 48 of 167 priced readings sit inside a persistent run, which makes the deepest and most expensive remediation band the most frequent of the three by reading count — and we cannot yet say whether that is a few long runs or many short ones. Those two worlds draw on the reserve very differently, and the arithmetic to tell them apart is a histogram we have not published. That is the top item of work this document generates.
# The published print, and the register of hardening documents beside it
curl -s https://thetadriven.com/.well-known/semantic-vega.json
curl -s https://thetadriven.com/hardening/index.json
# Recompute the same statistics from the same series, locally
npx -y thetacog-mcp@latest advisory-premium
cat attest-out/advisory-premium.json
# Compute a day-one series over any git history — including yours — and price it
# with the same curve and the same floors. Only the input moves.
npx -y thetacog-mcp@latest vega-backtest
To become a constituent: publish /.well-known/semantic-vega.json from your own repository and open a pull request adding a row to data/benchmark/vega-registry.json. Nothing is asked in return, and you leave by deleting a file we do not control. An index built by fetching rather than by asking is the only kind whose constituent list a stranger can verify.
n = 1 constituent. Every distributional claim here characterises one repository. The five-constituent threshold in R5 is the line we have drawn for ourselves, and we are on the wrong side of it, in public, at the top of the document.
The 200-row series is one codebase over one period. It is not a random sample of anything. Its vega, its breach rate and its tail shape are properties of how we work, and the correct prior is that another repository will differ substantially.
The sensor is calibrated and the calibration is bad. sensorCalibrated: true attaches three documented false negatives (HD-3). The premium is honest about the sensor it was computed over; the sensor still misses the contradiction class 3 out of 3.
Run lengths are unpublished (R8, PARTIAL). The B3 event count is measured; the shape that produced it is not.
The moments assume independence the series does not have (R7). Drift readings are serially correlated; a sample standard deviation over correlated draws understates uncertainty about the mean and the tail.
Curve v0 is unfitted (R10). The weights wBreach = 2.0, wVega = 0.10, wKurt = 0.05 are published so they can be attacked with arithmetic; they are not derived from loss experience, because there is none.
The scope-gate correction is known and unapplied (R9). The raw reading overstates; the sealed rate will be lower; the size of the gap is not yet measured.
Excess kurtosis at −0.88 is not a comfortable number to lean on at this sample size and this dependence structure. It is reported because R2 requires it, not because we think it is stable. This line read −0.26 from publication until 2026-08-19, against a table three screens above it publishing −0.88 for the same statistic on the same series — two numbers for one number, in one document, authored in one commit. Corrected to the table's value, which is the one the print carried on 2026-08-17. The live reading has since moved further and changed sign; see “What has moved since this reading”, and note that the instability this hole warns about is now demonstrated rather than hypothesised.
The calibration annex depends on a local artifact. The print reads sensorCalibrated: true only where attest-out/trigger-battery.json is present on the machine that generated it. A constituent who never runs the battery publishes an uncalibrated print, honestly marked — but the marker is the only thing standing between an uncalibrated print and a reader who does not look.
rows.map((r) => Number(r.offPct)).filter(Number.isFinite). Number(null) is 0 and 0 is finite, so all 33 rows the panel had REFUSED to grade entered the priced series as a flawless 0% off-lane reading. The published breach-at-kill read 29.0% against a priced reality of 40.1%; vega, kurtosis and the advisory rate all moved with it. This is HD-1 R1 — the requirement that UNMEASURED propagate to every consumer as a non-verdict — violated by the consumer that this document publishes, within a day of publishing both. The floor held; the cast undid it. Fixed, and ratcheted by tests/pmu-simulator/absence-is-never-a-pass.test.mjs, which prices a fixture of 35 breaching rows plus 20 unmeasured ones and fails unless the breach reads 100%. Found by an adversarial read of our own rate, which is the only reason it is in this list instead of in someone else's diligence memo.This is v1, corrected in place on 2026-08-19: hole 8 was publishing −0.26 for a statistic the table published as −0.88, and the live print had inverted the sign of that statistic and withdrawn the sensor calibration since the reading was taken. Both are recorded above rather than quietly overwritten, on the same principle as hole 10 — a correction that leaves no trace is indistinguishable from a number that was always right.
Attacks on the method, on the statistic set, or on the thresholds in R5: elias@thetadriven.com. An actuary or model-validation quant who takes this document apart is doing the work the T5 gate exists to invite, and the corrections publish here with credit.
Three hardening documents, and they are ordered because each one is a precondition for the next. Read them this way:
sensorCalibrated field in the table above is HD-3's annex reaching into this document.Around them: the live board at /benchmark is where these statistics are published as they move, self-metric first. The browser check at /verify-receipt recomputes a single receipt with nothing installed, which is the smallest unit of this whole apparatus a stranger can falsify. The vocabulary is fixed at /definitions — every load-bearing term above links into it, so a disagreement about a word can be settled without settling it in prose. And the register itself is /hardening/index.json: fetch each document, sha256 the served file, compare the first sixteen hex characters to the row. If a hash there stops matching its artifact, the build is already failing.