AI-AGENT EXCEEDANCE · holder 3f0bc417daa3 · public, no login
On 638 of this holder's AI-agent tasks, 6.1% (95% CI 4.5–8.2%) used more than four times the tokens of the median task.
A task is one prompt to the commit it produced; its tokens are summed over every model call in it. Steered tasks were dispatched one at a time, each with its scope written down before it ran, and placed against it afterwards. Nothing was blocked. Plain tasks were not steered.
n plain 638 · steered 43 · 2026-08-29 to 2026-10-09 · line = k × the median plain task (3.9M tokens); no per-task budget was declared
Steered means a dispatched task: one spec row, its scope written down before that worker ran. 11 tasks that a chat session ran itself under a standing /steer or /goal prompt are in neither arm, because their scope was declared for the session, not the task. They are counted here and in the raw reading, never dropped.
The raw reading behind every number on this page, recomputed on each request: /api/backup/fat-tail?manifest=4bc9f62e171b4b36aa6fc8547820c96ae7daae46655564cb62a1e7df9a2b85bd. The math is one open-source function (token-tail.mjs, MIT). Same ticks, same seed, same numbers.
Exceedance frequency
| line | plain | rate | 95% CI | steered | rate | 95% CI |
|---|---|---|---|---|---|---|
| 2× (7.8M) | 132/638 | 20.7% | 17.7–24.0% | 3/43 | 7.0% | 2.4–18.6% |
| 3× (11.7M) | 63/638 | 9.9% | 7.8–12.4% | 1/43 | 2.3% | 0.4–12.1% |
| 4× (15.6M) | 39/638 | 6.1% | 4.5–8.2% | 0/43 | 0.0% | 0.0–8.2% |
| 10× (39.1M) | 10/638 | 1.6% | 0.9–2.9% | 0/43 | 0.0% | 0.0–8.2% |
| 20× (78.1M) | 5/638 | 0.8% | 0.3–1.8% | 0/43 | 0.0% | 0.0–8.2% |
| 50× (195.3M) | 1/638 | 0.2% | 0.0–0.9% | 0/43 | 0.0% | 0.0–8.2% |
Each cell is a count over its denominator with a Wilson 95% interval. An exceedance is a task whose tokens went past the line (k × the median plain task). These tasks declared no budget of their own, so the median is the reference; a deployment that declares a per-task limit is read against that limit instead. The single worst task was 79× its arm's median (plain) and 3.5× (steered). That is one order statistic, not a rate.
The exceedance curve
P(task > x) for both arms, with the null band. Live: it re-reads the holder's newest backup.
+The tail of your token spendworst task 15.4M declared vs 310.4M undeclared
Live from your newest backup: how often one task uses more tokens than a given size.
scope declared before the work · n 43 no declared scope · n 638
Worst task 15.4M declared vs 310.4M undeclared · the worst 5 % of tasks carry 17 % vs 39 % of all tokens · CV 0.55 vs 2.54
Tokens per day — every task that closed that day, summed
Computed 2026-10-09 22:36:30 UTC from the tasks your backup carries (backup 4bc9f62e17, received 2026-10-09 21:00:49 UTC). Re-queried every 60 s.
+How to read the tailthe median, the far right, and what it is not sufficient for
The typical task is the median; the price of cover is set by the rare one at the far right.
Not sufficient for: same work per tick (route and task mix differ); whether the work was good (Rice).
What the p-values say, and what they don't
In aggregate, steered tasks used 47% fewer tokens than decile-matched plain tasks would have (permutation p = 0.00185). The worst steered task against the median worst of 43 plain tasks: p = 0.01575.
Where p is under 0.05, the two arms are statistically distinguishable on that measure. That does not show steering caused the difference, because tasks were not randomly assigned to each arm. Tasks are matched by complexity decile; they were not drawn from one task bank, and route and task mix differ.
Why an underwriter would track this
A token budget is a limit an agent was given before it acted, the same way a signing authority is a limit a person was given before they acted. Both are measured the same way: count how often, and how far, the actual exceeded the declared. Here no limit was declared per task, so the line is a multiple of the median task, and an exceedance is a cost event. In a deployment with tools attached, the same loop becomes an action event. Whether token exceedance predicts actions outside delegated authority is not yet measured. That is the next reading, and it will be published here with its null.
UNMEASURED today: a fitted tail shape (GPD ξ above 4×) and a per-execution-hour rate. There are 39 exceedances above 4×, too few for a stable fit.
Not sufficient for
same work per tick (route and task mix differ); whether the work was good (Rice). Nor is it the frequency of AI agents in general: these are one holder's own coding tasks.
The breakthrough, and whose it is
Whether it was good is undecidable (Rice, 1953). What it actually did cannot be reconstructed from its own account (the data processing inequality). The only remaining move is to read a record the actor did not write.
Every number on this page is read from that kind of record: each task placed against a scope written down before it ran, signed by a key the agent did not hold, and recomputable by anyone from the stored archive. The method is described in US Patent Application 19/637,714 (36 claims, filed 2 April 2026, Track One; patent pending).
For a placement lead, the condition on the slip above arrives with its measurement, its record and its method already in place. There is nothing to invent, staff or defend in front of a panel. An in-house build of the same method is a licence conversation, and a short one (/pricing).
Make this page for your own agents
1. Install the free extension:
code --install-extension thetadriven.thetacog-mcp
2. Generate your attestation row: press 🔗 Countersign (backup) (under the + on THE SECOND READER card in the ThetaCog sidebar). Your page is then at thetadriven.com/tail/<your id>, and you can switch sharing off at any time.
3. Send the link to your broker.
Free to read and cite with attribution. Paid only if you want it attested under contract (/pricing). The measurement is open source (MIT).
Elias Moosman · ThetaDriven · speaking at the Royal Institution, London (Braintree, AI Beyond Scale), 10 November 2026 · elias@thetadriven.com