Watermarking Tags the Slop. It Does Not Remove It.
Published on: August 11, 2026
Ready for your "Oh" moment?
Ready to accelerate your breakthrough? Send yourself an Un-Robocall™ • Get transcript when logged in
Send Strategic Nudge (30 seconds)Published on: August 11, 2026
Ready to accelerate your breakthrough? Send yourself an Un-Robocall™ • Get transcript when logged in
Send Strategic Nudge (30 seconds)Green in-lane · amber a little out · red drift. Every panel is a real commit, byte-identical on recompute. Tap any panel to open its shareable receipt.
Open a paragraph a model helped you write and read it again. It is grammatical, it is fluent, it is not the thing you meant, and you cannot point at the defect — which is why it costs you an hour instead of a minute. That gap is slop: prose that occupies space without carrying a claim. And here is the whole claim, meant to be swung at: watermarking tags the slop. It does not remove it. A mark counts which tokens the model picked. It tells you a machine was in the room and nothing whatsoever about whether the paragraph means what you meant — and nothing you own measures that. So stop scoring the text and go read it.
Every other ruler measures the surface too. Grammar checks grammar. Readability counts syllables, and pays you to chop a hard idea into short pieces. Not one of them asks the only question you had: did the thing I meant arrive in another mind?
So stop scoring the text and go read it. A local model walks your document three paragraphs at a time, knows nothing about what you intended, cannot fill in your gaps, and reports the exact sentence where it stopped following. Four tracks then race to repair that sentence — two fenced by a semantic lattice, two not, prompts byte-identical between the twins — and every change you accept lands as one atomic commit tagged with the track that earned it. Whether the fence writes better prose stops being an opinion and becomes a win rate a stranger recomputes on their own machine.
Three of the four rulers have no model in them at all. It runs on your hardware and nothing leaves it: npx -y thetacog-mcp@latest attest-demo. And the ruler that failed — the one measurement of ours that turned out not to work — is published beside the ones that did.
The rest of this pays that off in the order the argument was found. A label is not neutral: colouring the vocabulary green and red hands the model a second thing to optimise while it is still choosing your word, which widens the very gap the label was meant to disclose. The rereading hour is a real line-item nobody invoices. Your rejections are the calibration — a click is a measurement. Even a seven-billion-parameter model is astonishingly good at noticing what does not make sense, which is the asymmetry the whole instrument stands on: reading is cheaper than writing. One number we built did not work, and it is printed here with its failure — deliberate word salad outscored a genuine rewrite on half the sampled paragraphs, and the cause is a byte count. Three of the four rulers contain no model at all. The win-rate ledger is the first honest attempt to settle whether a semantic fence beats a raw model, and it is running now. A system that can tell when it is not doing what you asked is not in the class the transparency rules were written to catch. And the sources arrive last, raw, with the concluding left to you.
One house habit, printed up front: the thought each section is built to fire was written down before publication and committed to the open repo. The win condition is recomputation — you run the command and check a number — never agreement.
The maître d', presenting: Two Plates, One Cold — the left plate arrives with a serial number pressed into the crust, beautiful, untouched, and nobody at the table can tell you whether it tastes of anything. The right arrives cold with one bite taken out of it and a note saying exactly where the diner stopped chewing.
Before any of the argument below, run it and see what comes back. npx -y thetacog-mcp@latest attest-demo — one minute, on your machine, nothing leaves it. What returns is a placement and a receipt you can recompute, not a score you have to take on faith. If it does not do what this page says it does, you have found that out in a minute rather than in a quarter, and you should stop reading.
The label and the reader are not competing. They answer different questions, and only one of them is yours. A mark establishes provenance: a machine was here. Useful to a regulator, useful to a platform, and it settles nothing you were actually worried about at eleven at night. The reader establishes reception: this is where you lost me. Everything people want from the first — the paper that reads like a person wrote it, the draft that stops sounding generated — falls out of the second as a consequence, because the tell and the gap are the same sentence. Fix where it stopped meaning something and you have removed the thing the label was there to warn about.
WHAT THE LABEL MEASURES WHAT YOU ACTUALLY ASKED
─────────────────────── ───────────────────────
which words the model picked did the reader follow it?
↓ ↓
a keyed hash over the a paragraph read cold by a
token sequence model with no stake in it
↓ ↓
survives copy-paste "wait — what is this
dies under heavy editing referring to?"
↓ ↓
catches: unedited pasting catches: the exact sentence
catches: nothing about meaning where comprehension broke
↓ ↓
a serial number on the slop 4 rewrites, 1 commit each
↓
a ledger a stranger recomputes
Do you worry about $1.2B in AI liability?
If the property is trivial, software can check it — and why are you paying to check trivial properties? If it isn’t trivial, Rice’s theorem says nobody can. So we fixed the math.
a number we can call — or whatever you would actually ask
Who did this make you think of? We’d love to know.
Take the left column seriously first, because it is solving a real problem and solving it competently. Statistical text marking — the green-list construction published by Kirchenbauer and colleagues in 2023 and since deployed in various forms — biases the model, at each token, toward a pseudo-randomly selected half of its vocabulary. Read the finished text with the key and an improbable density of "green" choices proves which system produced it. It survives copy-paste, which is exactly the property that formatting metadata never had, and that is a genuine engineering achievement aimed at a genuine harm: nobody wants a world where you cannot tell whether the essay, the review, or the affidavit came out of a machine.
Now the sentence that column cannot reach. A provenance label answers "who wrote this." It has no opinion whatsoever on whether what was written is what you asked for. Those are different questions, and only one of them is the reason you hired the machine. Attribution is a question about the author; slop is a question about the distance between intent and artefact — and it is the second one that costs you the rereading hour, ships the paragraph that says nothing, and quietly makes your document worse than the outline it came from.
So the right column measures the thing the left one skips, and it does it the crudest possible way: it reads. A local model walks the document three paragraphs at a time, scores only the middle one, and reports its running comprehension as an internal monologue in its own words — "I have lost what 'the aperture' refers to; the previous paragraph called it a window." Below a threshold, that sentence becomes a card. Four tracks then rewrite the flagged sentence: local model alone, local model fenced by the semantic lattice, cloud model alone, cloud model fenced. The prompt is byte-identical between a fenced track and its unfenced twin — if it ever differs in any other way, the win rate measures prompt quality instead of the fence and the experiment is worth nothing. You pick one, hit Enter, and the change lands as one atomic commit tagged with the track that earned it.
Breaking a statistical watermark is the inevitable consequence of rewriting a sentence, not a feature anyone built. The keyed hash chains on the preceding tokens; change the sentence and the chain is gone. We mention it because it will be the first thing someone accuses this of being for — and because a tool whose side effect is destroying a label should say so in its own launch post rather than wait to be asked.
The whole argument, as rungs you can refuse one at a time: (1) slop is the delta between meant and shipped, not a style defect; (2) a label on the output measures authorship and not that delta; (3) the delta is measurable by reading — a cold model reports where comprehension breaks; (4) which repair is better cannot be settled by the models that produced them, so a human click is the measurement; (5) three deterministic rulers can veto a repair without any model in the loop; (6) the accumulated clicks are the first real evidence on whether a semantic fence writes better prose. Reject any one and you know precisely which sentence we disagree on — rung 6 is the one worth fighting over, and course H is where that fight is. Rungs 3 and 5 are claims about running code, and you cannot settle those by thinking: npx -y thetacog-mcp@latest attest-demo is a one-minute local install, MIT-licensed, that shows you the placement half of this working on your own machine before anyone asks you to believe anything.
The maître d', presenting: Reduction, Boiled Grey — stock reduced past the point of flavour: every ingredient still in the pot, nothing left in the bowl you could name. The kitchen labels it, correctly, as its own.
Here is the belief this whole build rests on, and it is not a technical claim. Slop is not the machine writing badly. It is the machine writing safely — the statistically averaged sentence that offends nobody, commits to nothing, and reads fine right up until you try to act on it. Every writer using these tools has had the same small horror: the paragraph is grammatical, fluent, well-organised, and it is not the thing you meant. You cannot point at a defect. That is precisely what makes it expensive.
Now watch what happens when you colour the vocabulary. At each token the model was answering one question: what most accurately continues this thought? Under a marking scheme it is answering two: what most accurately continues this thought and happens to fall in the permitted half of the dictionary right now. When the exact word for your meaning lands in the wrong half, the model is nudged — slightly, invisibly, thousands of times per document — toward a word that is nearly right. That is not a paraphrase of slop. That is the definition of it, implemented as a scheduled bias. The scheme discloses that the output is machine-made by making it very slightly more machine-like.
The people who design these schemes have a good answer, and it deserves to be said in their words: the bias is tuned small on purpose, the measured quality loss is minor, and losing a little precision is worth living in a society that can tell synthetic text from human text. Every claim in that argument is true, and we are not calling it stupid. It is the rung the rule does not reach that concerns us. Attribution regimes assume the danger is not knowing who wrote it. The danger the working writer actually faces is not knowing whether it says what they meant — and no amount of labelling touches it, because the label is computed over which words appeared, and meaning is not a property of which words appeared.
The maître d', presenting: L'Addition, Salted Twice — two bills in one folder: the printed one for the meal, and beneath it the handwritten one for the hours somebody spent tasting everything before it left the pass, salted into payroll where nobody itemises it.
You already know the first bill: tokens, seats, subscriptions, and it arrives monthly with a number on it. The second bill is the rereading, and it does not arrive at all — it is absorbed into salary, which is where costs go to become invisible. Count it once honestly. Take a document you shipped in the last month that a model helped write. Time how long you spent not writing and not editing for style, but hunting: reading your own output looking for the places where it had quietly gone generic. For most people who do this for a living the honest answer is somewhere between a third and half of the time the model was supposed to save.
The reason that hour does not shrink as models improve is the part worth sitting with. It is not priced against model quality; it is priced against the absence of a measurement. A better model produces better prose and gives you exactly the same job to do afterwards, because you still have no instrument that says here, paragraph nine, this is where it stopped tracking you. You are the instrument. You are an excellent one and you are extremely expensive, and — the part that stings — you are worst at this on your own writing, because you know what you meant and your eyes supply it.
So the check, and it costs one afternoon rather than a procurement cycle: run a cold reader over something you already published and see whether the sentences it flags are sentences you would defend. If they are, the instrument is wrong and you have learned that cheaply. If a stranger's model puts its finger on the three paragraphs you already privately suspected, you have just watched the rereading hour become a measurement — and a measurement, unlike a habit, can be delegated, scheduled, and eventually deleted from your week.
If you ship prose for a living and a model is now in your drafting loop, this page was set for you. You do not need convincing that the rereading is real work — you did it this week. What may not be priced yet is that the hour exists because of a missing instrument, not a missing skill.
The maître d', presenting: The Diner's Notes, Written in Vinegar — sharp, sour, and kept: the kitchen files what came back on the plate, not what was said politely at the table.
Here is what you bring that nobody in this industry can buy: a preference that was not generated by a language model. The central problem in evaluating machine writing is that the graders are the same class of system as the writers — a model scoring a model's prose is a system grading its own homework in a slightly different font. Every benchmark built this way inherits the bias it is trying to detect.
The escape is embarrassingly simple and it is why this tool is shaped the way it is. Four candidates for the same sentence, generated from a byte-identical prompt except for the fence, and a human picks one. That click is the measurement. Not a score, not a rubric — a revealed preference from the one party who knows what the sentence was supposed to say. Every pick appends to a local, append-only ledger: which track produced the winner, what the deterministic rulers said about it, what the original was. Twelve picks tell you nothing and the ledger will say so; a few hundred across several documents start to separate the tracks, and that is the honest shape of this evidence, stated before we have it rather than after.
Which makes early users something better than users. The instrument is built and calibrated; what does not exist yet is volume, and volume is the only thing that can answer the question the whole thing was built to ask. You are not being asked to believe the lattice writes better prose. You are being handed the apparatus that could prove it does not — running locally, with your ledger on your disk, and the aggregate submitted only if you choose to.
The maître d', presenting: The Nose Before the Hands — anyone at the pass can smell that the sauce has broken; three people in the building can fix it. The kitchen stopped asking the nose to cook.
The single most useful thing found while building this was not in the specification. A small local model is a mediocre writer and a startlingly good reader. Ask a seven-billion-parameter model to produce a sentence worth publishing and you get filler. Ask the same model to read three paragraphs and say, in its own voice, where it lost the thread, and it puts its finger on the exact clause — reliably, cheaply, and without the flattery a larger model brings.
That asymmetry is the whole architecture. Detection is a fundamentally easier problem than generation, and this is not a hopeful claim about language models — it is the ordinary structure of verification, the same reason a compiler is smaller than a programmer and a proof-checker is smaller than a mathematician. So the expensive engines are pointed at the small job (rewriting one flagged sentence) and the cheap engine is pointed at the big one (reading everything). The document gets scanned continuously by the thing that costs nothing.
Two hard-won details, since they cost real time. The aperture is three paragraphs, scoring only the middle one. One sentence in isolation has no context and the reader hallucinates confusion; a whole chapter and the monologue goes vague and managerial. Preceding paragraph, target, following paragraph — score the centre, slide down one. And the scans run four at a time: a diagnosis takes around twenty seconds per paragraph, and sequential scanning was measured at roughly five minutes to surface a single card — a cadence at which no writer will ever keep the tool open. Parallelism here is not an optimisation, it is the difference between an instrument and a demo.
The maître d', presenting: Gristle, Served Openly — the one cut the kitchen cannot chew either, brought to the table on its own plate rather than swept off the board on the way out.
The specification this was built from made a promise: if the lattice-guided tracks consistently beat the raw ones, that is mathematical proof that grounding meaning to a physical placement produces better prose. We built the measurement, then calibrated it against known inputs before trusting it, and it did not survive. node scripts/rewrite/calibrate.mjs pushes four controlled cases through the drift path — an untouched paragraph, a synonym swap, a genuine rewrite, and deliberate word salad ("Purple bicycles enumerate the quarterly banana futures of Tuesday"). The controls are perfect: an untouched paragraph reads 100 every time, a synonym swap 97 to 100, deterministically, on repeat runs.
Then the part that matters. On two of the four sampled paragraphs, the word salad scored higher than the genuine rewrite. The ranges overlap; the measurement does not discriminate quality at single-sentence scale. The cause is not mysterious and it is not a bug — it is a byte count. The comparison needs roughly 220 bytes of compressible mass to produce a stable walk, and one edited sentence inside one paragraph does not carry that mass. The signal is real and the aperture was wrong.
Where the provenance matters, since a reader is entitled to ask what kind of process ships a broken number: this was never in a release and never in a user's hands. It failed on the bench, to the guard written specifically to try to break it, before a line of it reached the console. The consequences were applied throughout rather than papered over — drift is now reported only as a displacement magnitude (identical, minimal, moderate, large), the composite score weights it at exactly zero, nothing auto-rejects on it, and the card says so on its face. The honest claim today is narrower than the specification's and it is the one being tested: not that the drift number proves prose quality, but that the fence — lattice coordinates and voice rules injected before generation — produces candidates a human picks more often. That is what the ledger in course D is recording, and it is not yet answered. Anyone telling you a semantic grounding layer is proven to write better prose is selling you a number we just told you does not work at that scale.
The maître d', presenting: The Copper Mould, Handed Over — bake it in your own kitchen from the same mould and compare crumb against crumb; the house keeps no secret oven and no secret temperature.
A rewrite tool that can silently corrupt a manuscript is worse than no tool, so the guarantees are structural rather than promised. Every paragraph and sentence carries exact character offsets into the source file, and the parser asserts that slicing the raw file at those offsets reproduces the extracted text — verified across all 1,700 sentences of a 430-paragraph chapter. An edit is a pure splice. After each write the file is re-parsed and every queued card re-anchored by exact text match; a card that cannot be re-bound is dropped rather than guessed at, because guessing is how a manuscript gets quietly wrong.
Then the rulers, and the point of this course is what is not in them. The slop detector is fifty-four named patterns, deterministic, LLM-free — the hedge stacks, the "it's not just X, it's Y" construction, the throat-clearing openers, the abstraction-density spike — and it scores a candidate against the original so a rewrite that buys clarity by adding filler is rejected on arithmetic, not taste. The readability delta pairs Flesch-Kincaid with character entropy, deliberately, because Flesch-Kincaid alone rewards chopping a hard idea into staccato fragments; if the grade improves while entropy collapses, the model is dumbing the sentence down and the pair catches it. The third is the lattice displacement, reported and weighted zero, for the reasons course F just published. Three of the four require no model to reproduce: run them yourself, on the same input, and get the same numbers.
And the default is refusal. The batch mode is dry-run — it sweeps a document, fires all four tracks, and writes a standalone report without touching a byte. Committing requires an explicit flag, and even then any candidate that increases slop is refused regardless of how well it scores elsewhere. Certainty here never means trust: the ledger is append-only, every commit carries the track that earned it, and the arithmetic is reproducible by the person with the most interest in catching us.
The maître d', presenting: The First Crumb Trail — a line of crumbs left deliberately across the floor of an argument nobody has walked yet; whoever leaves theirs first is the only one who can say where the path actually went.
State the open question as plainly as it deserves. Everyone building on top of language models is now claiming that some structuring layer — a grounding grid, a retrieval scheme, a constitution, a rulebook — makes the output better. Almost none of them can tell you how they would know. The comparisons run against benchmarks the same class of model wrote, or against a rubric a model applied, or against a demo where the guided version simply had the better prompt. The question is asked constantly and settled nowhere.
The apparatus for settling it is not exotic. Identical prompts, one variable, a human picking blind, one commit per decision, an append-only tape. That is not novel science; it is the ordinary structure of an experiment, applied to a question that has so far attracted assertions. The reason it does not exist yet is that it is unglamorous and it can embarrass whoever runs it — which, as course F demonstrates, it already has.
And here is the null standing between us and any of that, found on 12 August by reading our own instrument instead of trusting it. The scoreboard reported four tracks with success rates of 75, 17, 0 and 0 per cent — and none of those numbers were about the fence. Every track was calling the model without passing a time budget, so all four inherited a fallback ceiling instead of the one the roster actually granted: 45 seconds where 120 were budgeted, 180 where a slower model was allowed 420. The fenced arm sends a longer prompt by construction, so it spent its extra tokens against a limit roughly three times too tight and lost cards it might have won. A win rate produced that way is not a measurement of guidance. It is a measurement of which arm had less to say. One rule — how long a model may take — had been written down in two places, and the stricter one silently won.
That is the whole state of it, published rather than tidied. The apparatus works; the experiment has not started; the first honest run begins when every arm can finish. Which is the moment this stops being ours. The invitation is open in the plainest sense: run it on your own document, with your own rules as the fence, and the tape is yours whatever it says. Reporting a run back to a shared table is a thing we intend to offer and have not built — it will be opt-in when it exists, never a default, because a ledger that phones home by surprise is not a ledger anyone should trust. Until then the tape stays on your disk, which is the correct place for it.
So here is what the reader can own, and it does not require our tool to be good. Run the same experiment on your own writing, with your own rules injected as the fence, and you have a ledger nobody else has — evidence about how your documents respond to structure, produced by your own preferences rather than a vendor's benchmark. If our fence loses on your tape, the tape is still yours and the finding is still real. That is the difference between a measurement and a marketing claim: a measurement is still worth something when it comes out against the person who built it.
The maître d', presenting: Glass Kitchen, Ash and All — service in full view, every scorch mark public, the failed batch left on the counter where the room can see it. A speakeasy cannot sell a standard.
Now the turn this release actually justifies, and it is not about writing. Transparency obligations exist because a generative system is assumed to be an opaque emitter — text arrives, nobody can say whether it did what was asked, so the sane regulatory move is to stamp it. The stamp is a proxy for the missing capability. Article 50 of the EU AI Act, whose transparency obligations came into application on 2 August 2026, is that proxy written into law, and the drafting is not foolish: given systems that cannot report their own alignment to intent, marking the output is close to the only lever available.
Which is exactly why the capability, not the exemption, is the interesting object. A system that continuously measures where its output diverges from what was asked — and hands you the divergence, per sentence, with the repairs and a ledger of which repair you accepted — is a different kind of artefact from an opaque emitter. It is not making a legal argument about itself and neither are we; nothing here is a compliance opinion and no exemption is being claimed. The point is narrower and more durable: the reason a stamp was needed is a measurement gap, and gaps get closed by instruments. Rules follow instruments — the whole history of measurable industry runs that direction, and the instrument arrives first every single time.
Which is why the failed calibration is in course F instead of a drawer. A standard nobody can inspect is not a standard; it is a product with a press release. The specification, the transcripts that produced it, the delta between what was specified and what shipped, the guard that broke our own number, and the intent behind every section of this page are in a public repository — including the parts that make us look worse. The survival framing, for anyone weighing whether to wait: this instrument gets built by somebody, the first tape is already appending, and precedence, not persuasion, is what a second mover negotiates against.
The maître d', presenting: Digestif, Bitter and Raw — nothing sweet poured over it to cover the taste; the receipts folder comes with the glass and you settle the bill by checking the arithmetic yourself.
Ingredients, not conclusions — what is on the record, for you to cook with. The marking mechanism is public science: A Watermark for Large Language Models, Kirchenbauer, Geiping, Wen, Katz, Miers and Goldstein, 2023 — the green-list construction, the detection statistics, and, in the authors' own follow-up work, its degradation under paraphrase. The regulation is filed text, not our characterisation: Article 50 of Regulation (EU) 2024/1689 sets the transparency and marking obligations for generative systems, with application from 2 August 2026 — read the article itself rather than anyone's summary of it, ours included. The readability instruments are old and openly criticised: Flesch-Kincaid dates to 1975 and rewards short words and short sentences by construction, which is why it is never used here without an entropy term beside it. Our own numbers — the 100-on-a-no-op control, the two-of-four discrimination failure, the roughly 220-byte mass threshold, the twenty-second diagnosis, the 1,700-sentence offset assertion, the fifty-four slop patterns — all come out of scripts/rewrite/ and its calibration guard in this repository, and the spec-versus-reality delta document that drove the build is committed beside them. Re-run them and expect the numbers to move, which is what a live ledger does.
The companion arguments, if you want the thesis underneath the tool: the meat-mechanic post on why a machine that manufactures its own meaning at runtime is a different risk class from one whose goal is set upstream, and you never insure the catastrophe on what becomes possible once a crossing is readable rather than arguable. The book's chapters on the two halves of today's claim: how your brain grounds symbols on why meaning resolves to positions rather than to more words, and countable, not accountable on the difference between a number you can defend and a number you can only assert.
The to-do, and the bill — what is shipped and what is not, stated plainly. The engine, the four tracks, the deterministic rulers, the offset-safe parser, the console, the append-only ledger and the headless benchmark all run today out of the repository. The packaged command — the one that will let you point this at your own document without cloning anything — is the next release of npx thetacog-mcp, and it is not there yet. What you can run this minute is npx -y thetacog-mcp@latest attest-demo, after reading the source, which is what the MIT licence is for: it shows you the placement half — an action read against a declared lane, no model in the verdict — which is the same machinery the fenced tracks use. Expect several more tools to arrive on that same command; this is the one that saves you from slop, and the world-saving is a side effect we are content to leave as a side effect.
And the win condition, as declared before the first plate: this piece wins if you recompute something — the calibration, the slop patterns, the article text, the demo on your own machine — and fails if you leave merely agreeing with it. Count how many of the eleven claims above you could check without asking us for anything. The sentences each section was built to fire are committed in this repo's cook-rounds folder, timestamped before publication; compare them against what actually fired in your head. You are the only one who can run that check — which was, all along, the point.