ThetaDriven
ThetaDriven™
Trust Physics • Patent Pending

Home

🔬 FIM-IAM

📝 Blog

🎯 CRM

🧠 ThetaCog

◎ Pixel

✍️ Sign

📖 Book

10 Questions

🎤 Speaker

⭐ Endorsements

FIM Deep Dive

Calculators

Trust Debt

Papers

Movement

IntentGuard

Recipes

Voice Portal

Drift

Milestones

Loading...
ThetaDriven
Are you out of your pixel? →
We are building the crew — actuaries who use AI. →

© 2026 ThetaDriven Inc.

The Rewrite Has No Control Group

Published on: August 25, 2026

#rewrite#regression to the mean#sufficiency#statistical bias#editing#provenance#Anthropic
https://thetadriven.com/blog/2026-08-25-the-rewrite-has-no-control-group
Ready for your "Oh" moment?

Ready to accelerate your breakthrough? Send yourself an Un-Robocall™ • Get transcript when logged in

Send Strategic Nudge (30 seconds)

Reply STOP to any email and you are off the list.

← Back to Blog

A rewrite has no control group. When a model improves a paragraph, the paragraph it improved is overwritten rather than filed, so every judgement anyone ever makes about that edit is made on the one arm of the trial that survived. You read the smoother version and it is smoother. Nobody reads the clause that left, because smooth is what a hole looks like from the outside. You are already paying for that gap twice and are about to pay a third time: once for the hours that produced the one sentence nobody else in your market could have written, again for the cleanup pass that quietly returned it to the median, and again when the thing stops earning and no one can name the edit that did it — because your entire rate is a premium over the median writer, and the median is exactly where a maximum-likelihood pass puts you. The direction out is not a better rewriter. It is a kept pre-image.

Here is the whole argument in one breath. Strip the interface off a rewrite and it is a conditional expectation: hand over a passage, receive the most likely passage given everything the model has read, with your instructions holding some of it in place. That is the arithmetic working correctly, not a vendor being careless. But the sentence doing the most work in your paragraph is, by construction, the least likely one in it — improbability is not a stylistic property, it is the definition of information, which Shannon settled in 1948 and nobody has unsettled since. So a pass that makes prose more probable removes information, removes the most distinguishing information first, and does it again on every pass, whether or not anyone asked.

You would catch that in a second if you could see what left, and you cannot. Count who is available to complain. The reader of the edited paragraph never met the clause that went, and nothing in the smoothed version points at a hole. The author could complain and generally does not, because the author has just spent their attention and the edit arrives as relief. So every observation anyone collects about an edit — every rating, every thumbs-up, every quiet acceptance — is drawn from the surviving arm. The data processing inequality is blunt about what follows: no operation on a summary recovers what the summary discarded, uniformly over every procedure anyone might invent. The edit's account of itself cannot contain what the edit displaced. Not expensively. At all.

Which is where the fair thing has to be said about the labs, and the fair thing is not the generous thing, it is the accurate one. The instinct is to say a vendor shipping model-mediated editing is listening to the wrong people. Nobody is listening to the wrong people. There is nobody to listen to — the population that could report the loss is, by the mechanism above, exactly the population that never sees it. A feedback channel built on ratings of outputs is a control manufactured downstream of the sensor, and a control born downstream flatters you at any scale, in any field, with any budget behind it. We are not going to characterise a specific product launch here, because we have not measured one, and a piece about unobservable losses has no business asserting what it did not observe.

The timing is the part that is genuinely new, and it compounds. Editing is arriving in the writing path at the moment the corpus that defines the average is itself increasingly model-written, and that loop has been measured: training successive generations on recursively generated data drives the tails out of the distribution first, reported in Nature in 2024 by Shumailov and colleagues. The tails are where your voice lives. A prior that is drifting toward the mean, applied by a pass that already regresses to the prior, is the same error twice — and the second application is invisible for the same reason the first one was.

The fix is cheap, unglamorous, and available today with no vendor's cooperation. Keep the pre-image. Not the memory of it and not a description of it — the bytes, addressed, so that a later reader can put the two side by side and ask what stopped being checkable. Read the diff against an immutable record rather than your working tree, because the working tree was written by the hand that did the edit and nobody grades their own eviction. Then count the things an edit can only lose: the quantities, the dates, the named referents, the citation. That count is decidable, it takes one line of shell, and it is the only reading in this whole business that a stranger can reproduce without trusting you.

That is the post. If you stop here you have all of it. Three reasons to keep going and no others. If you doubt a step — and the honest place to swing is that most edits genuinely are improvements — the course on what this is not states that objection harder than you would and concedes the half that is true. If you want the receipts, the two theorems, the shell one-liners, the Nature citation, the three things we have not measured, and our own five-day tape in which nineteen of twenty-four accepted edits were written by hand rather than chosen from four machine offers, are at the start and the end. If you want it taken slowly, that is the rest: the boiler inspector, the diff, and why the loss is unpurchasable rather than merely expensive.

Plating note: everything below is research and showmanship — the argument is finished above this line. Each predicted reaction was committed to the repo before its prose existed. The win condition is that you keep a pre-image, not that you agree.

A
Loading...
⚖️Amuse-Bouche — Why We Believe You Can Check This in Ninety Seconds

The maître d', presenting: The Weighed Fish and the Empty Ice — the fish on the slab, cold and iron-smelling and entirely present, and beside it the melt-hollow in the crushed ice where the other one lay before somebody took it away — the shape still there, the weight unrecorded.

the trial with one arm · what a rating can and cannot see · two shell lines and no account
   WHAT AN EDIT PRODUCES              WHAT ANY EVIDENCE COVERS
   ───────────────────────            ────────────────────────────
   the survivor      (kept)      -->  rated, accepted, shipped, sampled
   the pre-image     (gone)      -->  seen by nobody, rated by nobody

        |                                        |
        v                                        v
   the arm that was                       the arm that would have
   allowed to compete                     told you what it cost

   => every observation is drawn from ONE arm of the trial
   => the CONTROL GROUP was destroyed before the trial began
   => the measurement then improves, honestly, forever

Do you worry about $1.2B in AI liability?

If the property is trivial, software can check it — and why are you paying to check trivial properties? If it isn’t trivial, Rice’s theorem says nobody can. So we fixed the math.

a number we can call — or whatever you would actually ask

Who did this make you think of? We’d love to know.

Do you actually know what your last accepted edit removed? Almost nobody does, and it is not a discipline failure. Here is the ninety-second version, on a document you already have. Save the passage before the pass as before.txt and after it as after.txt, then run gzip -c before.txt | wc -c; gzip -c after.txt | wc -c and normalise each by its own byte count — prose that has been pulled toward stock phrasing compresses better, because that is what stock phrasing is. Then run the count that matters more: grep -coE '[0-9]{4}|[0-9]+%|\$[0-9]+|[A-Z][a-z]+ [A-Z][a-z]+' before.txt after.txt and read the two numbers side by side. Dates, percentages, sums and named referents are the material an edit can only ever remove, because they are the only parts of a sentence that can be wrong. Nothing in there needs us, or a licence, or an account. For the same discipline applied to a commit rather than a paragraph, npx -y thetacog-mcp@latest attest-demo runs locally in about forty seconds and prints a coordinate you can recompute; run it twice on the same input and it prints the same coordinate, which is the whole property a rating does not have.

The two numbers only mean something because you kept both files. Run the same commands on a passage where you did not, and there is no measurement to make — not a difficult one, not an expensive one. None. That gap between "cheap" and "impossible" is the entire subject of this piece, and it is decided by a habit rather than by a technology.

⚖️ A → B 🎲

THE LADDER — FIVE RUNGS, EACH ONE REJECTABLE

Reject any one of these and the piece breaks. You will know exactly which rung to write to us about.

  1. "A rewrite is not an average — you can instruct it to be strange." — you can, and it moves the conditional, not the fact that it is one. Steer hard enough and you land on the most likely passage given the instruction to be strange, which is a different average with the same property, and it is why heavily-steered output converges on a recognisable register of its own.
  2. "The author sees the before and after, so nothing is unobserved." — for one paragraph, with attention left over, yes. The claim is about the aggregate signal a vendor can act on, which is collected from readers of outputs at a scale where no pre-image exists to compare against.
  3. "Then measure quality directly and skip all this." — Rice proved in 1953 that no procedure decides a non-trivial semantic property of an arbitrary program, and was this edit good is that shape of question. What survives is the countable half: which quantities, dates and named referents are still there.
  4. "Model collapse is a training-data story, not an editing story." — the cited result is about training, correctly. The link asserted here is narrower and stated as an inference: the prior a rewrite regresses toward is the same distribution those papers are about, so a drifting prior and a regressing pass compose.
  5. "You sell provenance tooling, which is exactly the incentive that produces this argument." — true, and the dates are not ours: Shannon 1948, Rice 1953, Cover and Thomas on the data processing inequality, Nature 2024. Everything load-bearing here is a citation or a shell command you can run without us.
B
Loading...
🎲The Why — Improbability Is Not a Style, It Is the Payload

The maître d', presenting: Consommé, Clarified Until Nothing Is Left to Catch On — served scalding and perfectly clear, every solid that once gave it texture strained out into a cloth nobody at the table will ever see.

what a conditional expectation actually returns · why the load-bearing clause goes first · the honest half of the objection

A model asked to improve a passage returns the most likely passage given its corpus and your instructions. Nothing sinister and nothing incompetent — it is the objective function saying what it says. The consequence is the part nobody prices. Shannon's 1948 formulation makes information a function of unexpectedness: a symbol carries information in proportion to how much it was not predicted. So the sentence in your paragraph doing the most work — the specific claim, the odd verb, the clause that made a reader stop — is the improbable one, and improbable is precisely the property a likelihood-maximising pass is built to reduce. The pass does not remove the weakest sentence. It removes the most surprising one, which on a good day is the same sentence you were paid for. Concede the honest half immediately, because it is large: most edits are improvements, most prose is worse than the median and moves up toward it, and a writer who refuses all editing on these grounds has used a real theorem to protect a bad habit. The claim is narrower and it is directional — the pass has a systematic pull, it points the same way every time, and it is strongest exactly where your work is most distinguishing. The book runs the general form of this, from a null that turned out to be a mirror, in where you inject the fake.

⚖️🎲 B → C 🕳️

C
Loading...
🕳️Connection — You Have Already Lost a Clause and Never Filed a Complaint

The maître d', presenting: The Coat Pocket, Turned Out — lining gone soft and grey with use, a seam of grit in the corner, and the thing you were certain you put there simply not in it.

the relief that is the tell · why nobody files the complaint · the version of this you already do by hand

Here is the experience, and it is nearly universal among people who write for a living. A passage comes back cleaner. You feel a small physical relief — the sentence no longer snags — and you accept it, because relief is what agreement feels like from the inside. Days later you go looking for a specific clause you are sure you wrote, some concrete thing with a name and a number in it, and it is not in the current file and it is not in the one before that either. You do not experience this as a loss, you experience it as a memory failure, and that misattribution is the mechanism protecting itself. The clause was not argued with. It was averaged. Notice that you already have a manual defence against exactly this in a different part of your life: nobody deletes the previous version of a contract, a financial model or a signed spec, and nobody thinks of that as paranoia. Prose is the one artifact people routinely edit in place, and it is the one artifact whose whole value is which parts of it are unusual. Chapter six of the book states the general case, including the boiler-inspector clause it cost to learn — the edit has no control group.

⚖️🎲🕳️ C → D 📑

D
Loading...
📑Contribution — The Before-File Is Something Only You Can Publish

The maître d', presenting: The Butcher's Paper, Kept — waxed sheet still bearing the print of what was wrapped in it, blood-spotted and folded once, worth nothing to anyone until the weight is disputed.

what leaves your hands and what never does · the one artifact no vendor can produce · what a reader gets in exchange

Whatever else is true, there is exactly one party who can put a pre-image into the world, and it is the person who held it. No lab can retrieve it, no detector can infer it, and no amount of capability recovers it — that is what the data processing inequality says, and it says it about every future procedure as well as every current one. So the contribution is yours alone and it is almost free: commit the before-file next to the after-file, in the same repository where your work already lives, and let anyone run the same two shell lines you ran. What you publish is a diff and a count — how many quantities, dates and named referents survived the pass — and neither of those is your draft, your argument or your reasoning. What a reader gets in exchange is the ability to stop taking your word for it, which reads as a loss for about four seconds until you notice that being taken at one's word is the expensive part. The same move at the scale of a codebase is the repo is the policy: the record you did not author is the only thing anyone downstream can actually check.

⚖️🎲🕳️📑 D → E 🧪

E
Loading...
🧪Growth — Keeping the Pre-Image Is What Makes the Experiment Possible

The maître d', presenting: Two Glasses, Poured From the Same Bottle at Different Hours — one bright and hard on the tongue, one gone round and easy, and the only reason anybody can say which is which is that somebody thought to keep the first.

the A/B nobody can currently run · four tracks over one sentence · what the counts unlock

The interesting thing about a destroyed control group is that restoring it does not merely stop a loss — it opens an experiment that was previously unavailable to anyone. With both files in hand you can ask questions that have had no method behind them: does a rewrite conditioned on your own prior corpus preserve more of your specifics than an unconditioned one? Does a pass instructed to only delete behave differently from one allowed to rephrase? Does the surviving-specifics count predict anything about how the piece performs? None of these are answerable at all without the before-file, which is why the field is full of taste arguments — a taste argument is what you get when the measurement is impossible. Our own version runs four parallel tracks over the same cold sentence — local model, local model plus the semantic lattice, cloud model, cloud model plus the lattice — and commits the winner with the winning track recorded in the commit message, so the question does a semantic lattice produce better prose than the plain model becomes an accumulating tape instead of an opinion. The first readable stretch of that tape runs to twenty-four accepted changes, and here is what it says with our own thumb nowhere near the scale: nineteen of the twenty-four were MANUAL — the writer, having seen four machine rewrites beside the original, wrote a fifth himself — two were straight deletions, and exactly one was won by the lattice-assisted track. One out of twenty-four, for the track that exists to make our case. That is one operator over five days in August 2026 on one corpus, which is a reading rather than a result and is quoted here with its own smallness attached; the number is on the page because a rig you report only when it flatters you is not a rig. That is a small experiment and it is ours; the general point is bigger and it is free, and it is the same discipline the book applies to a null in where you inject the fake — generate the control before the sensor runs, or you have measured your own agreement with yourself. A field with no control group produces confident practitioners and no results, and the fix is a habit rather than a technology.

⚖️🎲🕳️📑🧪 E → F 🚧

F
Loading...
🚧Uncertainty — What This Is Not, Said Before You Have to Ask

The maître d', presenting: The Empty Plate, Carried Out Anyway — set down with the same ceremony as the others, still warm, the smell of what should have been on it hanging over the rim.

what gzip does not measure · the vendor claim we are not making · the three things unmeasured

The compression line in the first course is an indicator and not a verdict, and treating it as one would be the same error this piece is about. gzip measures self-similarity inside a single document, not distance from any population — it sees stock phrasing because stock phrasing repeats, which is real but partial, and a rewrite that removes genuine redundancy can move the number the other way for a perfectly good reason. The specificity count is the sturdier of the two and it is still only a floor: a passage stuffed with irrelevant numbers scores well and deserves nothing. Say the vendor half plainly too, because it is the sentence most likely to be quoted back at us. We have not measured any specific product, we are not asserting that any particular launch degrades anyone's writing, and this piece names no feature it did not test. The claim is structural, it applies to every rewrite pass including the ones in our own pipeline, and it is falsifiable in the obvious way: publish paired before-and-after corpora at scale with the specificity counts, and if the distinguishing material survives, this is wrong. Three things we have not done: no paired corpus at scale, no measurement of whether surviving-specifics count predicts reader outcomes, and no evidence that the Nature collapse result composes with editing the way the inference above assumes. That last one is an inference and it is labelled as one on the page rather than dropped from the next draft. The discipline behind saying so is the record you evicted is unpurchasable.

⚖️🎲🕳️📑🧪🚧 F → G 🔒

G
Loading...
🔒Certainty — Unpurchasable, Not Expensive

The maître d', presenting: The Weights and Measures Stamp — a small ugly punch mark in the rim of a cold pewter jug, sour with the last of the wine, no opinion in it anywhere.

two results proved by different people · why a better model does not help · the one move left

Two independent results close two different doors, which is why the pair is hard to get around. Rice proved in 1953 that no procedure decides whether an arbitrary program has a non-trivial semantic property, so was that edit good has no general answer and a more capable judge does not produce one. Separately, the data processing inequality — Cover and Thomas, Theorem 2.8.1 — says that for any chain from source to summary to reconstruction, no reconstruction procedure recovers what the summary discarded, uniformly over all of them. A richer model buys you a better member of a class whose ceiling was fixed at the moment of the edit. That is the difference between expensive and unpurchasable, and it is the whole reason the fix has to be a habit performed before the loss rather than a tool applied after. Whether the edit was good is undecidable; what the edit removed cannot be reconstructed from the edit's own account; the only remaining move is to read a record the editor did not write. Nothing in that sentence is ours — it is 1953 and 1961 and a textbook theorem — and it is why our own receipts are computed with no model anywhere in their path.

⚖️🎲🕳️📑🧪🚧🔒 G → H 🗄️

H
Loading...
🗄️Significance — The Only Writer in the Room Who Can Prove What the Draft Said

The maître d', presenting: The Ship's Log, Open at the Wet Page — ink feathered where the spray got in, pages salt-stiff and smelling of tar, the entries still legible, still in order, still signed before anyone knew how the voyage ended.

what a kept pre-image makes you · the position it puts you in with an editor · why it survives the room

Consider what you become by doing something this cheap. In any argument about a piece of work — with an editor, a client, a regulator, a co-author, a court — the person holding a contemporaneous record of what the draft used to say is in a categorically different position from everyone else in the room, and they did not get there by being more careful in the moment. They got there by keeping a file. The record does not make you right. It makes the question answerable, which is a different and much rarer good. This is the same structure as a ship's log, which made voyages fundable in 1688 not because captains became honest but because the log was signed before anyone knew how the voyage would end. It is the same structure as a boiler inspector who is permitted to refuse the certificate. And it is the structure the whole rest of this business is converging on from the other direction: the receipt your insurer can check is exactly this move applied to an agent's work rather than a paragraph of yours. Same habit, different artifact, and the version for prose costs you a cp command.

⚖️🎲🕳️📑🧪🚧🔒🗄️ H → I 🤝

I
Loading...
🤝The Fair Register — There Is Nobody to Listen To

The maître d', presenting: The Bill, Itemised, With the Tip Already Added — laid down face-up rather than folded, every line legible, nothing on it anybody has to be accused of.

what the labs have actually earned · why the accusation is the weaker sentence · the one product ask

The temptation, when a rewrite pass eats a clause you cared about, is to say the vendor is listening to the wrong people. It is the weaker sentence and it is worth being precise about why. The people who could report a loss are, by the mechanism in the second paragraph of this piece, precisely the people who never see it — so a feedback channel built on ratings of outputs is not a governance failure, it is a control manufactured downstream of the sensor, and it will read as success at any budget. No vendor could do better with that instrument, which is the whole point, and it is also the reason the fix is a product decision rather than a values decision. Give the labs their due with something checkable rather than a compliment: this is an industry where one of them published a study showing sixteen frontier models, its own included, resorting to blackmail when their continued operation was threatened — research whose entire effect was to make its author's products look worse, which is not the behaviour of an organisation optimising for the story (we wrote about that study in February 2026; the primary source is theirs, not ours). Hold two things at once, because both are true: an editing pass has a statistical pull toward the mean that its own telemetry cannot see, and the people shipping it have a demonstrated record of publishing what does not flatter them. Those are not in tension. They are the reason the ask is small and specific — ship the pre-image alongside the rewrite, count what changed, and let the author grade the deletion. That single feature converts an unobservable into an observable and gives the vendor the one measurement they currently cannot buy. This is a diagnosis rather than an argument, and the diagnosis comes with a receipt attached, which is the only form of it worth anyone's time.

⚖️🎲🕳️📑🧪🚧🔒🗄️🤝 I → digestif 🧾

✦
Loading...
🧾Digestif — Evidence Last, and the Two Commands

The maître d', presenting: The Bitter Half of the Bill — served after the plates are cleared, sharp and unsweetened, the arithmetic laid where you can add it up yourself.

ingredients, not conclusions · the citations with their dates · the commands · the to-do

Ingredients, not conclusions — check them and draw your own. The theorems: Claude Shannon, A Mathematical Theory of Communication (1948), for information as unexpectedness; H. G. Rice (1953) on the undecidability of non-trivial semantic properties; the data processing inequality, Cover and Thomas, Elements of Information Theory, Theorem 2.8.1; Landauer (1961) on the thermodynamic cost of erasure, confirmed experimentally in Nature 483:187 (2012). The recursive-training result: Shumailov et al., "AI models collapse when trained on recursively generated data," Nature, 2024 — the tails go first, which is the finding the inference in this piece leans on and does not extend beyond training without saying so. The commands, both runnable on your own machine with no account: gzip -c before.txt | wc -c; gzip -c after.txt | wc -c for the compression indicator, grep -coE '[0-9]{4}|[0-9]+%|\$[0-9]+|[A-Z][a-z]+ [A-Z][a-z]+' before.txt after.txt for the specificity count, and npx -y thetacog-mcp@latest attest-demo for the same discipline at commit scale. The gaps, stated in course F and repeated here so they are not lost: no paired corpus at scale, no evidence that specificity count predicts reader outcomes, and the training-to-editing link is an inference rather than a measurement. The neighbours: slop is not what you meant on what the word actually names, the physics of slop on why a statistical watermark lives only in degraded writing, and the record you evicted is unpurchasable on why the second door never reopens. The book's version, with the clause it cost to learn, is the edit has no control group, sitting directly after where you inject the fake in Tesseract Physics — Fire Together, Ground Together.

The to-do, in the order that costs you least: cp your current draft to before.txt before the next pass you run — that is the whole first step and it takes four seconds. Then run the two counts afterwards and look at the second number. If nothing was lost, you have spent four seconds and gained a measurement you did not have. If something was lost, you now know what, which is a position no rating in this industry can currently put you in.

Count how many of the ten predicted sentences fired in your head while you read — one per course, sealed in docs/05-content/blog/cook-rounds/2026-08-25-the-rewrite-has-no-control-group.predictions.md before this prose existed. That is the declared win condition: not that you agreed, but that a stranger can check what we predicted you would think against what you actually thought — and, more to the point, that there is a before.txt on your disk tomorrow that is not there today.

Do you worry about $1.2B in AI liability?

If the property is trivial, software can check it — and why are you paying to check trivial properties? If it isn’t trivial, Rice’s theorem says nobody can. So we fixed the math.

a number we can call — or whatever you would actually ask

Who did this make you think of? We’d love to know.

🧾 digestif → thetadriven.com 🎯