Proof · measure every change

We don’t say trust us. We hand you receipts.

every paid fix records a before · a verdict at day 28 · and it is allowed to say the fix failed
  • Every paid fix records what the page was doing first — captured at the moment it publishes, from your own Search Console. Not reconstructed afterwards, when it would be too late to be honest.
  • The verdict can be “it didn’t work” — improved, holding, regressed, or too early. A system that only reports success is not measuring; it is advertising.
  • Small pages don’t get a verdict at all — on a page with a click a day, a 28-day comparison cannot tell a real effect from ordinary variation. It says so instead of guessing.
  • “No data” and “zero traffic” are different claims — a page with no history at the moment of the fix records no baseline and waits, rather than being written down as a zero.
  • Nothing here has produced a verdict yet — the fixes are recorded and measuring, and not one of them has come back. When they do, they will say what they say.
what we found when we tested our own method

The acid test

19–20 july 2026 · production search console · read-only

We asked one question of our own method: on a page of a given size, how large does a real effect have to be before a before-and-after window can detect it at all?

Page sizeSmallest effect detectableReading
0–1 clicks a day~74%a verdict here is noise
1–3 clicks a day~34%the tier boundary
3–8 clicks a day~25%measurable
8+ clicks a day~20%measurable, extrapolated

On pages below that floor, 76% of our own naive day-28 verdicts were false positives — the same verdict a placebo would have produced.

three bands measured, the fourth extrapolated beyond the range the test covered · this is the number that changed the product · naive before-and-after maths no longer decides anything here

What we did about it

A verdict now has to beat a placebo.

If you fix a page and its traffic rises, the honest question isn’t “did it go up”. It’s “did it go up more than it would have anyway” — because the whole site moves together, seasons turn, and Google updates land on days nobody chose.

  1. 01

    Build a counterfactual

    Fit on the eight weeks before the fix, using your untouched pages as controls, to model what this page would have done without it.

  2. 02

    Test against placebos

    Re-run the same comparison on dates when nothing happened. If the method finds an effect there, it would have found one here too.

    Automatic check
  3. 03

    Judge, or decline

    Verified, no effect, or negative — and where there isn’t enough history, or the page is too small, neither.

  4. 04

    Write it in plain words

    The numbers decide the verdict; the model only explains it. It cannot overturn the maths.

  • at least eight untouched control pages are needed to build the null — with fewer, the engine reports that it cannot judge rather than judging badly
  • a page we have recently fixed is never used as a control — measured: a treated control absorbs the effect being measured
Four answers

Two of them are bad news.

A receipt resolves to one of four things, and only one of them is the outcome anybody wants. That is the design working, not the design failing — a measurement instrument that can only return “improved” is not an instrument.

And where a fix is reversible and you asked for it in advance, a severe regression at day 28 can trigger the undo automatically. The measurement is wired to the remedy.

VerdictWhat it meansWhat follows
Improvedbeat its own counterfactualGoodkeep it
Holdingno detectable effect either wayNeutralwait
Regressedmeasurably worse than doing nothingBadundo available
Too early28 days isn’t long enough to tell yetUnknownhold to day 60
Refusals

Three times it declines to give you a number.

Most of the engineering here is in the cases where an honest answer is unavailable. Each of these could have been a verdict; each is deliberately not.

  • The page is too small

    Below roughly one to three clicks a day, a 28-day window cannot separate a real change from ordinary variation. The receipt is marked unmeasurable and the fix’s evidence is pooled with your other fixes instead of standing alone.

    The fix still ships. It just doesn’t get a lone verdict it can’t support.

  • There isn’t enough history

    The engine fits on the eight weeks before the fix, and needs more history still to find placebo dates to test itself against. A recently connected store has months of data, not years, and where that run doesn’t exist the receipt reports insufficient data — measuring, not concluding.

    An honest “not yet” beats a confident number built on four weeks.

  • There’s no baseline to compare

    A page with no Search Console rollup at the moment the fix publishes records no baseline at all and is retried later. It is never written down as a zero.

    “We have no data” and “you had no traffic” are different sentences.

Patience, encoded

Day 28 is a check-up, not a sentence.

The instruction a verdict is written under says that search effects settle over six to twelve weeks and that a substantial share of recoveries land after the first month. Judging everything at day 28 and acting on it would throw away work that was about to succeed — so the default at day 28 is to hold, and rollback is reserved for regressions severe enough to be worth not waiting on.

That expectation is written into the instruction itself, not left to the model’s temperament on the day.

What day 28 is allowed to conclude

Clearly better
improved
No signal either way
hold to day 60
Slightly down
hold to day 60
Severely down
consider rolling it back

effects settle over 6–12 weeks · the instruction says so, so the verdict does too

Where this stands today

Built, running, and not yet proven.

We would rather tell you this than let you find out. Fixes are shipping and recording baselines across production stores right now, and the receipts are sitting in the measuring state where they belong. No day-28 verdict has been returned yet.

When those verdicts do land, some of them will say the fix did nothing and some will say it made things worse — and that is exactly what this page will show, because a proof system you only publish when it flatters you isn’t one.

  • The measurement runs whether or not we like the answer — a daily sweep, no manual step
  • Verdicts land whole — the numbers and the explanation together, or neither
  • We will publish what comes back — including the negative ones

As of 4 August 2026

Fixes recorded with a baseline
61
Across production stores
3
Still measuring
55
Retired unjudged, the page re-fixed
6
Verdicts returned so far
0

counted by hand on 4 August 2026 · this counter is the one to check us on

Before you ask

The honest answers.

  • Why show me a feature with no results yet?

    Because the alternative is showing you results we don’t have. The measurement apparatus is real, it’s running against production stores, and you can read exactly how it decides. When the verdicts arrive we’ll show those too, including the ones that say the fix did nothing.

  • Isn’t 76% false positives an argument against you?

    It’s an argument against the method almost everyone uses, including us until we tested it — a plain before-and-after on a page too small to carry one. Publishing that number is the reason to believe the ones that come after it.

  • What if my store is too small to measure?

    Then per-page verdicts won’t be honest and you won’t be given them. Your fixes are pooled instead, so the evidence comes from the whole programme rather than one noisy page.

  • Does a bad verdict cost me anything?

    The fix was already paid for; the measurement is free. If a fix regressed and it’s reversible, you can undo it — and where you asked for that in advance, a severe regression can undo itself.

  • Who writes the verdict?

    The numbers do. A model writes the sentence explaining it and cannot change the judgment — and if that call fails, the receipt stays measuring rather than shipping numbers without their explanation.

Anyone can show you a graph
that goes up.

61 fixes recorded · 55 still measuring · 0 verdicts so far · counted by hand on 4 August 2026