Stickybit.← VerificationPortuguêsAudit · verification · 2026
Yardstick audit

Can the yardstick fail?

Every technical decision has a yardstick: the test suite that ships the release, the benchmark that says it got faster, the AI panel that grades, the report that accepts the equipment. Almost nobody checks the yardstick. Here are five tests that work for any of them, what each one has already caught in our measurements, and how to do it in practice. The card alongside tells you what your yardstick can claim today.

Specimen · your yardstick's card

Answer based on what was actually done, not what is planned. "Don't know" counts as "no": a yardstick nobody has tested gets no benefit of the doubt. Nothing you mark leaves your browser.

Why audit the yardstick

The result inherits the yardstick's defects.

A thermometer that always reads 36.5 degrees will never flag a fever. The problem does not show in any reading, because every reading looks normal. It only shows when someone puts the thermometer in hot water and sees it does not rise.

Yardsticks for software, data and AI are the same. A suite that cannot fail stays green. A benchmark that mistakes noise for improvement invents winners. A panel of AIs that fail alike agrees with itself. None of these defects shows up by looking at the result. All of them show up when you test the yardstick on purpose.

The five tests below are that "hot water". They are cheap next to what they protect: none requires a new measurement of the system, only using the yardstick in a way it is not usually used.

THE SAME THERMOMETER, THREE SITUATIONS 36.5no fever 36.5with a 39 fever 36.5in water at 60 degrees only the deliberate test, on the right, shows it does not rise
Sketch. The first two readings look normal; the defect is in the yardstick, not the patient.
1Test one

Compare the system with itself.

Run the yardstick twice on the same thing, changing nothing. The right answer is "nothing changed". If the yardstick declares a difference, it is reading the machine's noise as news, and it will do the same when there is a real change.

We did this with Go's garbage collector: 198 comparisons of a program with itself. The rule most used in pipelines, "is the new average more than 5% off the old one?", declared a winner in 33 of 198, 16.7%. The rule that checks the whole range against the limit declared zero.

  1. Pick a case you know did not change: the same code, the same equipment, the same answer.
  2. Run it through the yardstick the normal way, twice or more, as if it were a real comparison.
  3. Count how often it flags a difference. That is the yardstick's false-alarm rate; anything below it is noise.
198 IDENTICAL PAIRS · INVENTED WINNERS Average against the limit33 Range against the limit0 same program, same machine, 5% limit
Our measurement, Apple M2, Go 1.26.5. Full method in The draw made the call.
2Test two

Plant defects on purpose.

Make a copy of what is being judged with a small, known mistake: flip a sign in the code, make the system 10% slower, put a wrong answer among the right ones, measure an item already known to be out of spec. Run it through the yardstick. If it stays green, it cannot see that kind of mistake.

In the specimen of the entry Green is silence, the suite checked by default catches 5 of 8 planted defects and lets the real defect through. On the real changes we tried to audit, off-the-shelf defect-planting tools almost never reached the lines the change had touched. That is why it pays to plant by hand, where the risk is.

  1. List three or four mistakes that would really hurt in your case: the wrong threshold, the inverted sum, the rounding.
  2. Make one copy per mistake and run the yardstick on all of them, without telling the yardstick anything.
  3. The share caught is the yardstick's grade. The ones that passed green are its blind spot, and go into the report in writing.
EIGHT COPIES, ONE MISTAKE EACH caught by the yardstickpassed green 5 of 8is the yardstick's grade, not the system's.The three outside are its blind spot.
The result of the specimen in the entry Green is silence, with the default suite.
3Test three

Change the draw and the order.

Many yardsticks hide a piece of chance: the seed of the random draw that estimates the error margin, the order the tests run in, the order the answers reach whoever grades them, the day the report was measured. Change only that, no data at all, and see whether the verdict moves.

We redid 198 comparisons with 200 different seeds. 3.6% of the decided verdicts flipped with no new data. And the study's main number, 81.6% of differences that did not support a decision, ranged from 78.9% to 84.2%. The number survives, but is now quoted with its range.

  1. Find out what, besides the data, goes into the calculation: seed, order, who measures, when.
  2. Redo the verdict changing only that, a few dozen times. It costs minutes of computer time.
  3. Every verdict that changes is being decided by the coin. Publish the number with its range, and treat those cases as undecidable.
LIMIT seed 1meets seed 2undecidable seed 3meets seed 4undecidable same data; only the seed changed
Sketch. Full measurement in The draw made the call.
4Test four

See whether the judges fail together.

Three witnesses who read the same newspaper are not three witnesses. When a yardstick combines several judges (several tests, several people, several AIs), what matters is not how many there are, but how differently they fail.

A 2026 study measured nine AI models from seven families judging together. Their errors had an average pairwise correlation of 0.391. In practice, the nine were worth 2.18 independent votes, and the panel was right about as often as the best judge alone. Going beyond five judges barely changes that.

  1. Set aside a batch of cases where you know the right answer.
  2. For each pair of judges, check how often both fail on the same case.
  3. If they fail together, count them as one. The page AI checking AI does this math for you.
NINE JUDGES WHO FAIL ALIKE 2.18 actual votes average pairwise error correlation: 0.391
Data from "Nine Judges, Two Effective Votes" (arXiv, 2026). The interactive math is in AI checking AI.
5Test five

Who can touch the yardstick?

An exam is worth little if the student can rewrite the answer key. Ask who wrote the yardstick, who can change it and whether anyone sees when it changes. A test written from the code itself, a report produced by the supplier, an AI grading a model from its own family: in all of them, whoever is judged is holding the yardstick.

We measured today's most common case. An AI wrote tests by looking at the code of 19 changes we knew contained a defect. The tests caught 0 of 19: they recorded what the code does, mistake included. On the most used leaderboard for AI coding agents, the problem was filed in its own repository: in some setups, the agent could edit the tests that would judge it.

  1. Write down who can change each piece of the yardstick: tests, limits, reference data, the judge's instructions.
  2. Separate: whoever is judged does not change the yardstick for their own change without review by someone else.
  3. Treat any change that edits the code and its test together as an automatic warning.
19 CHANGES WITH A KNOWN DEFECT 0 of 19 caught by tests writtenby looking at the code itself
Our measurement, August 2026, two public Go projects. Details in Green is silence.
Limits

What the card does not catch.

It depends on who answers

The card is a self-assessment. An optimistic answer produces an optimistic card. In a real audit, every "yes" comes with the test that was run and its result.

Passing all five does not guarantee the right measure

A yardstick can fail when it should, be stable and independent, and still measure the wrong thing: speed when the customer feels freezes, a style grade when accuracy was what mattered.

A planted defect is not a real defect

Planted mistakes are the ones someone thought of. The ones that really escape tend to live where nobody thought, and where it is not even easy to test.

The numbers come from our own cases

33 of 198, 3.6% and 0 of 19 come from measurements on open-source code and our own machine. They show the defects exist; their size in your case only shows up by measuring your case.

Keep going

To use in practice.

Talk to us

Does this apply to your case?

Tell us in two lines what you need to decide or measure. The first conversation is to see whether measurement solves your case, and if it does not, we say so.

Talk on WhatsApp algorithms@stickybit.com.br Stickybit · Porto Alegre, Brazil, since 2004

← Verification · stickybit.com.br

Sources