Four words of this section.
Whatever decides whether something passed: the test suite, the benchmark, the acceptance report, the panel of AIs that scores. Almost nobody checks the yardstick. →
Meets, fails and undecidable. The third is not a flaw in the report: it is the report saying where the measurement cannot settle it. →
How far past the line the result must be to decide. Declared before measuring, it protects; chosen afterwards, it is cheating. →
Reaching a verdict without seeing the audited party's confidential data: through sizes, timestamps, encrypted computation or proofs. →
Producing became free. Checking did not.
A 40-page report, a technical assessment, a thousand lines of code, an evaluation score: with AI, each comes out in minutes, and costs the same whether it is right or wrong. Checking any of them still takes someone reading carefully, running the test, looking at the data.
The math is simple. If producing gets a hundred times cheaper and checking costs the same, with the same team the share that gets checked falls a hundredfold. It is nobody's carelessness; it is the ratio changing.
That is why value moved to whoever checks. But checking with another AI has the same problem: it gets cheap and loses its validity along the way. What remains is a real judge: something that does not depend on whoever produced the work, that is expensive to build once and cheap to run every time.
Three verdicts, not two.
Almost every report, dashboard and test pipeline answers yes or no. When the measurement has an error margin and the result lands near the line, that yes or no is decided by chance, and nobody finds out. Our report has a third answer, and it says exactly where the measurement stopped deciding.
Meets
The result, with its whole error margin, sits on the right side of the limit, with the margin agreed beforehand. It holds up when measured again.
Fails
Even at the best end of the error margin, the limit was exceeded. It holds up when contested: no reasonable measurement changes that.
Undecidable
The measurement is not precise enough to decide here. That is an answer, not a flaw: it says you need to measure better, or that the contract should not depend on this.
When what is being checked is the presence of a problem (a dead sensor, a defect, a fraud), the three become refuted, no alarm and undecidable. "No alarm" is never written as "approved": it is a detector's silence, worth what the detector is worth.
The yardstick is often worse than the result.
Everything below is our own measurement, with the prediction and the losing criterion recorded before running. Each number has a notebook entry with the full method.
comparisons of a program with itself in which the rule most used in CI pipelines declared a winner.
Notebook · The draw made the call 3.6%of decided verdicts changed just by changing the seed of the random draw, with no new data.
Notebook · The draw made the call 0 of 19real defects caught by tests an AI wrote by looking at the code. The test is born agreeing with the mistake.
Notebook · Green is silence 2.18truly independent votes in a panel of nine AIs judging together. They make similar mistakes.
2026 study · AI checking AI 100%of comparisons, in the stretch where the robot was moving, in which an anomaly detector found the dead sensor more normal than the live one.
Notebook · Judge what leaks AcquitsThe cheap AI may clear on its own what has no problem. Never convict on its own. The accusation goes to the expensive one.
Notebook · The cheap one acquitsFour conditions for the yardstick to count.
It sits outside whoever produced the work
The yardstick cannot be written, tuned or chosen by whoever will be judged by it. Tests generated from the code and an AI grading its own family fail here.
It rests on something the producer does not control
A calculation that can be redone, external data, a law of physics, a timestamped record. Another AI's opinion does not count: it makes similar mistakes.
It is expensive to build and cheap to run
The work is in building the yardstick once: the edge cases, the planted defects, the limit and the margin. After that it runs on every change, with no people.
It has an owner who pays when it is wrong
A slack yardstick costs nobody anything until the defect reaches the customer. If nobody pays for slackness, it goes slack. The report says who answers for it.
For use in practice.
One painful decision, one yardstick at a time.
Pick an expensive decision
Shipping a release, paying a supplier, accepting a construction job, approving a model. A decision where erring either way costs money.
Find the yardstick that decides today
It almost always exists already: a suite, a report, a dashboard, a spreadsheet. We do not start by building; we start by measuring what is there.
Measure whether it can fail
Compare the system with itself, plant defects, change the random draw, see whether the judges agree too much. Each test has its own page here.
Deliver the three-valued report
With the limit and margin written down beforehand, the blind spot declared and a receipt a third party can check. If the numbers change, we publish the correction.
What we do not do, and where this can mislead.
We do not certify by silence
"No test failed" becomes "no alarm", never "approved". If the yardstick cannot fail, the report says so before saying anything else.
We do not choose the margin afterwards
Limit and margin go into the report before the result. And the right size of the margin is still an open question: we measured three ways of deriving it, and they do not fully agree.
Undecidable can be large
In one of our studies, 81.6% of the differences a statistical test called "real" did not support a decision. An honest report can come out with a lot of undecidable, and that is information, not failure.
The cases are ours, not clients'
The measurements in this section use public and in-house data. We have not yet published a client case in this area; when there is one, it goes here in the same format.
Does this apply to your case?
Tell us in two lines what you need to decide or measure. The first conversation is to see whether measurement solves your case, and if it does not, we say so.
- Our own measurements (2026), with prediction and losing criterion recorded beforehand: the program-against-itself control and the 200-seed audit (Go's garbage collector, Apple M2); AI-generated tests for 29 changes in two public Go projects; an anomaly detector on the public LeRobot ALOHA dataset. Full methods in the notebook entries linked above.
- Nine judges, two effective votes: "Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels" (arXiv, 2026). Details in AI checking AI.
- The specimen at the top is an in-browser simulation: 20 measurements, each the average of 10 runs with Gaussian noise, 95% margin.