Stickybit.PortuguêsSection · verification · 2026
Verification · contestable verdicts

Verdicts that hold up when contested.

AI has made reports, assessments, code and evaluation scores almost free to produce. Checking them is still expensive. We audit the yardstick a decision is made with (the test suite, the benchmark, the acceptance report, the panel of AIs) and return three answers, not two: meets, fails or undecidable, with the margin declared before measuring.

Specimen · the same system measured 20 times, two reports

The contract says: respond within 200 ms. Each measurement has an error margin. Pick the system's real time and the measurement noise.

Two-valued reportcompares only the number
Three-valued reportcompares the range
meetsfailsundecidable
–side switches, two-valued report
–side switches, three-valued report
–undecidables, said out loud
–margin of each measurement

Simulation: each measurement is the average of 10 noisy runs, and the range is that average plus or minus the error margin (95%). "Switching sides" means saying meets in one measurement and fails in the next. Undecidable does not count as a switch: it is the report saying the measurement does not decide there.

First things first

Four words of this section.

Yardstick

Whatever decides whether something passed: the test suite, the benchmark, the acceptance report, the panel of AIs that scores. Almost nobody checks the yardstick. →

Three verdicts

Meets, fails and undecidable. The third is not a flaw in the report: it is the report saying where the measurement cannot settle it. →

Margin

How far past the line the result must be to decide. Declared before measuring, it protects; chosen afterwards, it is cheating. →

Checking without opening

Reaching a verdict without seeing the audited party's confidential data: through sizes, timestamps, encrypted computation or proofs. →

Why now

Producing became free. Checking did not.

A 40-page report, a technical assessment, a thousand lines of code, an evaluation score: with AI, each comes out in minutes, and costs the same whether it is right or wrong. Checking any of them still takes someone reading carefully, running the test, looking at the data.

The math is simple. If producing gets a hundred times cheaper and checking costs the same, with the same team the share that gets checked falls a hundredfold. It is nobody's carelessness; it is the ratio changing.

That is why value moved to whoever checks. But checking with another AI has the same problem: it gets cheap and loses its validity along the way. What remains is a real judge: something that does not depend on whoever produced the work, that is expensive to build once and cheap to run every time.

SAME CHECKING TEAM Before almost everything that shipped got checked After output exploded; what gets checked stayed the same size checked shipped with nobody checking
Sketch. The checking team did not shrink; output grew around it.
What we deliver

Three verdicts, not two.

Almost every report, dashboard and test pipeline answers yes or no. When the measurement has an error margin and the result lands near the line, that yes or no is decided by chance, and nobody finds out. Our report has a third answer, and it says exactly where the measurement stopped deciding.

the whole range on the good side

Meets

The result, with its whole error margin, sits on the right side of the limit, with the margin agreed beforehand. It holds up when measured again.

the whole range on the bad side

Fails

Even at the best end of the error margin, the limit was exceeded. It holds up when contested: no reasonable measurement changes that.

the range crosses the line

Undecidable

The measurement is not precise enough to decide here. That is an answer, not a flaw: it says you need to measure better, or that the contract should not depend on this.

When what is being checked is the presence of a problem (a dead sensor, a defect, a fraud), the three become refuted, no alarm and undecidable. "No alarm" is never written as "approved": it is a detector's silence, worth what the detector is worth.

What we have measured

The yardstick is often worse than the result.

Everything below is our own measurement, with the prediction and the losing criterion recorded before running. Each number has a notebook entry with the full method.

What makes a good judge

Four conditions for the yardstick to count.

  1. It sits outside whoever produced the work

    The yardstick cannot be written, tuned or chosen by whoever will be judged by it. Tests generated from the code and an AI grading its own family fail here.

  2. It rests on something the producer does not control

    A calculation that can be redone, external data, a law of physics, a timestamped record. Another AI's opinion does not count: it makes similar mistakes.

  3. It is expensive to build and cheap to run

    The work is in building the yardstick once: the edge cases, the planted defects, the limit and the margin. After that it runs on every change, with no people.

  4. It has an owner who pays when it is wrong

    A slack yardstick costs nobody anything until the defect reaches the customer. If nobody pays for slackness, it goes slack. The report says who answers for it.

The section

For use in practice.

How an engagement starts

One painful decision, one yardstick at a time.

  1. Pick an expensive decision

    Shipping a release, paying a supplier, accepting a construction job, approving a model. A decision where erring either way costs money.

  2. Find the yardstick that decides today

    It almost always exists already: a suite, a report, a dashboard, a spreadsheet. We do not start by building; we start by measuring what is there.

  3. Measure whether it can fail

    Compare the system with itself, plant defects, change the random draw, see whether the judges agree too much. Each test has its own page here.

  4. Deliver the three-valued report

    With the limit and margin written down beforehand, the blind spot declared and a receipt a third party can check. If the numbers change, we publish the correction.

Limits

What we do not do, and where this can mislead.

We do not certify by silence

"No test failed" becomes "no alarm", never "approved". If the yardstick cannot fail, the report says so before saying anything else.

We do not choose the margin afterwards

Limit and margin go into the report before the result. And the right size of the margin is still an open question: we measured three ways of deriving it, and they do not fully agree.

Undecidable can be large

In one of our studies, 81.6% of the differences a statistical test called "real" did not support a decision. An honest report can come out with a lot of undecidable, and that is information, not failure.

The cases are ours, not clients'

The measurements in this section use public and in-house data. We have not yet published a client case in this area; when there is one, it goes here in the same format.

Talk to us

Does this apply to your case?

Tell us in two lines what you need to decide or measure. The first conversation is to see whether measurement solves your case, and if it does not, we say so.

Talk on WhatsApp algorithms@stickybit.com.br Stickybit · Porto Alegre, Brazil, since 2004

← stickybit.com.br

Sources