Stickybit.← VerificationPortuguêsAI checking AI · verification · 2026
AI checking AI

Nine judges are worth two votes.

Having several AIs grade the same work looks prudent: if one gets it wrong, the others correct it. That only works if they fail each in their own way. They do not. In a 2026 study, nine judges from seven different model families were worth, together, 2.18 independent votes, and the whole panel scored almost the same as the best judge alone.

Specimen · how many judges, and how much they fail together
if each failed in their own wayvotes the panel is really worthceiling, with infinitely many judges
–independent votes, in fact
–ceiling, even with infinite judges
–what the next judge adds
–of what was paid for, delivered

The calculation is the one the study uses: with n judges that fail together at a rate φ, the panel is worth n ÷ (1 + (n − 1) × φ) independent votes. With 9 judges and φ = 0.391, that is 2.18. The ceiling, with infinitely many judges, is 1 ÷ φ.

The idea

Nine people who read the same summary.

Picture a jury of nine where everyone learned about the case from the same summary, written by someone who got one detail wrong. Nine votes, one opinion. If the summary was wrong, all nine are wrong together, and the unanimity only raises confidence in the mistake.

AI judges are similar. Models from different makers learned from very similar text, in very similar ways. When one trips on a question, the others tend to trip on the same question, in the same direction. That is what "fail together" means, and it is the number the specimen lets you move.

If they failed each in their own way, adding judges would work very well: the majority would almost always be right, because some judges' mistakes would land on different questions from others'. That is the promise behind every panel. What the study did was measure whether it holds.

EACH ROW IS A JUDGE · EACH COLUMN, A QUESTION Each fails in their own way They fail together judge's mistake column where the majority was wrong
Sketch. On top, each judge misses three questions, all different, and the majority is never wrong. Below, each judge misses the same three, and the majority misses all three.
What the study measured

72% against 71.8%.

The study had nine judges, from seven model families, assess the same cases on three public reading-comprehension test sets, where the right answer is known. Then it measured how much each pair of judges' mistakes landed together: on average, 0.391, on a scale where 0 is "each in their own way" and 1 is "always together". The most similar pair reached 0.603; the most different, 0.161.

At that rate, the nine judges are worth 2.18 independent votes on the first set (between 2.07 and 2.31), and 2.35 and 2.48 on the other two. You pay for nine opinions and get the equivalent of just over two: a ratio of 24.2%.

The effect shows up in accuracy. If the nine failed each in their own way, the majority should be right 94% of the time. It was right 72% of the time. The best judge, alone, scored 71.8%. And the cleverest ways of combining the votes recovered at most 11% of that gap.

ACCURACY ON THE SAME SET OF CASES Best judge, alone71.8% Panel of nine, measured72% Panel of nine, if each failed in their own way 94% the missing 22 points are the price of failing together
"Nine Judges, Two Effective Votes" (arXiv, 2026), on the first of the three test sets. The eight extra judges bought two tenths of a point.
The same lesson, 40 years earlier

27 programs written separately failed in the same places.

In 1986, John Knight and Nancy Leveson tested the idea behind panels: if several teams write the same program without talking to each other, one team's defects should not coincide with another's, and voting among versions covers the mistakes. They had students at two universities write 27 independent versions of the same program and ran all of them against a million test cases.

The failures coincided far more than chance would explain. Different people, with no contact, tripped on the same hard parts of the problem. Their conclusion holds for today's AI judges: diversity of origin does not guarantee diversity of error. The hard spot is hard for everyone.

What actually helps

A judge of a different nature is worth more than one more of the same kind.

The study shows that past five judges, the gain from adding another is negligible. Switching maker helps little, because the models learned from similar sources. What changes the picture is bringing in a judge that does not fail the same way because it does not work the same way:

  1. A calculation you can redo

    If the work has a number, recomputing it is a judge no language model imitates. A sum that does not add up is a refutation, with no opinion involved.

  2. Outside data

    Check the quote against the original text, the amount against the record, the address against the registry. The anchor sits outside the model that generated the work.

  3. A test that runs

    Code run against cases written from the rule, not from the code itself. It fails for reasons that have nothing to do with what the generator "thought".

  4. A person, in the right place

    Too expensive for everything, right for the accusation. It is the top rung of the ladder below.

The ladder that worked

The cheap one acquits. The expensive one checks the charge.

In our own test, with summaries of laws where we knew which were faithful and which had been tampered with, what worked was not stacking judges to vote but giving each one the decision it is good at. The full method is in the notebook entry The cheap one acquits.

Rung 1 · free
Fixed rules and a small model

They read everything. When they say "faithful", the case moves on: of the 79 tampered summaries that reached the small model, it let none through. It over-accuses, but almost never acquits wrongly.

Rung 2 · paid
A large model, only on accusations

It reviews only what was accused. Of the 84 correct accusations, it changed none; of the 13 unfair ones, it reversed 7. It worked as a court of appeal.

Result
Almost 8 times cheaper

With 5% tampered summaries in the flow, the ladder comes out almost 8 times cheaper than sending everything to the large model. Inverting the ladder, with the cheap one convicting on its own, raises false alarms and costs more.

What to do

Four rules for a panel of AIs.

  1. Count the votes the panel is worth, not the judges

    Measure how much your judges fail together, on a set of cases with known answers. With that number, the specimen's calculation says what the panel is worth, before you pay for the tenth judge.

  2. Stop at a few

    Past five judges of the same kind, the gain vanishes. The money goes further on a judge of a different nature than on a sixth model.

  3. Unanimity is not certainty

    Nine AIs agreeing can be a mistake in chorus. Treat unanimity among similar judges as a single vote, with the confidence of a single vote.

  4. Separate who acquits from who convicts

    The cheap judge decides what is clearly fine; the accusation goes up to someone expensive and different. Never let the cheap one convict on its own.

Limits

Where this can mislead.

One study, text tasks

The numbers come from one study on three reading-comprehension test sets. On other tasks, the shared-error rate may be higher or lower; what the study shows is that it is not small, and that it has to be measured.

The calculation assumes uniform overlap

The formula uses an average rate. In practice, judges may agree a lot on easy cases and split on hard ones. That is why the right measurement is on your own set of cases.

The ladder was measured on one kind of work

"Almost 8 times cheaper" holds for law summaries with 5% tampering. With another error rate or another kind of content, the math changes, and the notebook entry shows how to redo it.

A judge of another nature also fails

A recomputed number only catches arithmetic errors; outside data only catches what it records. They do not replace the panel: they cover its blind spot.

The section

To go deeper.

Talk to us

Does this apply to your case?

Tell us in two lines what you need to decide or measure. The first conversation is to see whether measurement solves your case, and if it does not, we say so.

Talk on WhatsApp algorithms@stickybit.com.br Stickybit · Porto Alegre, Brazil, since 2004

← Verification · stickybit.com.br

Sources