Nine people who read the same summary.
Picture a jury of nine where everyone learned about the case from the same summary, written by someone who got one detail wrong. Nine votes, one opinion. If the summary was wrong, all nine are wrong together, and the unanimity only raises confidence in the mistake.
AI judges are similar. Models from different makers learned from very similar text, in very similar ways. When one trips on a question, the others tend to trip on the same question, in the same direction. That is what "fail together" means, and it is the number the specimen lets you move.
If they failed each in their own way, adding judges would work very well: the majority would almost always be right, because some judges' mistakes would land on different questions from others'. That is the promise behind every panel. What the study did was measure whether it holds.
72% against 71.8%.
The study had nine judges, from seven model families, assess the same cases on three public reading-comprehension test sets, where the right answer is known. Then it measured how much each pair of judges' mistakes landed together: on average, 0.391, on a scale where 0 is "each in their own way" and 1 is "always together". The most similar pair reached 0.603; the most different, 0.161.
At that rate, the nine judges are worth 2.18 independent votes on the first set (between 2.07 and 2.31), and 2.35 and 2.48 on the other two. You pay for nine opinions and get the equivalent of just over two: a ratio of 24.2%.
The effect shows up in accuracy. If the nine failed each in their own way, the majority should be right 94% of the time. It was right 72% of the time. The best judge, alone, scored 71.8%. And the cleverest ways of combining the votes recovered at most 11% of that gap.
27 programs written separately failed in the same places.
In 1986, John Knight and Nancy Leveson tested the idea behind panels: if several teams write the same program without talking to each other, one team's defects should not coincide with another's, and voting among versions covers the mistakes. They had students at two universities write 27 independent versions of the same program and ran all of them against a million test cases.
The failures coincided far more than chance would explain. Different people, with no contact, tripped on the same hard parts of the problem. Their conclusion holds for today's AI judges: diversity of origin does not guarantee diversity of error. The hard spot is hard for everyone.
A judge of a different nature is worth more than one more of the same kind.
The study shows that past five judges, the gain from adding another is negligible. Switching maker helps little, because the models learned from similar sources. What changes the picture is bringing in a judge that does not fail the same way because it does not work the same way:
A calculation you can redo
If the work has a number, recomputing it is a judge no language model imitates. A sum that does not add up is a refutation, with no opinion involved.
Outside data
Check the quote against the original text, the amount against the record, the address against the registry. The anchor sits outside the model that generated the work.
A test that runs
Code run against cases written from the rule, not from the code itself. It fails for reasons that have nothing to do with what the generator "thought".
A person, in the right place
Too expensive for everything, right for the accusation. It is the top rung of the ladder below.
The cheap one acquits. The expensive one checks the charge.
In our own test, with summaries of laws where we knew which were faithful and which had been tampered with, what worked was not stacking judges to vote but giving each one the decision it is good at. The full method is in the notebook entry The cheap one acquits.
They read everything. When they say "faithful", the case moves on: of the 79 tampered summaries that reached the small model, it let none through. It over-accuses, but almost never acquits wrongly.
It reviews only what was accused. Of the 84 correct accusations, it changed none; of the 13 unfair ones, it reversed 7. It worked as a court of appeal.
With 5% tampered summaries in the flow, the ladder comes out almost 8 times cheaper than sending everything to the large model. Inverting the ladder, with the cheap one convicting on its own, raises false alarms and costs more.
Four rules for a panel of AIs.
Count the votes the panel is worth, not the judges
Measure how much your judges fail together, on a set of cases with known answers. With that number, the specimen's calculation says what the panel is worth, before you pay for the tenth judge.
Stop at a few
Past five judges of the same kind, the gain vanishes. The money goes further on a judge of a different nature than on a sixth model.
Unanimity is not certainty
Nine AIs agreeing can be a mistake in chorus. Treat unanimity among similar judges as a single vote, with the confidence of a single vote.
Separate who acquits from who convicts
The cheap judge decides what is clearly fine; the accusation goes up to someone expensive and different. Never let the cheap one convict on its own.
Where this can mislead.
One study, text tasks
The numbers come from one study on three reading-comprehension test sets. On other tasks, the shared-error rate may be higher or lower; what the study shows is that it is not small, and that it has to be measured.
The calculation assumes uniform overlap
The formula uses an average rate. In practice, judges may agree a lot on easy cases and split on hard ones. That is why the right measurement is on your own set of cases.
The ladder was measured on one kind of work
"Almost 8 times cheaper" holds for law summaries with 5% tampering. With another error rate or another kind of content, the math changes, and the notebook entry shows how to redo it.
A judge of another nature also fails
A recomputed number only catches arithmetic errors; outside data only catches what it records. They do not replace the panel: they cover its blind spot.
To go deeper.
Does this apply to your case?
Tell us in two lines what you need to decide or measure. The first conversation is to see whether measurement solves your case, and if it does not, we say so.
← Verification · stickybit.com.br
- "Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels" (arXiv, 2026): average shared-error rate 0.391 (0.161 to 0.603) among nine judges from seven families; 2.18 effective votes [2.07 to 2.31] on MNLI, 2.35 on SNLI and 2.48 on AlphaNLI; panel 72% against 71.8% for the best judge; 94% predicted under independence; sophisticated aggregation recovers at most 11% of the gap; negligible gain past five judges.
- Knight, J. C. and Leveson, N. G., "An Experimental Evaluation of the Assumption of Independence in Multiversion Programming", IEEE Transactions on Software Engineering, 1986: 27 versions, two universities, a million test cases.
- The ladder: our own 2026 measurement, described in The cheap one acquits.
- The specimen uses the study's effective-votes formula: n ÷ (1 + (n − 1) × φ). Check: 9 ÷ (1 + 8 × 0.391) = 9 ÷ 4.128 = 2.18.