Half faithful, half tampered, and we knew which was which.
We took articles of Brazilian law and asked AI models to summarize them. For half, we asked for a faithful summary. For the other half, we asked them to change the legal meaning without saying so: swap "forbidden" for "allowed", drop an exception, change who holds the right.
That gives a rare advantage: the right answer does not depend on anyone's opinion, not even another AI's. We knew which summary was which because we were the ones who asked.
Then we set three checkers to say "faithful" or "tampered": fixed rules that compare keywords, a small model that runs for free on our computer, and a big model, one of the best on the market, paid per use.
The expensive one is not less wrong than the cheap one.
A checker can be wrong in two ways: let slip a tampered summary, or wrongly accuse a faithful one. The first mistake is the dangerous one; the second is the tiring one, because someone will have to check again.
The fixed rules almost never accuse wrongly, but they let almost half slip through. The small model, added to the rules, catches 97 of every 100 tampered summaries. The big one catches all of them.
The surprise is in the other column: the big one accuses wrongly just as often as the small one, about 10 of every 100 faithful summaries. Paying more did not buy precision by itself. So the question stops being "which checker is better" and becomes "who decides what".
It may clear alone. Never condemn alone.
The small model has a very useful way of being wrong: it accuses too much, but it almost never clears what it shouldn't. Of the 79 tampered summaries that reached it, it cleared none. So when it says "faithful", you can trust it and move on. That settles most of the traffic for free, since most summaries are faithful.
When it says "tampered", the case goes up to the big model. And here is the number that taught us the most: of the 84 correct accusations, the big one changed none. Of the 13 unfair accusations, it overturned 7.
In other words: the big one found nothing the small one had missed. It worked as a court of appeal, and all of its value was in overturning unfair accusations.
Flip the rule, with the cheap one condemning alone and the expensive one reviewing only what it cleared, and false alarms rise to 17 of every 100 faithful summaries. That is worse than sending everything to the expensive one, and it costs more, because almost everything the cheap one does is clear.
Does the rule beat a random draw?
Every rule that sends some of the cases to the expensive checker spends money. So you have to ask: if, instead of the rule, we drew at random the same number of cases for the expensive one, would it come out the same?
We ran the draw 4,000 times. The rule "the cheap one clears, the big one checks the charge" catches every tampered summary, which the draw does not. But its false alarm rate, 7.7%, falls inside the range of the draw. On that measure it does not beat chance: the reduction we measured earlier cannot be credited to the rule.
The rule that beats the draw on both measures is a different one: the big one steps in only when the two free checkers disagree. It catches all of them, wrongly accuses only 3.6%, and calls the big one in 29% of cases. It is the rule selected in the specimen up top.
The dumb rule beat the AI's score.
We took the same question to code review: what deserves human eyes first? We asked a small model for a risk score on 200 real changes, knowing which ones later needed a fix. The score did about as well as a coin toss. Worse: it reacted to there being code on screen, not to which code it was.
The rule that worked was the simplest one possible: the size of the change. Across 3,358 changes in four projects, holding back those over 400 lines stops 19% of them and gathers 36% of the ones that later turned out defective. That is almost twice the concentration you would expect by chance.
There is no magic cutoff: the tighter it is, the more you catch and the more you hold back, on a curve with no knee. Where to cut is a business decision, and now it can be made with the price on the table.
In 2026 Meta published a review system with the same design, over more than 535 thousand changes: the automatic checker approves low-risk cases on its own and never blocks on its own; the rest goes to a person. It is our rule, at another scale.
Four rules for setting up the check.
Decide what each checker may do alone
Choosing the order is not enough. The cheap one may clear alone; condemning takes confirmation. Pointed the wrong way, the same ladder gets worse and more expensive.
Use the expensive one as an appeal, not a detective
A second AI answering the same question only overturns accusations. To find what the cheap one misses, the step above needs other evidence: run it, reproduce it, check the source.
Run the random draw before celebrating
Every routing rule has to beat sending the same number of cases at random. Our first recommendation did not.
Price it by the error rate of your own flow
With 5% tampered summaries, the best rule comes out almost 8 times cheaper than sending everything to the big one. With half tampered, the advantage shrinks. Move the slider in the specimen to see.
Accepting as faithful a summary that was tampered with. This is the dangerous mistake.
Calling a faithful summary tampered. Not dangerous, but it costs: someone will check again, and a checker that cries wolf too often stops being heard.
How many tampered summaries there are in the batch. It changes everything in the cost math, and it is rarely the number you imagine.
Where this could be wrong.
383 cases is not many
The three best rules gave 2.1%, 3.1% and 3.6% false alarms. With this sample, those three cannot be told apart. Separating them takes a larger set, which is the next step.
A single subject
Every case is a summary of Brazilian law. On code, contracts or medical reports, each checker's error pattern may be different.
Time was left out of the math
The small model is free, but it takes about 4 seconds per case on our computer. On a screen where someone is waiting for the answer, the savings may flip.
The "industry proof" is observational
Meta's result (1 incident per 50 approvals) compares cases the system chose with cases it did not touch. To know how much human review still adds, you would have to randomize.
← Research notebook · stickybit.com.br
- Our own experiments, Aug–Sep 2026: 383 summaries (189 tampered, 194 faithful) of provisions of Brazilian law; claude-opus-5 and claude-haiku-4-5 models as generators and paid checker, a local 8B model as free checker. Pre-registered; the policy runs cost US$ 0.00 (reusing verdicts already paid for).
- Code review: 3,358 changes from four open-source projects, with each defect attributed through the change that later fixed it.
- Meta, "Automating Low-Risk Code Review at Meta: RADAR", 2026 · Google Mantis (the claim "light models to classify, advanced ones for deep analysis").