Four words for this page.
The size of difference from which someone changes what they do. Here, 5%. It has to be declared before seeing the result.
Redoing the calculation many times, picking at random among the measurements that already exist, to see how much the result wobbles. The technical name is bootstrap.
The number that starts the draw. Same seed, same draw, same result. It looks like solidity; it is just repetition.
The share of verdicts that changes when only the seed changes. Those were not decided by the measurement, but by chance.
Comparing averages invents winners.
We measured the switch of Go's garbage collector to a new version: 198 comparisons across programs and metrics, on one machine. Before looking at the result, we ran the test almost nobody runs: comparing the program with itself. The right answer is "nothing changed", every time.
The rule most used in CI pipelines, "is the new average more than 5% off the old one?", declared a winner in 33 of 198 comparisons, 16.7%, between a program and itself. That is the machine's normal noise becoming news.
The rule we use instead looks at the range the difference probably lies in and compares that range with the limit. If the whole range is past the limit, it decides. If it is entirely inside, it decides "same". If it crosses the line, it says undecidable. In the same test, it invented zero winners.
And where does the range come from?
The range is computed by drawing. Take the 12 runs of each version, draw 12 with replacement, compute the difference, and repeat that hundreds of times. The drawn differences show how much the result wobbles. The range is the middle stretch of them.
Like every draw on a computer, this one starts from a seed. Our tool always used the same one, 12345. It is common practice and for a good reason: running again gives the same number, and nobody gets confused.
But repeating the same draw is not the same as the draw not mattering. In comparisons where the end of the range touches the limit, one seed puts that end on one side of the line, another seed puts it on the other. The verdict flips. The data is the same; only the coin changed.
200 seeds, and 3.6% of the decided verdicts flipped.
We took the 198 comparisons already published and redid only the draw, with 200 different seeds. No new data, no new runs. Before running, we recorded the prediction: among decided verdicts, the coin share is above zero and below 5%. Small enough not to topple anything, big enough to exist.
| Published verdict | Comparisons | Changed under some seed | Coin share |
|---|---|---|---|
| Decided (faster, slower or same) | 110 | 4 | 3.6% |
| Undecidable | 88 | 4 | 4.5% |
| All | 198 | 8 | 4.0% |
The prediction held: 3.6% is inside the range recorded beforehand. Our measurement, August 2026 (the seed audit): 198 program-metric pairs from Go's garbage collector, Apple M2, Go 1.26.5, each recomputed with 200 seeds.
We also redid, with the 200 seeds, the most quoted number of that study: of the comparisons where a statistical test says "there is a difference", 81.6% do not support a decision against the 5% limit. It ranged from 78.9% to 84.2%, and landed exactly on 81.6% in 170 of the 200 seeds. The number survives. But now that the variation is measured, quoting only the central digit would be picking the prettiest version. The honest form is "81.6% (78.9% to 84.2% across draws)".
And one case deserves a name. In a comparison of reading HTTP requests, counting collector cycles, the published verdict was undecidable. Across the 200 seeds, 52% say "same". It is the only case where the published verdict is the minority side: it was literally heads or tails, and the coin landed on the side that got printed. It changes no conclusion, because both verdicts say "you cannot act on this". But it is the clean example of the phenomenon.
The coin lands where the decision matters.
We ran the same audit on two more of our tools, and each needed a different test, because in each one something different wobbles the boundary. In one it is the seed. In another, the error margin the manufacturer declares in the spec. In another, which stretches of data were set aside for calibration.
The result ranged from 0.18% to 10.5%, and what explains the difference is not the tool: it is how close the cases live to the limit. The antennas sat, at the median, 7 dB from the line; almost nothing flips. Benchmarks pile up near the action limit, because that is exactly why someone wanted to compare them.
The consequence is uncomfortable. The coin share is largest exactly where the decision is hard, because that is where cases accumulate. A verdict that never touches the boundary was not deciding anything that needed measuring. In the specimen above, it is the difference between options a and b.
Requiring a margin works. Picking the margin afterwards does not.
The obvious fix is to require a margin: decide only if the result cleared the line with room to spare, and call everything close undecidable. In sensor attestation, requiring double wiped out the coin zone, in a tool built before that rule existed. That is a good sign the rule was not tailored to one case.
The trap is picking the size of the margin by looking at the table. A margin of 1.5 times changed nothing; 2 times wiped it out. Taking 2 because it gave the pretty result is the same mistake under another name: a limit chosen after seeing the verdict. We tried three ways of arriving at the value. From the instrument's own variation, it came out at 1.4. Looking for where the coin disappears, 2.0. The third could not point to any value. The margin has to exist; its right value is still an open question.
Four rules for a verdict that survives being redone.
Ask what, besides the data, moves the line
The seed of the draw, the error margin of a spec, which samples were left out. Each tool has its own. If nobody can say which it is, nobody measured it.
Draw again before publishing
Change only what is not data and see which verdicts flip. It costs minutes of computer time, with no new measurement.
Quote the number with its range
"81.6%" and "81.6% (78.9% to 84.2%)" say different things. The second tells how much of the number is measurement and how much is the draw.
Declare the limit and the margin first
The action limit and the required margin go into the report before the result. Chosen afterwards, they are just another way of reaching the answer you wanted.
Where this could be wrong.
One machine, one pair of versions
The 3.6% comes from one Apple M2, one Go version and one collector switch. The mechanism is general; the size of the share depends on where the cases live, and that changes from study to study.
200 seeds are also a sample
A comparison that flips in one seed out of a thousand may not have shown up. The measured share is a floor, not the exact count.
The antennas were not real
The antennas' 0.18% was measured on a synthetic sweep generated by the tool itself. It describes the method, not an installed antenna.
The margin's value is open
We know requiring a margin works and that picking it from the result is cheating. We do not yet know how to derive the right value from outside, for any tool.
← Research notebook · Certified telemetry · stickybit.com.br
- Our measurements, 2026: Go's new garbage collector against the old one (Apple M2, 8 GB, Go 1.26.5), 198 program-metric pairs, the program-against-itself control and the 200-seed audit, with the prediction recorded before running. The same audit on the antenna report and on sensor attestation, and the test of three ways to derive the margin.
- The draw used to estimate the range: Efron, "Bootstrap methods: another look at the jackknife" (The Annals of Statistics, 1979).
- The specimen is our own simulation: 48 comparisons with 12 runs of each version, generated in the browser with fixed data; the range is the stretch between 5% and 95% of 400 drawn differences; only the seed changes with each click.