Stickybit.← NotebookPortuguêsMeasurement · deciding under uncertainty · Sep 30, 2026 ·
Measurement

The draw made the call.

To say whether a new version of a program got faster, measurement tools draw at random: they redo the calculation many times, picking at random among the measurements they already have, to size up the doubt. The draw has a fixed seed, so the result repeats and looks solid. We changed the seed 200 times: 3.6% of the decided verdicts flipped.

Specimen · 48 comparisons between the new version and the old
fasterslowersameundecidableverdict differs from the published one

seed: 12345 · draws so far: 0

–decided in the published run
–decided ones that have flipped
–undecidable ones that have flipped
–published verdict in the minority

A simulation to show the mechanism, not our data (that comes further down). Each comparison has 12 runs of each version. Each bar is the range the difference probably lies in, computed by drawing at random. The dashed lines are the action limit: 5% faster or slower. The data never changes; only the seed of the draw.

Before the measurement

Four words for this page.

Action limit

The size of difference from which someone changes what they do. Here, 5%. It has to be declared before seeing the result.

The draw

Redoing the calculation many times, picking at random among the measurements that already exist, to see how much the result wobbles. The technical name is bootstrap.

Seed

The number that starts the draw. Same seed, same draw, same result. It looks like solidity; it is just repetition.

Coin share

The share of verdicts that changes when only the seed changes. Those were not decided by the measurement, but by chance.

Where the problem came from

Comparing averages invents winners.

We measured the switch of Go's garbage collector to a new version: 198 comparisons across programs and metrics, on one machine. Before looking at the result, we ran the test almost nobody runs: comparing the program with itself. The right answer is "nothing changed", every time.

The rule most used in CI pipelines, "is the new average more than 5% off the old one?", declared a winner in 33 of 198 comparisons, 16.7%, between a program and itself. That is the machine's normal noise becoming news.

The rule we use instead looks at the range the difference probably lies in and compares that range with the limit. If the whole range is past the limit, it decides. If it is entirely inside, it decides "same". If it crosses the line, it says undecidable. In the same test, it invented zero winners.

INVENTED WINNERS · 198 IDENTICAL PAIRS Average against the limit33 Significance test2 Range against the limit0 and the min-max envelope, also 0
The same program against itself, 5% limit. Our measurement: Apple M2, Go 1.26.5.
The next question

And where does the range come from?

The range is computed by drawing. Take the 12 runs of each version, draw 12 with replacement, compute the difference, and repeat that hundreds of times. The drawn differences show how much the result wobbles. The range is the middle stretch of them.

Like every draw on a computer, this one starts from a seed. Our tool always used the same one, 12345. It is common practice and for a good reason: running again gives the same number, and nobody gets confused.

But repeating the same draw is not the same as the draw not mattering. In comparisons where the end of the range touches the limit, one seed puts that end on one side of the line, another seed puts it on the other. The verdict flips. The data is the same; only the coin changed.

LIMIT 5% −5% seed 1same seed 2undecidable seed 3same seed 4undecidable seed 5same same data, five draws: the right end dances around the line
Sketch. A comparison whose real difference is about 1% slower, with the range stretching to almost 5%. Which seed was the published one decides what goes into the report.
What we measured

200 seeds, and 3.6% of the decided verdicts flipped.

We took the 198 comparisons already published and redid only the draw, with 200 different seeds. No new data, no new runs. Before running, we recorded the prediction: among decided verdicts, the coin share is above zero and below 5%. Small enough not to topple anything, big enough to exist.

Published verdictComparisonsChanged under some seedCoin share
Decided (faster, slower or same)11043.6%
Undecidable8844.5%
All19884.0%

The prediction held: 3.6% is inside the range recorded beforehand. Our measurement, August 2026 (the seed audit): 198 program-metric pairs from Go's garbage collector, Apple M2, Go 1.26.5, each recomputed with 200 seeds.

We also redid, with the 200 seeds, the most quoted number of that study: of the comparisons where a statistical test says "there is a difference", 81.6% do not support a decision against the 5% limit. It ranged from 78.9% to 84.2%, and landed exactly on 81.6% in 170 of the 200 seeds. The number survives. But now that the variation is measured, quoting only the central digit would be picking the prettiest version. The honest form is "81.6% (78.9% to 84.2% across draws)".

And one case deserves a name. In a comparison of reading HTTP requests, counting collector cycles, the published verdict was undecidable. Across the 200 seeds, 52% say "same". It is the only case where the published verdict is the minority side: it was literally heads or tails, and the coin landed on the side that got printed. It changes no conclusion, because both verdicts say "you cannot act on this". But it is the clean example of the phenomenon.

Not the method, the place

The coin lands where the decision matters.

We ran the same audit on two more of our tools, and each needed a different test, because in each one something different wobbles the boundary. In one it is the seed. In another, the error margin the manufacturer declares in the spec. In another, which stretches of data were set aside for calibration.

The result ranged from 0.18% to 10.5%, and what explains the difference is not the tool: it is how close the cases live to the limit. The antennas sat, at the median, 7 dB from the line; almost nothing flips. Benchmarks pile up near the action limit, because that is exactly why someone wanted to compare them.

The consequence is uncomfortable. The coin share is largest exactly where the decision is hard, because that is where cases accumulate. A verdict that never touches the boundary was not deciding anything that needed measuring. In the specimen above, it is the difference between options a and b.

HOW MUCH FLIPS WHEN ONLY NON-DATA CHANGES 5G antennaswhat wobbles: the spec's error margin 0.18% Go benchmarkswhat wobbles: the seed of the draw 3.6% Sensor attestationwhat wobbles: which stretches calibrate 10.5%
Coin share in three of our tools. For the antennas, the sweep was synthetic. In sensor attestation, the 10.5% shows up only at the setting where the two sides of the calculation differ by 2%; at every other setting, it was zero.
The fix, and its trap

Requiring a margin works. Picking the margin afterwards does not.

The obvious fix is to require a margin: decide only if the result cleared the line with room to spare, and call everything close undecidable. In sensor attestation, requiring double wiped out the coin zone, in a tool built before that rule existed. That is a good sign the rule was not tailored to one case.

The trap is picking the size of the margin by looking at the table. A margin of 1.5 times changed nothing; 2 times wiped it out. Taking 2 because it gave the pretty result is the same mistake under another name: a limit chosen after seeing the verdict. We tried three ways of arriving at the value. From the instrument's own variation, it came out at 1.4. Looking for where the coin disappears, 2.0. The third could not point to any value. The margin has to exist; its right value is still an open question.

What to do

Four rules for a verdict that survives being redone.

  1. Ask what, besides the data, moves the line

    The seed of the draw, the error margin of a spec, which samples were left out. Each tool has its own. If nobody can say which it is, nobody measured it.

  2. Draw again before publishing

    Change only what is not data and see which verdicts flip. It costs minutes of computer time, with no new measurement.

  3. Quote the number with its range

    "81.6%" and "81.6% (78.9% to 84.2%)" say different things. The second tells how much of the number is measurement and how much is the draw.

  4. Declare the limit and the margin first

    The action limit and the required margin go into the report before the result. Chosen afterwards, they are just another way of reaching the answer you wanted.

What we do not know yet

Where this could be wrong.

One machine, one pair of versions

The 3.6% comes from one Apple M2, one Go version and one collector switch. The mechanism is general; the size of the share depends on where the cases live, and that changes from study to study.

200 seeds are also a sample

A comparison that flips in one seed out of a thousand may not have shown up. The measured share is a floor, not the exact count.

The antennas were not real

The antennas' 0.18% was measured on a synthetic sweep generated by the tool itself. It describes the method, not an installed antenna.

The margin's value is open

We know requiring a margin works and that picking it from the result is cheating. We do not yet know how to derive the right value from outside, for any tool.

← Research notebook · Certified telemetry · stickybit.com.br

Sources