Stickybit.← NotebookPortuguêsMeasurement · the Go language · Sep 28, 2026
Measurement

Faster, different bits.

Processors know how to do the same arithmetic on several numbers at once. The Go language now offers this in a way that runs on any chip. We measured where it speeds things up (2.2× to 2.7×), where it gets worse (10× slower) and the side effect nobody advertises: the same sum gives a slightly different result depending on the machine.

Specimen · sum of squares of 4,096 vibration readings
Σx² · one number at a time…
the 32 bits of the result: sign · exponent · digits0 differs from "one at a time"

largest value among the same readings… same in every mode

The arithmetic runs right here in your browser, with the 32-bit precision the processor would use: the numbers are split across the lanes and the partial sums are combined at the end, as the compiler does. The exact bits vary with the implementation; the fact that they vary does not.

What SIMD is

One instruction, several numbers at once.

SIMD stands for "single instruction, multiple data". Instead of adding one number at a time, the processor takes a block of 4, 8 or 16 numbers and does the same arithmetic on all of them with a single command. Each position in the block is a lane.

Think of a supermarket checkout. The regular cashier scans one item at a time on a single belt. The SIMD cashier has four belts side by side, and a single motion scans four items. At the end, the cashier combines the four subtotals into one total. That is why it is faster: fewer motions for the same number of items.

And that is why the result changes. Decimal numbers in a computer have limited precision: after about seven significant digits (at 32-bit precision), the rest is rounded. Every addition rounds a tiny bit. Adding in a different order, in four subtotals instead of one line, rounds in different places, and the last digit comes out different.

one at a time one belt: 8 motions Σ added in a line SIMD four belts: 2 motions 4 subtotals Σ another order,another last digit
The same cart at both checkouts. In math the total is the same; in a computer, with limited precision, the order in which the additions happen changes the last digit.
What changed in Go

One program, several block sizes.

Up to version 1.26, Go only offered SIMD for Intel and AMD chips. Version 1.27 added ARM chips (the ones in Macs and phones), WebAssembly (which runs in the browser) and, most importantly, a portable package: you write it once, and the compiler generates one version per block size, chosen when the program starts running.

Each chip family has its own way of doing SIMD. NEON is the ARM one and handles 4 32-bit numbers at a time. AVX2 is the one in most Intel and AMD servers, with 8. AVX-512 is the one in the most expensive servers, with 16.

That is convenient: the same program runs with 4 lanes on a Mac, 8 on a typical server and 16 on a high-end server. And that is exactly where the problem lies: how many numbers each lane adds, and in what order, depends on the machine.

Pieces are still missing: there is no ready-made command to combine the lanes into a single sum (it arrives in version 1.28), nor to scatter results to arbitrary memory positions. The package is experimental and may change.

one at a time a single line x₀x₁x₂x₃x₄… Σ 4 lanes alternating terms, combined later x₀x₄ x₁x₅ x₂x₆ x₃x₇ Σ another order,another rounding
With decimal numbers, (a+b)+c does not always equal a+(b+c) in a computer. With 8 or 16 lanes, the way the subtotals are combined changes again.
Measured

Where it speeds up, where it gets worse.

We tested three pieces of code that run millions of times in two of our systems. The first is the check in our telemetry codec, TUBE: it confirms that no reconstructed value drifted from the original by more than the agreed tolerance. The second is a histogram of the changes in the same signal, which counts how many changes fall into each size bracket. The third is the distance between vectors in our search index, SIEVE. We used real data: 182,976 vibration readings from a bearing (CWRU, 12 kHz).

SIMD made the check 2.2× faster and the distance 2.7× faster. For the histogram, nothing: the heavy work there is dropping each value into its bin, a different memory position each time, and the package cannot do that in blocks.

The surprise came from the emulation mode. When the chip has no SIMD, or when someone turns it off with an environment variable, Go emulates the instructions in software. The documentation promises this works correctly. We measured it 10× slower than the plain version for the check and 4× for the histogram. A forgotten setting on the server becomes a large, silent slowdown.

Measured on a Mac with an ARM chip, median of 6 runs, without isolating the processor from other tasks: there is noise. The histogram difference fits within that noise. We have not yet measured on Intel/AMD servers with AVX2 or AVX-512.

with SIMD emulation mode
Codec check · largest gap between original and reconstructed
Histogram of changes · 128 bins
Distance between vectors · 128 dimensions, 1,024 groups
Speed compared to the plain version, one number at a time, on a scale that grows by multiplication. Hover or tap a point to see the timings. Emulation mode for the distance was not measured.
CodePlain versionWith SIMDEmulation modeResult
Codec check3.87 ns/reading1.74 (2.2×)37 (≈10× worse)identical: finding the largest value does not depend on order
Histogram of changes6.6 ns/reading7.2 (≈1×)26 (≈4× worse)identical: each calculation is done on a single number, with no adding to others
Distance between vectors335 µs126 (2.7×)not measureddifferent: 32-bit sum, off by ~1.5 parts in 10 million
The risk

The same program, three fingerprints.

Our telemetry codec (TUBE, part of the certified telemetry family) promises one thing above all: the same file, byte for byte, on any machine. That is what lets us seal the file with a fingerprint (a mathematical summary that changes if a single bit changes) and check it later somewhere else.

We moved the sums the codec computes for each block of readings to SIMD and compared them with the plain version: the result changed in 98.7% of the blocks. The numbers that actually go into the file came out the same in all 304 blocks, because the final rounding absorbed the difference. But that was luck with this signal, not a guarantee: a sum right on a rounding boundary becomes a different number and changes the file.

The option to multiply and add in a single step makes things worse. The same program, with the same data, produced one fingerprint on the ARM chip, which does that single step, and another on an Intel chip emulated by Rosetta 2 (the Mac's translator for Intel programs), which lacks that feature and does the multiplication and the addition separately. A program that adapts to the block size also adapts the result of any sum of decimal numbers.

Σ x·x in a single step one program, one dataset ARM chip (Mac)4 lanes · single step 302e… Emulated Intelno single step dbb3… High-end server16 lanes · AVX-512 not measured same program, a result that depends on the machine
The first characters of the fingerprint (SHA-256) of the sums of all blocks. We did not measure the third path: the Mac's translator does not emulate AVX-512, and the result on a real server is left for the next round.
The rule

What we decided.

  1. SIMD only for arithmetic that does not depend on order

    Finding the largest or smallest value, comparing, subtracting, dividing each number by another: the result is the same in any order. Sums of decimal numbers that end up in a file or a fingerprint stay out. It is the same rule that already banned "multiply and add in a single step" in the codec, extended to the order of additions.

  2. Emulation mode is not a plan B

    Keep the plain version, one number at a time, written and chosen on purpose. Relying on emulation means accepting a 10× slowdown nobody sees until the bill arrives.

  3. Where the sum is unavoidable, check near the boundary

    The index distance needs a sum. So SIMD decides on its own only when the answer is far from the boundary; near it, it recomputes at full precision (section below).

  4. Measure on the real machine

    The block size changes with the chip, and so does the result. Speed and bit tests have to run on the production server, not just on a laptop.

What about the index guarantee?

Decide fast far from the edge, double-check near it.

Our search index (SIEVE) promises never to leave a neighbor out: if an item is within the search radius, it shows up. With SIMD, the distance comes out at 32-bit precision, with a small error. We calculated the maximum size of that error, about 8 millionths of the distance itself for 128-dimension vectors, and tested it against the exact calculation on 3,000 pairs, with four kinds of data, three vector sizes and three block sizes, with and without the single step: the error never went past the limit.

Then we deliberately planted 20,000 points right at the search radius. The 32-bit calculation alone got wrong whether 9 to 38 of them were inside or outside, which would break the promise. With a double-check band (near the radius, recompute at full precision), it got zero wrong.

The cost of the band: 16 recomputations in 200,000 (0.008%). The 2.7× gain survives almost whole.

search radius band: recompute exactly (0.008% of cases) inside: the fast calculation decides outside: the fast calculation decides distance to the query point →
The band's width is exaggerated in the drawing: it is about 8 millionths of the distance, and in practice the calculated limit was 8 to 250 times looser than the largest error we observed.
Where it pays off, in general

Repeated code, heavy arithmetic, small numbers.

SIMD pays off most when the same arithmetic repeats over lots of independent data that is already in the fast memory close to the processor, and when the arithmetic is the bottleneck, not waiting on memory or "if this, do that" decisions. The closer the code gets to that, the bigger the gain.

Pays off a lot

  • Text: parsing JSON, validating accented characters (UTF-8), finding commas and line breaks in CSV, base64.
  • Compressing and checking files: packing numbers, checksums (CRC), encryption (AES).
  • Similarity search: distance between vectors, especially with whole numbers.
  • Images, audio, huge spreadsheets and analytics databases.

Pays off little or gets in the way

  • Calculations where each step depends on the previous one: entropy coders, parsers with many states, adaptive prediction.
  • Code that waits on memory and scattered access (graphs, lookup tables).
  • Sums of decimal numbers that must give the same result on any machine.
  • Short pieces of code, where preparing the block costs more than the arithmetic.

If the code repeats a lot, the arithmetic is the bottleneck and you can use whole numbers, it pays off a lot. If any of the three is missing, change the algorithm or the data layout first: that gain is usually bigger and survives a change of machine. In our index, the distance between vectors made of whole numbers from 0 to 255 is exact in 32 bits, because the sum fits entirely within the available precision; the distance to the center of a group, which has decimals, is not.

Three words from this piece
Lane

Each position in the block the processor handles at once. A Mac chip has 4 lanes for 32-bit numbers; a high-end server, 16.

Limited precision

The computer stores a decimal number with about seven significant digits (in 32 bits). The rest is rounded, and the order of the arithmetic decides where.

Fingerprint

A mathematical summary of the file that changes if a single bit changes. It proves two files are identical without comparing them in full.

What we have not measured yet

Where this could be wrong.

A noisy Mac

Without isolating the processor from other tasks or pinning the clock speed. The next round is on a Linux server, with the test pinned to one core and enough repetitions to separate signal from noise.

A real high-end server

The Mac's translator does not emulate AVX-512. How the "single step" behaves and how much 16 lanes gain on a real server remains open.

From code snippet to system

The 2.7× comes from a piece of code with the data already in fast memory. In the whole index, with millions of vectors, waiting on memory may eat a good part of it. Only measuring the complete system will tell.

Experimental package

The package will change until it stops being experimental. None of this goes into production before then; what already applies is the rule of not summing with SIMD what must give the same result.

← Research notebook · stickybit.com.br

Sources