One instruction, several numbers at once.
SIMD stands for "single instruction, multiple data". Instead of adding one number at a time, the processor takes a block of 4, 8 or 16 numbers and does the same arithmetic on all of them with a single command. Each position in the block is a lane.
Think of a supermarket checkout. The regular cashier scans one item at a time on a single belt. The SIMD cashier has four belts side by side, and a single motion scans four items. At the end, the cashier combines the four subtotals into one total. That is why it is faster: fewer motions for the same number of items.
And that is why the result changes. Decimal numbers in a computer have limited precision: after about seven significant digits (at 32-bit precision), the rest is rounded. Every addition rounds a tiny bit. Adding in a different order, in four subtotals instead of one line, rounds in different places, and the last digit comes out different.
One program, several block sizes.
Up to version 1.26, Go only offered SIMD for Intel and AMD chips. Version 1.27 added ARM chips (the ones in Macs and phones), WebAssembly (which runs in the browser) and, most importantly, a portable package: you write it once, and the compiler generates one version per block size, chosen when the program starts running.
Each chip family has its own way of doing SIMD. NEON is the ARM one and handles 4 32-bit numbers at a time. AVX2 is the one in most Intel and AMD servers, with 8. AVX-512 is the one in the most expensive servers, with 16.
That is convenient: the same program runs with 4 lanes on a Mac, 8 on a typical server and 16 on a high-end server. And that is exactly where the problem lies: how many numbers each lane adds, and in what order, depends on the machine.
Pieces are still missing: there is no ready-made command to combine the lanes into a single sum (it arrives in version 1.28), nor to scatter results to arbitrary memory positions. The package is experimental and may change.
Where it speeds up, where it gets worse.
We tested three pieces of code that run millions of times in two of our systems. The first is the check in our telemetry codec, TUBE: it confirms that no reconstructed value drifted from the original by more than the agreed tolerance. The second is a histogram of the changes in the same signal, which counts how many changes fall into each size bracket. The third is the distance between vectors in our search index, SIEVE. We used real data: 182,976 vibration readings from a bearing (CWRU, 12 kHz).
SIMD made the check 2.2× faster and the distance 2.7× faster. For the histogram, nothing: the heavy work there is dropping each value into its bin, a different memory position each time, and the package cannot do that in blocks.
The surprise came from the emulation mode. When the chip has no SIMD, or when someone turns it off with an environment variable, Go emulates the instructions in software. The documentation promises this works correctly. We measured it 10× slower than the plain version for the check and 4× for the histogram. A forgotten setting on the server becomes a large, silent slowdown.
Measured on a Mac with an ARM chip, median of 6 runs, without isolating the processor from other tasks: there is noise. The histogram difference fits within that noise. We have not yet measured on Intel/AMD servers with AVX2 or AVX-512.
| Code | Plain version | With SIMD | Emulation mode | Result |
|---|---|---|---|---|
| Codec check | 3.87 ns/reading | 1.74 (2.2×) | 37 (≈10× worse) | identical: finding the largest value does not depend on order |
| Histogram of changes | 6.6 ns/reading | 7.2 (≈1×) | 26 (≈4× worse) | identical: each calculation is done on a single number, with no adding to others |
| Distance between vectors | 335 µs | 126 (2.7×) | not measured | different: 32-bit sum, off by ~1.5 parts in 10 million |
The same program, three fingerprints.
Our telemetry codec (TUBE, part of the certified telemetry family) promises one thing above all: the same file, byte for byte, on any machine. That is what lets us seal the file with a fingerprint (a mathematical summary that changes if a single bit changes) and check it later somewhere else.
We moved the sums the codec computes for each block of readings to SIMD and compared them with the plain version: the result changed in 98.7% of the blocks. The numbers that actually go into the file came out the same in all 304 blocks, because the final rounding absorbed the difference. But that was luck with this signal, not a guarantee: a sum right on a rounding boundary becomes a different number and changes the file.
The option to multiply and add in a single step makes things worse. The same program, with the same data, produced one fingerprint on the ARM chip, which does that single step, and another on an Intel chip emulated by Rosetta 2 (the Mac's translator for Intel programs), which lacks that feature and does the multiplication and the addition separately. A program that adapts to the block size also adapts the result of any sum of decimal numbers.
What we decided.
SIMD only for arithmetic that does not depend on order
Finding the largest or smallest value, comparing, subtracting, dividing each number by another: the result is the same in any order. Sums of decimal numbers that end up in a file or a fingerprint stay out. It is the same rule that already banned "multiply and add in a single step" in the codec, extended to the order of additions.
Emulation mode is not a plan B
Keep the plain version, one number at a time, written and chosen on purpose. Relying on emulation means accepting a 10× slowdown nobody sees until the bill arrives.
Where the sum is unavoidable, check near the boundary
The index distance needs a sum. So SIMD decides on its own only when the answer is far from the boundary; near it, it recomputes at full precision (section below).
Measure on the real machine
The block size changes with the chip, and so does the result. Speed and bit tests have to run on the production server, not just on a laptop.
Decide fast far from the edge, double-check near it.
Our search index (SIEVE) promises never to leave a neighbor out: if an item is within the search radius, it shows up. With SIMD, the distance comes out at 32-bit precision, with a small error. We calculated the maximum size of that error, about 8 millionths of the distance itself for 128-dimension vectors, and tested it against the exact calculation on 3,000 pairs, with four kinds of data, three vector sizes and three block sizes, with and without the single step: the error never went past the limit.
Then we deliberately planted 20,000 points right at the search radius. The 32-bit calculation alone got wrong whether 9 to 38 of them were inside or outside, which would break the promise. With a double-check band (near the radius, recompute at full precision), it got zero wrong.
The cost of the band: 16 recomputations in 200,000 (0.008%). The 2.7× gain survives almost whole.
Repeated code, heavy arithmetic, small numbers.
SIMD pays off most when the same arithmetic repeats over lots of independent data that is already in the fast memory close to the processor, and when the arithmetic is the bottleneck, not waiting on memory or "if this, do that" decisions. The closer the code gets to that, the bigger the gain.
- No repetition depends on the previous one. That is why the heart of our codec, where each prediction depends on the previous value, does not benefit from SIMD, and that is where the cost is.
- Data lined up and of the same type, side by side in memory. Reading or writing scattered positions kills the gain.
- Few decisions along the way, or decisions that can become a mask applied to the whole block.
- Heavy arithmetic for each number read. If the code only reads and adds once, memory speed is in charge.
- Small numbers. With bytes, 16 to 64 fit in a block, versus 2 to 8 with 64-bit numbers. That is why text processing gains more than scientific computing.
Pays off a lot
- Text: parsing JSON, validating accented characters (UTF-8), finding commas and line breaks in CSV, base64.
- Compressing and checking files: packing numbers, checksums (CRC), encryption (AES).
- Similarity search: distance between vectors, especially with whole numbers.
- Images, audio, huge spreadsheets and analytics databases.
Pays off little or gets in the way
- Calculations where each step depends on the previous one: entropy coders, parsers with many states, adaptive prediction.
- Code that waits on memory and scattered access (graphs, lookup tables).
- Sums of decimal numbers that must give the same result on any machine.
- Short pieces of code, where preparing the block costs more than the arithmetic.
If the code repeats a lot, the arithmetic is the bottleneck and you can use whole numbers, it pays off a lot. If any of the three is missing, change the algorithm or the data layout first: that gain is usually bigger and survives a change of machine. In our index, the distance between vectors made of whole numbers from 0 to 255 is exact in 32 bits, because the sum fits entirely within the available precision; the distance to the center of a group, which has decimals, is not.
Each position in the block the processor handles at once. A Mac chip has 4 lanes for 32-bit numbers; a high-end server, 16.
The computer stores a decimal number with about seven significant digits (in 32 bits). The rest is rounded, and the order of the arithmetic decides where.
A mathematical summary of the file that changes if a single bit changes. It proves two files are identical without comparing them in full.
Where this could be wrong.
A noisy Mac
Without isolating the processor from other tasks or pinning the clock speed. The next round is on a Linux server, with the test pinned to one core and enough repetitions to separate signal from noise.
A real high-end server
The Mac's translator does not emulate AVX-512. How the "single step" behaves and how much 16 lanes gain on a real server remains open.
From code snippet to system
The 2.7× comes from a piece of code with the data already in fast memory. In the whole index, with millions of vectors, waiting on memory may eat a good part of it. Only measuring the complete system will tell.
Experimental package
The package will change until it stops being experimental. None of this goes into production before then; what already applies is the rule of not summing with SIMD what must give the same result.
← Research notebook · stickybit.com.br
- The Go Blog: the portable SIMD experiment · Hacker News discussion
- Our pages: Certified telemetry (the family) · TUBE, telemetry within an agreed tolerance · SIEVE, certified similarity search.
- Data: Case Western Reserve University Bearing Data Center, 12 kHz accelerometer (182,976 readings).
- Our own measurements, Go 1.27.0 with the experimental package enabled (
GOEXPERIMENT=simd), Sep 25, 2026. The error limit for the 32-bit sum follows the classic analysis of floating-point sums (γₙ = nu/(1−nu), u = 2⁻²⁴).