A locked box with gloves on the outside.
Picture a locked glass box with rubber gloves built into its wall. You put the pieces inside, lock it and hand the box to a jeweler. They assemble the ring through the gloves, never touching the pieces, and give the box back still locked. Only you have the key.
That is what homomorphic encryption (FHE) does with data: the server computes on the encrypted data, never opening it, and returns a result only the key holder can read. It lets you process health, financial or personal data in a cloud you do not have to trust.
The limit shows up in the analogy too. Through the gloves you can add and multiply. You cannot look: the jeweler cannot see which stone is bigger. Every decision ("which is largest?", "did it cross the limit?") is hard under encryption, and the cheapest way out is almost always to send it back to the client to decide.
Three documents say it became practical. None publishes the numbers side by side.
Google’s announcement presents HEIR and four demos (recommendation, card fraud, intrusion detection and voice wake-word), all on a single processor core and with no timings published. HEIR is the code. The HE-LRM paper does private lookup in a recommendation table and reports 24 s to 489 s.
The observation that organizes everything: HEIR is not a new scheme nor an accelerator. It is a compiler that emits code for the same libraries we use, Lattigo included. Its performance ceiling is, literally, what this bench measures. It automates tedious, valuable rewrites; it does not change the physics. That is why the four demos chosen are shallow models on already-summarized data: exactly the first column of the board below.
Add, score, aggregate
- adding and multiplying in batch
- dot product, dense layer
- smooth curves on a guaranteed range
- voting, counts, federated gradients
Deep, divided, large
- many multiplications in a row
- dividing by an encrypted value
- the cleanup (bootstrapping)
- bandwidth and server keys
Deciding under encryption
- compare, find the largest, sort
- exact neural activation (ReLU)
- lookup by encrypted position in a big table
- any "if" that depends on encrypted data
In one sentence: FHE works today when the computation is a fixed, shallow sequence of additions and multiplications, applied to a large batch, with every comparison and decision pushed to the client, who holds the key.
Multiplying is cheap. Moving is expensive.
The line "FHE is slow because encrypted multiplication is expensive" is wrong. Multiplying takes 2.7 ms. What costs is moving data between the positions of an encrypted package: summing a 2,048-number vector needs 12 moves and takes 112.8 ms, about 40 times the multiplication.
That is where the bench’s biggest gain comes from. Arranging numbers in the right positions before computing (packing) made catalog scoring 251 times faster and a neural-network dense layer ~126 times, with the same library, the same parameters and the same machine.
"Just parallelize" buys much less: with 4 cores, multiplying scales 3.92 times, but moving data, which is memory-bound, stays at 2.59. With 8 cores, throughput drops.
In practice: optimizing FHE is, first of all, minimizing moves. It is what HEIR automates, and where it is worth most.
The cliff of deciding.
Under encryption there are only two operators: sum and product. Comparing two values becomes a chain of nine polynomials that spends 40 levels of the budget, against the 4 to 8 of a normal parameter set. Without the cleanup it is not slow: it is impossible. And there is a measured blind zone: values that differ by less than 2⁻³⁰ (one billionth) have no defined comparison.
With the real cleanup, at real 128-bit, picking the largest of 4 took 23 min 31 s and returned one position with only 3 bits of precision. Among 1,000 candidates, the anchored estimate is 130 hours. It is slow and wrong.
The other dialect, TFHE, works bit by bit and compares naturally (1.74 s per 8-bit comparison), but it becomes millions of times slower at batch arithmetic. There is a bridge between the two: 13.9 s per position. Choosing the dialect is the most expensive decision in the project.
In practice: send the encrypted scores back and let the client decide. On the key holder’s side, the same decision costs microseconds and no privacy at all.
The cleanup, the bandwidth and the table size.
The cleanup takes 1 min 18 s at real 128-bit, after the client sends 10.26 GB of keys. It spends 15 levels to give back 10. And the "toy" parameter set, which appears in much published material, is ~40 times faster than the real-security one.
Bandwidth shows up before CPU. In a full batch, encrypted data takes 12 to 36 times the original; for a single number, 229 thousand times. Fetching a table row without revealing which one requires touching the whole table: at 1 million rows, the client sends 1.91 GB to receive 16 numbers.
Matching three banks’ lists in secret costs according to the universe of identifiers, not list size: in the universe of Brazilian national IDs, ~9 GB per bank, per round. For pure matching, the right tool is a different one.
In practice: FHE rewards big batches and punishes single lookups. Size by bandwidth, not only by time.
Attacking the price, and publishing what each shortcut charges.
"Works, but costs" does not mean impossible: it means expensive, and price can be attacked. The four new tests attack the four expensive items by changing the computation, the protocol or what travels over the wire. Each one publishes the gain and what it charges, because a shortcut that shows only the gain is advertising.
The result has a pattern: no gain came from making FHE faster; they came from doing less FHE. The biggest is skipping the cleanup: sending the package back to the client, who re-encrypts in 1.04 s with no cleanup key at all, against 1 min 18 s and 10.26 GB.
And the remaining cost has a pattern too: the three cheapest shortcuts fail silently. Cheap and silent is the combination that costs most in production.
Before the attacks, a correction to this bench
We checked the parameters against the security ceiling in Lattigo’s own documentation. Two older numbers were not 128-bit: the "8 multiplications in a row" (T2) and the 89 ms division (T3). The measurements were right for those parameters, but weaker than the label. At the real ceiling, logN=13 holds 2 levels, not 8; and division with the real range of the SUS hospital data costs 2 min 36 s. The cleanup (T4) and picking the largest (T9) already used real 128-bit and do not change.
| Test | Shortcut | Measured gain | What it charges |
|---|---|---|---|
| T13 | tree instead of queue | 28.4× more factors | 128× of memory |
| T14 | divide at the client | 57,690× | the quotient must be final |
| T15 | send back instead of cleanup | 75× · −10.26 GB | client online, and sees mid-computation |
| T15 | cleanup only when needed | 19.7× | depth known in advance |
| T16 | key travels as a seed | 2.0000× | 3.7 s of setup |
| T16 | trim the final answer | 10× | breaks above 2¹⁶ silently |
Design to fit, without cleanup.
None of the three sources is wrong: HEIR is a serious project and HE-LRM solves a real problem. What the bench adds is scale, measured on the same machine with the library HEIR itself uses. The practical conclusion is dull and useful: size the budget to fit the whole computation, pack into large batches, minimize moves and send every decision to the client. Anything that needs frequent cleanups leaves "API response" and becomes a "batch job".
Before signing an FHE pilot
The question that separates a viable project from an expensive one is always the same: how many multiplications in a row does the computation have, and where does the decision happen? If the answer involves comparing on the server, the budget is off by an order of magnitude. We write that assessment before you commit engineering.
The operation that wipes the noise built up in encrypted arithmetic and restores budget. Possible, and expensive: 1 min 18 s at real 128-bit.
How many multiplications in a row fit before the result turns into noise. It is fixed together with the keys, before starting.
One encrypted package carries thousands of numbers side by side. Full, the cost is shared; with a single number, the whole cost falls on it.
Where this could be wrong.
One machine, one laptop
Everything was measured on an Apple M-series laptop. Between sessions the same machine varies up to ~1.8×; within a session, under 10%. Ratios between operations carry to another machine; absolute times do not.
Declared projections
Picking the largest of 10, 100 and 1,000, the 10⁷-row table and the national-ID universe are projections anchored on a real measurement, and are marked as such. None was made without a basis.
Parameters below the label
Two older numbers (T2 and T3’s division) used parameters weaker than 128-bit. They are corrected above; any FHE number published by third parties deserves the same question: which logN?
What we did not measure
Transciphering (encrypting with AES and decrypting inside FHE), which would remove much of the bandwidth problem, has no implementation in our library. We did not measure it and cite no number.
← FHE: computing without opening the data · stickybit.com.br
- Google: How Google is making private AI practical with homomorphic encryption · google/heir · arXiv:2506.18150, HE-LRM.
- Our own bench (FHE repository,
12_reality_check): Lattigo v6.2.0 (BGV/BFV/CKKS) and go-tfhe v0.2.2 (TFHE), mean ± σ of 3 runs, August 2026 baseline. Every published number carries a provenance marker checked by the repository itself. - 128-bit security ceiling: Lattigo parameter documentation, following the HomomorphicEncryption.org standard. T14 real data: Brazil’s SIH/DATASUS, São Paulo, January 2024.