Three ways of not finding.
In the specimen above, the contract can be in one of three situations. With the right label, you just run your eyes over the labels and open one box. Stored with a generic label, like "misc", it is there, but you have to open box after box until you hit it. And if it was never stored, the search opens every box and comes back empty-handed.
The third case hides a cost that usually goes unnoticed: to be sure something is not there, you have to look at everything. "I couldn't find it" is expensive. And from the outside, "I couldn't find it" and "I don't have it" look like the same answer.
This difference, between information being there and it being reachable, shows up again in three very different places: inside AI models, in the searches they run, and in agents' memory.
Lost key or empty shelf.
In August 2026, Google Research tested 13 AI models on 2,150 facts, each one asked in ten different tasks. The strongest models had stored 95% to 98% of the facts: you could show the information was in there. Even so, they could not bring up 26% to 34% of them on the first try.
Letting the model "think more before answering" helped, and it helped in a revealing way. Among the facts that were stored but would not come up, thinking more brought back 40% to 65%. Among those that had never been stored, only 5% to 15%.
The difference becomes a cheap test. If thinking more brings the answer, it was a lost key, and a hint is enough. If it doesn't, it is an empty shelf, and the answer has to come from outside, from a search. They are different problems, with different remedies, and almost every system today treats them the same way.
The answer was on the same page.
Having a search at hand does not guarantee it gets used. In September 2026, a user asked Google whether a Canadian soccer team could still make the playoffs. The AI summary at the top answered, confidently, that the team had already qualified. It was false. The correct standings appeared a few centimeters below, in the page's own results. The summary spoke from memory before looking.
A study from the same month measured the effect at scale. AI models were asked to recommend doctors in the 100 largest US cities. Without search, only 4% to 11% of the recommended doctors existed in that city. With search, 64% to 71%. The question was the same; what changed was looking before asserting.
In both cases the information was present, on the page or in the public registry. What was missing was the step of going to get it, and the system preferred the shortcut of memory.
What gets lost, gets lost at writing time.
AI agents that work for days need memory: they note down what they learned and look it up later. The temptation is to keep a summary, which takes little space. A September 2026 study swapped the model that writes this memory and measured where quality was lost. In the text-summary format, 80% of the loss happened at writing time, not at search time.
And what was lost in writing does not come back. Trying to repair the memory using only what was stored did not reach 90% recovery in any of the 48 cases tested. Those that had also kept the original history got there in 34 of 48, in one direction of the swap. No smarter reader recovers what the writer did not write down.
Another paper, on memory tied to the code history, arrives at the same place by another road: the bottleneck is what gets captured, not how you search. What the agent thought and never wrote down goes nowhere. And it adds a lesson in humility: keeping the memory silent when it is not confident improved the score from 0.29 to 0.50.
It is on the boundary. Reading it is another matter.
Theoretical physics has an extreme version of this idea. Under the so-called holographic principle, everything that happens inside a region of space would be recorded on the surface that surrounds it. It is one of the most debated ideas in the field, with strong mathematical support in model universes and no test in ours.
The detail that matters here: even where the math works best, being recorded on the boundary does not mean it can be read. There are calculations showing that decoding certain parts of the interior from the boundary can take work that grows explosively. The information is present; access may be out of reach.
It works as an image, not as an argument. The holographic principle is about a ceiling: the most information that fits in a region. It is not a recipe for reading, and it proves nothing about AI. But it shows that "it is all there" and "it can be used" are separate claims even in the most fundamental laws we know.
Four rules for whoever stores and searches.
Tell "couldn't find it" from "don't have it"
Before searching outside, try again with a hint or with more thinking. If the answer shows up, it was a lost key. If it doesn't, it is an empty shelf and only an outside search will do.
If search is there, look before you assert
For whatever changes over time (scores, stock, deadlines, who practices where), the search comes before the answer, not after the correction.
Keep the original, not just the summary
The summary is convenient, and it is the one copy that cannot be repaired. Keeping the raw source is what lets you fix the memory when the reader or the writer changes.
Knowing when to stay quiet is part of memory
Without confidence, it is better to bring up no memory than the wrong one. Saying "I don't know" costs less than confidently making something up.
Putting information into memory. It can be done carefully, with the right label, or in a hurry, inside a summary that loses detail.
What lets you find what was stored without opening everything. When it goes missing, the information is still there, but it becomes expensive to reach.
When the information was never stored. No amount of thinking fixes it; you have to search outside.
Where this could be wrong.
The numbers are other people's
None of the results in this essay are our measurements. They come from studies published in 2026, some with small samples: 48 synthetic cases in the agent-memory study, 12 to 15 cases and a single evaluator in the code-history paper.
The "think more" test changes with the model
Some argue that newer models store fewer facts to gain reasoning. If that is true, the ratio of lost keys to empty shelves changes every generation, and the test has to be redone.
The physics is only an image
The holographic principle has not been tested in our universe, and the decoding cost holds for specific cases. We use it as a metaphor, not as evidence.
The shelf is a drawing
The specimen at the top shows the logic; it measures nothing. Real searches use indexes and shortcuts that make "opening box after box" cheaper, but not free, and proving absence remains the most expensive case.
← Research notebook · stickybit.com.br
- Google Research, "Empty shelves or lost keys? Recall is the bottleneck for parametric factuality", Aug 12, 2026 (WikiProfile benchmark: 2,150 facts, 13 models).
- "Understanding AI Provider Recommendations in Local Service Markets", arXiv 2609.18341, Sep 2026.
- "Does Your Agent's Memory Survive a Model Upgrade?", arXiv 2609.05339, Sep 2026 · "Why Git Is the Memory Solution for the Agentic Development Lifecycle", arXiv 2607.14390, Jul 2026.
- "When did Google get so f-ing weird?" and the Hacker News discussion, Sep 27, 2026 (the soccer team case comes from a comment).
- Charlie Wood, "Gravity Seems Holographic. What Does That Mean for Reality?", Quanta Magazine, Sep 25, 2026 · Harlow and Hayden (2013), on the cost of decoding.