CLAMP · questions with a guaranteed answer

The answer, margin included.

How many hours did emissions exceed the limit? How many times did the aircraft bank too far? CLAMP answers straight from the compressed file, reading only a small part of it, and returns a range that is guaranteed to contain the true answer. Not "95% likely": guaranteed, and anyone can redo the math.

Specimen · 24 hours of engine temperature, one reading per minute

Question: how many minutes were above 80.0 °C?

File tolerance
stored temperaturelimitcertainly abovemaybe (at the edge)hour that had to be opened
[0, 0]guaranteed answer (minutes)
0true answer
0hours opened, of 24
0%of the file read

Illustrative data, generated with a fixed seed. The rule is the same as in the measured CLAMP: a range from "certainly above" to "certainly above + maybe", and each hour is opened only if its minimum and maximum are not enough to decide.

In everyday terms

Speed cameras have a tolerance too.

On a 60 km/h road, the camera reads 61. Were you speeding? It depends on the device's tolerance. If it can be off by 2 km/h either way, 61 might be 59. The honest answer is: "certainly above? no; maybe".

A compressed file of measurements is in the same situation. TUBE stores each reading with an agreed tolerance: the stored value never strays from the original by more than, say, 0.5 °C. Anyone asking "how many times was it above 80 °C?" has to account for that tolerance, or they will be wrong exactly in the cases that matter: the ones close to the limit.

CLAMP splits the readings into three groups: those that certainly crossed (even after subtracting the tolerance), those that certainly did not, and those at the edge. The answer comes out as a range: at least the certain ones, at most the certain ones plus the edge ones. The truth is always inside.

limit: 80 °C ± tolerance certainly below certainly above: 5 maybe: 2 answer: [5, 7]
Seven readings above 80 °C? Maybe five, maybe seven. The range says exactly that, and never more.
How it works

Open only the hours you can't decide from outside.

TUBE stores the file in blocks of readings and, next to each block, notes its smallest and largest value in 2 to 4 bytes. It is like the label on an archive box: "2019 invoices, January to March".

For the question "above 80 °C?", CLAMP reads the labels first. If a block's largest value plus the tolerance is still below 80, the whole block is below, with proof, and it is never opened. If the smallest value minus the tolerance is already above, the whole block is above, and the label is enough. Only the blocks that cross the limit are opened.

When the event is rare, that is almost nothing: for a question about the top 0.1% of 100,000 readings, CLAMP opened 2 of 391 blocks. And the range it returns is identical to opening everything.

Proving that something did not happen costs even less: if no label comes near the limit, the answer is [0, 0] without opening a single block.

limit skipped: below, with proof opened: they cross the limit answered by the label each box: smallest to largest value in the block
Each block's label (its smallest and largest value) settles most cases without opening anything. Only blocks that straddle the limit need to be read.
What we measured

On real data, reading little of the file.

We ran the same kind of question on five public datasets, each with its own legal or safety threshold. In every case, the range contained the true answer.

At the coal plant, the US regulatory question (how many hours did the NOx rate exceed the 0.15 limit?) gave [162, 162]: an exact answer, with no hour at the edge. On the flight, "how many readings exceeded 25° of bank?" gave [612, 620], with the truth at 616, reading 2.3% of the recording.

In nuclear-test-ban monitoring, proving there was no signal on a seismometer cost zero blocks opened.

The contrast that matters: the usual way to answer fast is to sample part of the data and give a 95% margin. Over 2,000 questions, that method missed the truth 102 times (5.1%). CLAMP missed 0.

Nuclear treatydetection · 26 of 242 blocks
11%
Commercial flightbank angle · 4 of 176 blocks
2.3%
Crypto exchangesurveillance · 3 of 2,130 blocks
0.14%
Nuclear treatyproof of absence · 0 of 242
0%
Share of the compressed file that had to be read to answer. The rarer the event, the less gets opened.
Case (public data)QuestionGuaranteed answerTruthFile read
Barry plant (US), 2023 · demohours with NOx above 0.15[162, 162]162—
NASA DASHlink flight · demoreadings with bank above 25°[612, 620]6162.3%
Seismometer EKB5, Scotland · demoany signal above the threshold?[96, 102]10011%
Seismometer EKB10, Scotlandprove nothing crossed[0, 0]00%
Bitcoin, 1 day, 638,484 trades · demotrades outside the ±0.5% band[795, 795]7950.14%
Continuous glucose (Colás 2019) · demotime in the 70–180 mg/dL range, at the sensor's legal accuracy (±15)[84.2%, 95.3%]reported 90.3%—

The glucose case shows the other side: when the tolerance is large near the decision, the range gets wide, and that is information. A study announcing "time in range rose from 65% to 72%" may mean nothing once the sensor's accuracy is counted.

Where to use it

When someone will be held to the number.

  1. A limit written in law or contract

    Emissions, trade execution quality, vibration exposure. "Did it exceed the limit?" has consequences, and 0.1499 versus 0.1501 decides.

  2. An audit months later

    The auditor redoes the calculation from the stored file and reaches the same range, without having to trust whoever answered the first time.

  3. Proving something did not happen

    "No nuclear test", "no bank beyond the limit on this flight": the proof of absence comes from the labels, without opening the file.

  4. Where it isn't worth it

    If nobody will dispute the number, an ordinary dashboard is enough. And if the file's tolerance is large near the limit that matters, the range comes out wide; then the right move is to store with a smaller tolerance.

Three words on this page
Agreed tolerance

How far each stored reading may stray from the original. It is chosen upfront by whoever uses the data, and TUBE guarantees it is never exceeded.

Guaranteed range

The answer as [minimum, maximum], with the truth always inside. Not "probably inside": inside, and checkable.

Proof of absence

Showing that something did not happen, using only the block labels. It costs almost nothing and is worth as much as finding the event.

Limits

Where this could be wrong.

The range widens near the limit

For a sum over 100,000 readings, the range is 0.08% wide with no filter, but reaches 2.9% when the question is about the top 3%, because many readings sit at the edge there. The width is the honest price of the tolerance.

It needs a file with a guaranteed tolerance

CLAMP works on data stored with a per-reading guaranteed tolerance, like TUBE's. On a file that only promises "small average error", there is nothing to guarantee.

The competitor was measured by method, not product

The "5.1% error from sampling" comes from the method approximate-query systems use (sample and give a 95% margin), run by us. Running the installed competing product is still on the roadmap.

Not the first to give ranges

Earlier research already gave guaranteed ranges for some kinds of compression. What sets CLAMP apart is working with any compressor that guarantees a per-reading tolerance.

See also

← Certified telemetry · stickybit.com.br

Sources