Stickybit← TelemetryPortuguêsCase · bioacoustics · SIEVE
Case · fin whales on a fiber cable

Listening only to the whale's band, without missing a call.

Every whale call detector starts by filtering the call's frequency band. We show that this filter can become a proof: in a search built this way, no window the criterion asks for is left out. On 928,000 windows from a fiber cable off Oregon, the search discarded 95% of the data while losing nothing. We also publish the two places where the first version of this analysis was wrong.

Specimen · the same filter size, three choices of where to listen
record energy (ship and ocean noise)fin whale callwhere the filter listens
95.26%of the data safely discarded
27 mssearch time
100%of the filter inside the call's band
yesmatches the full scan

The tile numbers are measured on the 928,422 real windows. The spectrum drawing is illustrative. All three filters are equally safe: none loses a window. Only how much each one can discard changes.

The real problem

A detector that cannot tell you what it missed.

The fin whale sings very low, a pulse around 20 Hz, near the edge of what human ears can hear. To find these calls in hours of recording, the standard method compares the record with a template call and keeps the stretches that look similar enough.

It works and it is fast. But it is silent about its own misses. If a call was dropped because the threshold sat one notch too high, because the filter edge clipped it, or because a resolution step smoothed it away, nothing in the output says so.

In a population study, that gets absorbed into the statistics. But in the monitoring required during seismic surveys at sea or offshore wind installation, the regulator's question is exactly the one the detector cannot answer: did you miss one?

recordhours of fiber standard detectorsimilar enough? findingsstretches kept what was left out vanishes silently recordhours of fiber search with proofsame criterion findings+ guarantee:nothing the criterionasked for was dropped
The same criterion, with and without proof that nothing was dropped. The difference is not in what gets found: it is in what you can claim about what did not.
In everyday terms

At a noisy party, listen only for the low voice.

Imagine looking for a friend with a very deep voice at a crowded party. A clever trick is to "cover" the high notes and listen only to the low range. You hear less, but his voice is still there in full, because it was never in the high notes.

Filtering a signal's frequency band is exactly that. And it has a mathematical property that makes the trick safe: throwing away part of the signal can only make two stretches look more alike than they were, never less. If a stretch already looks little like the call in the low range alone, it looks even less like it overall.

That is what allows stretches to be discarded safely: if even in the call's band it does not come close to the template, it can be discarded without risk. Not with high probability. Always. The similarity search that does this with a guarantee is our tool SIEVE.

real distance inside the call's band only outsidethe band the in-band distance never exceeds the real one: if it is already too large, the stretch can go
Pythagoras, literally: the total squared distance is the in-band part plus the out-of-band part. The in-band part alone never overstates similarity.

d² = ‖in band(q) − in band(x)‖² + ‖outside(q) − outside(x)‖²The second term is unknown, but it is pinned from both sides by a single number stored per window: how much of that window's energy lies outside the band. With that number, the in-band distance becomes a guaranteed floor for the real distance.

Two conditions make this true, and outside them the guarantee does not hold: it is about distance in the declared spectrum, not the raw waveform; and it only holds when selecting bands in a well-behaved (orthonormal) frequency basis. An ordinary band-pass filter applied in time does not have this property.

What we measured

928,000 windows, 95% discarded with proof.

The data. A public, labeled set of fin whale recordings made with optical fiber turned into a microphone (DAS) on the OOI South cable, off Pacific City, Oregon. We took 1,563 stretches of the cable between 15 and 65 km from shore, cut them into 0.70-second windows, 928,422 windows in total, and used the authors' own template call (14 to 24 Hz, 0.70 s) as the query.

The same criterion, with and without proof. The authors' criterion is "similar enough to the call" (correlation of at least 0.2). With normalized windows, that becomes, with no approximation at all, a distance search. It is not our criterion against theirs: it is their criterion with proof that nothing was left out.

The result. Listening only to the call's band, the search discarded 95.26% of the windows in 27 milliseconds, and the answer matched the full scan exactly. We had declared, before running, that below 95% the idea would not pay its own cost. It passed by a hair, and we say so. With stricter criteria, the discard goes past 99.9%.

DATA SAFELY DISCARDED (SAME 14 DIMENSIONS) call band 95.26% PCA's choice 8.54% 16 references 0.15% all three match the full scan; only the work saved differs
All three filters are equally safe, because the guarantee only requires a well-behaved frequency basis. A poorly chosen filter costs speed, never a detection.
Filter (same 14 dimensions)Data discardedTimeMatches the full scan
Call band (the physics)95.26%27 msyes
PCA-chosen subspace8.54%145 msyes
16 reference windows (pivots)0.15%183 msyes
The warning that applies beyond this case

The statistics listened to the ship, not the whale.

The difference between listening to the call's band and letting PCA choose was 11 times in discard. The reason deserves to be said plainly.

PCA is a statistical method that finds the directions where the signal varies most. It did exactly that. But in a fiber record at sea, what varies most is ship noise, ocean noise and the equipment itself. The whale's call is weak by nature and barely weighs in the total variation. PCA spent 95% of its dimensions on frequencies where the animal is not.

The warning generalizes. Methods that summarize the signal on their own, without knowing what is being sought (PCA, autoencoders, representations learned from the raw record), optimize something other than detection. With a strong signal, the difference vanishes. With a weak call under broadband noise, it is an order of magnitude. If your pipeline compresses the signal before detecting, check where those learned dimensions fall relative to your species' band.

ENERGY INSIDE 14–24 Hz fiber recordwhat PCA optimizes 10.4% whale callwhat we look for 99.5% PCA's choicewhere it listened 4.6%
The call lives almost entirely in the 14–24 Hz band. PCA's subspace barely touches it.
Two corrections to what we first claimed

Where the first version was wrong.

After publishing the first version of this analysis, we ran two measurements we should have run before. Both contradict part of what was written, so both are here, next to the result.

  1. The filter alone would have failed

    Using only the in-band distance (the textbook technique, from 1994), the search discards just 1.87% of the data at the authors' criterion. What does the work is the extra number stored per window, the energy outside the band: with it, the discard goes to 95.26%. The reason is asymmetry: the call has almost all its energy inside the band, and a typical fiber window has almost all of it outside.

  2. The guarantee protects a step that was not bleeding

    We compared the criterion we certify with the one a bioacoustician actually runs (similarity on the already-filtered band). Ours is always a subset of theirs: zero windows only ours, and theirs has about 13 times more. Whoever filters the band was not losing calls at this step. The losses happen later: in peak picking, in the threshold and in other filters. We implied protection at a point that was not bleeding, and that was wrong.

Similarity thresholdIn-band distance only (1994)With out-of-band energyDeclared criterionPractitioner's criterionOurs only
0.201.87%95.26%13,709183,8410
0.306.47%99.27%–––
0.5072.43%99.94%36620,0540

What survives is narrower and still worth saying: band filtering can be made provably complete at negligible cost, and the extra number per window makes the search discard enough to be worth running. Whether that matters to you depends on where the losses you worry about happen. From the evidence we have, they happen after this step.

Why proof matters for science

One less term of uncertainty.

The regulatory use is obvious. What interests us more is estimating how many whales there are.

Estimating population from sound requires knowing the chance of detecting a call. Today that chance mixes two different things: physics (how sound travels, the animal's position, sea noise) and the algorithm (what the detector let slip). They are usually calibrated together.

If the algorithm's loss is provably zero with respect to the declared criterion, the two separate. Only the physical part remains, which the acoustics community already knows how to model. It is not a better estimate: it is one less source of uncertainty that nobody should have had to quantify in the dark.

today physicsalgorithm mixed, calibrated together with proof physicsalgorithm: zero only what acoustics already models remains
The algorithm's loss is zero only with respect to the declared criterion. That says nothing about biological correctness (see the limits below).
Three words on this page
Call band

The frequency range where the call happens: here, 14 to 24 Hz. Filtering means listening only to that range.

Provably complete search

A similarity search that returns every window meeting the declared criterion, with a mathematical guarantee, not with high probability.

Window

A short piece of the record (here, 0.70 seconds of one stretch of cable). The search compares each window with the template call.

Limits

Where this could be wrong.

No biological validation

The 0.2 threshold belongs to the authors' pipeline, which includes a filter we did not replicate. In our preprocessing it accepts 13,709 windows across the record, while the annotation marks one call. We prove completeness with respect to a declared criterion, not biological correctness.

No accuracy against annotations

The set was built to estimate source level and has three annotations in total, one per cable. Excellent for its purpose, insufficient to evaluate a detector.

Only the 20 Hz pulse

We did not test telling species apart, only the presence of the fin whale pulse.

A run that returned zero

An earlier run found no window at all: the raw fiber signal is dominated by a very slow drift. We added a filter from 5 Hz up, well below the call's band, and declared it before looking at the result.

What is still needed for this to become a detection result, not just an efficiency one: a set with many annotated calls, against which accuracy can be measured. Candidates we looked at: the Antarctic library of baleen whale vocalizations (~1,880 annotated hours, including the 20 Hz pulse) and the DCLDE sets. If you work in acoustic monitoring, we would like to know which you would use, and where this analysis is wrong: andrey@stickybit.com.br.

See also

← Certified telemetry · stickybit.com.br

Sources