Evidence
How the coverage figure is built
Three evidence layers, counted as a union rather than a sum, each labelled by strength.
The figure on the front page is the volume of anonymised address space the dataset can make a statement about. It spans the whole record rather than a recent window: an analyst reconstructing an intrusion discovered ten months late needs the addresses that were live then. Its three inputs differ in evidential strength and are reported separately.
| Layer | What it is | Method |
|---|---|---|
| Measured | Every exit address any of our sources knows — our probing, the enriched snapshot, scan certificates — across the whole record | probe, snapshot, scan |
| Interpolated | The rest of every announced prefix in which we know at least one address, at high or medium confidence | interpolated |
| From external lists | Netblocks published sources attribute to anonymising services | feed |
Why interpolate at all
Where 30 of the 256 addresses in a block are observed and the block demonstrably belongs to one provider, treating the remaining 226 as unknown understates the evidence. Every prefix with an observation is therefore interpolated, and the confidence says how far the inference reaches: above 50 % observed it is high, from 10 % medium, from 1 % low, below that very low. A lookup answers every band, with its confidence. The figure counts high, medium and low, but not very low: that band is by far the largest — over a billion addresses in the big residential prefixes in which a single proxy exit was seen — and in a headline it would read as a measurement it is not.
Union, not sum
The layers overlap substantially: most measured addresses fall inside an interpolated prefix, and many of those fall inside a listed netblock. Summing the three totals would count the same addresses up to three times. The ranges are therefore merged, and each layer reports only its marginal contribution. The rows sum correctly because they are differences, not independent totals.
What is left out
- Datacenter and cloud ranges. AWS, GCP, Cloudflare and comparable sources describe hosting, not anonymising services. A datacenter host is not by itself a VPN exit.
- Abuse feeds. Spamhaus DROP indicates a network is listed for abuse, which is context rather than evidence of anonymisation.
- IPv6. Not collected by probing. Some lists cover it, but the figure is IPv4 only.
- No time window. Addresses retired months ago still count. Whether one was live on a given date is a point-in-time question, answered per address rather than by filtering this total.
Recomputed by scripts/compute-coverage.ts whenever the data or feeds are refreshed. If it has not run, the page shows the measured count alone rather than an uncomputed estimate.