How we do this
Every figure in these reports comes from a scan we ran, not from a dataset someone else published. The method is the same each time: take a defined list of domains, fetch one page from each, extract its links, and request every destination — then report what came back.
Three rules keep the numbers honest:
- Say what was measured. Where a study covers a whole list we call it a census; where it covers the top of one we call it a sample and explain why.
- Exclude rather than assume. Pages we could not fetch, and pages that answered with no links at all, are left out of the totals instead of being counted as clean. That makes every rot figure conservative.
- Separate refusals from rot. A
403is usually a server declining an automated request. It is reported, but never counted as a dead page.
The aggregate data behind each report is published as JSON under a CC BY 4.0 licence. If you want to check our working, or use the numbers, please do — a link back is all we ask. If something looks wrong, tell us and we will look.
We publish hosts, never the specific dead URLs we find on other people's sites. The interesting question is what the web points at that has gone away, not which individual pages are broken today.