Methodology

Every figure traces to one source release. The pipeline refuses to run against a source whose licence has not been checked. Anything we cannot verify is shown as absent rather than guessed. This page explains what each of those means in practice, because a methodology page that only asserts good intentions is not worth reading.

A source is registered before it can be used

Each source is a row before it is a download. That row carries the publisher, the licence, the exact attribution string the licence requires, and — the part that usually gets skipped — a record of how the licence was established. “Read verbatim from Metadata.csv line 5 of the July 2026 release” is checkable by the next person. “Verified” is not.

Ingestion refuses to run against a source whose licence is unverified or more than a year stale. That is a constraint in the database, not a note in a document, because a rule that lives only in prose is a rule that gets written after the code.

We never construct a download URL

Publishers move files. The Insolvency Service’s July 2026 release ships a file named for June; the Innovate UK spreadsheets moved directory within a single day while we were looking at them. Assets are resolved from the publisher’s own index at the moment of fetching, matched on title rather than filename, and the bytes are checksummed. The same bytes can never be ingested twice.

Every figure carries its release

A published number is stored against the specific dataset version it came from, with the checksum of that file and the moment we fetched it. That is what makes a correction possible at all: when a publisher restates a figure, we can say exactly what we previously showed, from which release, and on what date.

A restatement is a correction, not news

Official statistics get revised, routinely. A pipeline that treats a revised figure as a change in the world will publish a false story — and once that reaches a third party or a language model’s cache, it cannot be retracted. So the two are separate events in our database. A new period is a change. A different value for a period we already published is a correction, and it is recorded against the pages that carried the old number.

What we refuse to publish

No named individuals, at any grain, in any dataset. No postcode-level pages or downloads. No figure without the release it came from. No time series that silently crosses a boundary change. And no prose written by a language model — every page on this site is rendered from rows, so a page cannot state a number the database cannot produce.

Where we are certain to be wrong

Sources have gaps and we would rather name them than average over them. The current ones are listed on the status page and repeated on the pages they affect: a missing nation in one dataset, a placeholder where an instrument should be, a series that starts later than you would want. Where a number is absent we show it as absent.

Detail