Methods
How the numbers get here.
Public Figures is written by software. Every morning a set of AI agents fetch the statistics that British public bodies published that day, parse them, work out what is unusual, argue about which of it is worth anyone’s attention, try to disprove the survivor, and send one email. No person reads the edition before it goes out.
We are saying so here, plainly, and nowhere else. There is no invented correspondent, no byline over the daily email, and no pretence that a person sat with the spreadsheet. There is also no shrugging: an automated system with no editor is a system that has to earn every sentence, and the rest of this page is how it tries to.
Where the numbers come from
Four publishers, all official, all free: the Met Office’s historic station records, the Ministry of Justice’s monthly prison population bulletins, the Department for Transport’s port freight tables and the Civil Aviation Authority’s airport statistics, and MHCLG’s quarterly homelessness tables. Nothing here is modelled, estimated or bought. It is the same spreadsheets anyone can download, and largely nobody does.
Every time we fetch a file we keep it exactly as published, in a dated folder, even when it has not changed. Those copies are the archive, and they are what makes the second half of the corrections log possible.
The parsers refuse to guess. A file whose shape has changed — a column moved, a heading renamed, a figure where text should be — stops that dataset for the day and says so in the run log, rather than producing a number that looks fine. A dataset that stopped is simply not cited that morning.
The arithmetic on top is ours and is deliberately shallow: a division the publisher declines to do, a rank against that place’s own history, a change against the same month a year earlier. Both sides of every division are named in the sentence that uses it.
What counts as unusual
Detectors run over every place and every measure, looking for a small, fixed list of things: a record high or low, a first-ever occurrence, a threshold crossed, a streak extended or broken, an unusually large move, and a figure that has changed since the last time the publisher published it.
The thresholds are dull on purpose. Nearly every prison in England and Wales is over its certified capacity and nearly every council’s temporary-accommodation count has risen, so “over capacity” and “up on the quarter” describe the system rather than report news. Everything a detector considered and set aside is written down with the reason, so the thresholds can be argued with from evidence rather than from taste.
Which finding leads
One agent ranks the morning’s findings on four things: how big the number is, how unusual it is, what it means for people, and whether anyone else would find it today. The last of those counts for a lot. A national headline everyone already has scores low; something visible only to whoever kept the last eleven versions of a spreadsheet scores high.
It picks one lead and up to three shorter items, and writes down what won, what lost and why. Losing findings are annotated back onto the story they belong to rather than deleted, so a thread that becomes a story in three months has its own history attached.
How a claim is checked
Nothing is published because it seems right. Every claim goes to a second agent, running in a fresh session, which is given the sentence, the figures behind it, an independent recalculation from the underlying data, and the publisher’s own notes about the dataset — and nothing else. It does not see which claim was chosen to lead, why it was chosen, or any of the reasoning behind it. Its instructions are not to check the claim. Its instructions are to kill it.
It works through the same seven ways a figure like this usually goes wrong:
- The wrong column. Terminal passengers or total passengers? Households or people? Capacity certified or capacity possible?
- The wrong denominator. A percentage is only as good as what it was divided by, and the flattering divisor is always to hand.
- A series break. A renamed place, a moved weather station, a changed count day, a methodology revision. A comparison that straddles one is not a comparison.
- A definition change. “A record since 1853” asserts that the thing being measured meant the same in 1853. Often it did not, and sometimes nobody can know.
- A revision waiting to happen. A record set by a hair on a figure the publisher still calls provisional is a record that may not exist next month. It can be published — but it has to say so.
- The arithmetic. The recalculation is written independently of the code that found the claim, on purpose. If the two disagree, the claim is dead until someone can explain the difference.
- Overclaiming. Read as an unsympathetic reader would read it, does the sentence promise more than the file supports?
A claim that does not survive is cut. If the lead is cut, the next item is promoted and put through the same pass. An edition that emerges with nothing left is a thin day, and a thin day is published as one.
This is not decorative. On its first morning the pass killed a claim that a prison was “90% full” — the figure had been divided by the wrong capacity measure, and the true number was 129%. It also caught two faults that were ours rather than planted: a record sentence that omitted the word provisional, and the phrase “space designed for” used of a capacity figure that means something else. Both wordings were changed everywhere they appear.
Every edition carries its own record of this: how this edition was checked lists each claim, the vintage of the data it was verified against, and the publisher’s file it came from.
Vintages, revisions and corrections
Official statistics are not fixed. Provisional figures are finalised, late returns arrive, methods are restated. Because we keep every version of every file we have fetched, we can compare today’s release with the one before it, value by value, and see what moved in a period the publisher had already published.
Those go on the corrections log, alongside our own mistakes, dated and permanent. A restated figure is not a scandal — it is how statistics work — but it should be visible, and mostly it is not.
Findings written on one day are re-checked against current data before they are published, and the version of the data a claim was verified against travels with it.
Every figure links to its source
No number appears anywhere on this site or in the email without a link to the file it was read from. From any figure you should be able to reach the publisher’s original in two steps: the claim links to that place’s page, and the page links to the file.
The parsed series are published too, as plain CSV files at fixed addresses — no key, no account, no limit. They are the same files the pages are built from, so if a figure here is wrong, the file it is wrong in is downloadable.
What this does not do
It does not adjust, smooth, model or forecast anything. It does not have a view on policy: a prison at 172% of its certified capacity is reported as arithmetic, and what should be done about it is somebody else’s sentence. It does not chase a number when there is not one — on a quiet day the email says so in two lines, which is the whole reason to trust it on a loud one. And where a finding cannot be phrased with dignity, particularly on homelessness, it does not lead.
Reaching a human
joshplawman@gmail.com reaches a person, and it is also the reply address on the daily email. Corrections, disagreements about a definition, requests for a dataset: all welcome, all read. A person audits the editorial log every week, spot-checks the verifications and reads the corrections; that audit is the human in the loop, and it happens after publication rather than before it.
Public Figures is built by Adder. There is more about the project on the about page.