The Ledger

What this deck could not tell you

Every news site publishes its confident half. This is the other one: the links refused, the stories it could not classify, the sources that went quiet, and the measured accuracy of its own instruments. Every number below was written by the pipeline that produced the front page, on the same run.

sources checked 21 Sep, 2:42 PM CDTwhere we pull fromhow it works

The reader's door

A story whose link does not open for a reader does not go on the page. The deck reads the publisher's own machine-readable access declaration, obeys robots.txt, and stores no article text. A link it cannot judge is kept, never quietly dropped: excluding on ignorance would delete good journalism while looking like diligence.

566

links checked

35

refused as walled

paywall, registration or dead

191

kept but unknown

no declaration, or a block we cannot read

13

dropped this run

before clustering

Of 566 links checked, 35 were refused and 191 could not be judged at all. The unjudged are on the site.

The instrument we are not showing you

The deck classifies every story it holds on several axes: whether a piece reports or argues, what kind of event it is, what field it touches, and whether the thing has actually happened. Each axis must clear a published accuracy floor, on a sample it was never tuned against, before a single label of it appears anywhere on this site. One of the four has.

axiscorrectagreement, against a 70% floorverdict
form54 / 6188.5%published
story type22 / 6136.1%held back
domain37 / 6160.7%held back
horizon17 / 4141.5%held back

Why the number is this low, stated rather than buried. Tuning the word lists against a first hand-labelled sample lifted it to 59, 70 and 69 percent. Scored again on a second sample the tuning had never seen, it came back at the figures above. That gap is not noise, it is what overfitting looks like from the inside: real improvement, and a number that cannot tell you so. The first figures are the ones most dashboards would have printed.

It also fails in a specific, honest direction. On 0 of 566 stories it declines to guess at all rather than picking the likeliest label. That abstention is the biggest single source of its error and it is the behaviour we would choose.

And the check on the one axis that passed. A source's form was set from these same hand labels for most sources, so for those rows the test is not independent: it measures whether the register agrees with the labels it was built from. Split apart:

how the source's form was setcorrectaccuracy
fitted to these labels48 / 5390.6%
judged independently6 / 875.0%

The independent row is the honest one, and its sample is small enough that the interval around it is wide. Both are here because printing only the flattering one would have been the easy thing to do.

Who made the labels, which changes what every number above means. They were made by an AI model reading headlines, and that same model wrote the classifier being scored. So these are not accuracies against ground truth, they are agreement with one judge marking its own work. The column is named agreement for that reason. Making them real needs a second reader labelling the same stories, which would give three figures instead of one: the machine against that reader, the machine against the first, and the two readers against each other. That last one is the ceiling, and nobody publishing a classifier accuracy usually tells you where theirs is.

Carried by nobody else

A primary source, reporting work somebody did rather than announced, that no other outlet in this deck's 56-source register picked up. The deck already measures how many outlets ran a story; this is the same number read the other way round.

14 of 566 stories in the store meet every condition. The picking uses the classifier above, which has not cleared its floor, and that difference is deliberate: choosing which story to show you with a weak instrument costs you nothing, while putting a label on it would be a claim. So there are no labels here.

What went quiet

A source that stops answering is easy to miss, because the page still fills with what it said last time. The deck fails a run rather than print a green line over a layer with no live source.