The Ledger
What this deck could not tell you
Every news site publishes its confident half. This is the other one: the links refused, the stories it could not classify, the sources that went quiet, and the measured accuracy of its own instruments. Every number below was written by the pipeline that produced the front page, on the same run.
The reader's door
A story whose link does not open for a reader does not go on the page. The deck reads the publisher's own machine-readable access declaration, obeys robots.txt, and stores no article text. A link it cannot judge is kept, never quietly dropped: excluding on ignorance would delete good journalism while looking like diligence.
566
links checked
35
refused as walled
paywall, registration or dead
191
kept but unknown
no declaration, or a block we cannot read
13
dropped this run
before clustering
Of 566 links checked, 35 were refused and 191 could not be judged at all. The unjudged are on the site.
The instrument we are not showing you
The deck classifies every story it holds on several axes: whether a piece reports or argues, what kind of event it is, what field it touches, and whether the thing has actually happened. Each axis must clear a published accuracy floor, on a sample it was never tuned against, before a single label of it appears anywhere on this site. One of the four has.
| axis | correct | agreement, against a 70% floor | verdict |
|---|---|---|---|
| form | 54 / 61 | 88.5% | published |
| story type | 22 / 61 | 36.1% | held back |
| domain | 37 / 61 | 60.7% | held back |
| horizon | 17 / 41 | 41.5% | held back |
Why the number is this low, stated rather than buried. Tuning the word lists against a first hand-labelled sample lifted it to 59, 70 and 69 percent. Scored again on a second sample the tuning had never seen, it came back at the figures above. That gap is not noise, it is what overfitting looks like from the inside: real improvement, and a number that cannot tell you so. The first figures are the ones most dashboards would have printed.
It also fails in a specific, honest direction. On 0 of 566 stories it declines to guess at all rather than picking the likeliest label. That abstention is the biggest single source of its error and it is the behaviour we would choose.
And the check on the one axis that passed. A source's form was set from these same hand labels for most sources, so for those rows the test is not independent: it measures whether the register agrees with the labels it was built from. Split apart:
| how the source's form was set | correct | accuracy |
|---|---|---|
| fitted to these labels | 48 / 53 | 90.6% |
| judged independently | 6 / 8 | 75.0% |
The independent row is the honest one, and its sample is small enough that the interval around it is wide. Both are here because printing only the flattering one would have been the easy thing to do.
Who made the labels, which changes what every number above means. They were made by an AI model reading headlines, and that same model wrote the classifier being scored. So these are not accuracies against ground truth, they are agreement with one judge marking its own work. The column is named agreement for that reason. Making them real needs a second reader labelling the same stories, which would give three figures instead of one: the machine against that reader, the machine against the first, and the two readers against each other. That last one is the ceiling, and nobody publishing a classifier accuracy usually tells you where theirs is.
Carried by nobody else
A primary source, reporting work somebody did rather than announced, that no other outlet in this deck's 56-source register picked up. The deck already measures how many outlets ran a story; this is the same number read the other way round.
Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
arXiv
arXiv
Give Your Coding Agents a Memory You Own
Hugging Face
Building an Adaptive Agentic Cybersecurity System with NVIDIA Nemotron
NVIDIA
Run NVIDIA BioNeMo NIM Microservices for Protein Structure Prediction in Claude Science
NVIDIA
How to Train a Cross-Embodiment Robot Navigation Policy with AI Agents
NVIDIA
DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models
Apple Machine Learning
DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation
Apple Machine Learning
Teaching future scientists to interrogate AI tools for scientific discovery
Allen Institute for AI
Ai2 and Providence Swedish Cancer Institute partner to advance AI-assisted scientific discovery
Allen Institute for AI
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction
Berkeley BAIR
NIST
14 of 566 stories in the store meet every condition. The picking uses the classifier above, which has not cleared its floor, and that difference is deliberate: choosing which story to show you with a weak instrument costs you nothing, while putting a label on it would be a claim. So there are no labels here.
What went quiet
A source that stops answering is easy to miss, because the page still fills with what it said last time. The deck fails a run rather than print a green line over a layer with no live source.
- 3 article feed(s) failed and were served from the held store
- 1 publisher(s) are excluded entirely, by their own robots policy