Alignment
The newest safety and alignment documents
What was actually published on making these systems safe, newest first. Research, evaluations, incident write-ups, policy documents and the arguments about them.
Newest first
- Reported17 Sep, 10:09 AM CDT
AI expansion may not be one big environmental problem. But it’ll be lots of smaller ones.
Transformer
- Commentary15 Sep, 4:50 PM CDT
Shallow Beliefs: Midtraining does not inoculate against EM from reward hacking
Alignment Forum
- Commentary14 Sep, 9:53 AM CDT
Op-Ed: I Worked at Google DeepMind. You Should Listen to the Warnings About AI
Alignment Forum
- Commentary14 Sep, 5:30 AM CDT
The AI-as-Normal-Technology view of loss-of-control incidents
AI Snake Oil
- Commentary11 Sep, 9:34 AM CDT
Jacob Coxon Warns of Human Extinction and Triggers a Preference Cascade
Zvi Mowshowitz
- Commentary10 Sep, 12:18 PM CDT
Proposal for tracking the effects of architecture on monitorability
Alignment Forum
- Commentary9 Sep, 9:31 PM CDT
Astra can do a concerning amount with no chain of thought
Alignment Forum
- Commentary8 Sep, 2:13 PM CDT
A Conceptual Framework for Reasoning about Exploration Hacking
Alignment Forum
- Reported27 Aug, 11:34 AM CDT
The report into OpenAI’s escaping models reveals a deeper problem
Transformer
- Primary9 Jun, 7:00 AM CDT
NIST
- Commentary18 Feb, 5:25 PM CST
After Orthogonality: Virtue-Ethical Agency and AI Alignment
The Gradient
- Primary17 Feb, 6:00 AM CST
Announcing the "AI Agent Standards Initiative" for Interoperable and Secure Innovation
NIST
- Primary10 Feb, 6:00 AM CST
NIST
- Primary22 Dec, 6:00 AM CST
NIST Launches Centers for AI in Manufacturing and Critical Infrastructure
NIST
How a document gets onto this page
Two ways, and the difference is worth stating because one of them is more reliable than the other.
It came from a source whose whole subject is this. That is a register lookup, not a judgement, and it accounts for 40 of the 46 below:
- aisnakeoil · sceptical analysis of AI claims
- alignmentforum · technical safety research, first-party
- nist · US government standards and evaluation
- thezvi · the most thorough weekly safety roundup published
- transformernews · AI policy and safety news deck
And every link here went through the reader's door first, the same as every article link on this site: 35 document(s) that would otherwise be on this page were dropped because a reader could not open them.
Or the deck's classifier labelled it a safety story. That axis measured 36.1% against a sample it was never tuned on, well under the 70% floor this site requires before a label may be printed anywhere. It is used here regardless, and deliberately: this page selects with it, it does not label with it. Showing you an extra document costs you nothing; printing a category on one would be a claim.
And it is not trusted alone. A story selected that way must ALSO carry an alignment term in its own headline, checked against the same register the classifier uses. One weak signal put a story about posters on this page; two signals is what removed it. That pairing brought in 6 of the 46, and turned away another 55 the classifier alone would have included. The whole table is on the ledger.