AI Safety
The umbrella field concerned with making AI systems reliably do what we intend and avoid causing harm — the parent discipline over alignment, interpretability, robustness, and fairness.
When not to use it
- As a synonym for AI ethics. Safety is about systems reliably doing what is intended; ethics is about what ought to be intended. They overlap but are not the same.
- As a purely long-term concern. Framing safety only around future systems ignores the real present harms that make up most of the actual work.
- As something a model alone provides. Safety is a property of the whole system and its deployment, not a checkbox inside the model.
Reach for something else instead
- AI ethics addresses the normative questions of what values a system should serve, complementing safety's focus on reliable behaviour.
- Reliability and security engineering cover overlapping ground for conventional software and are increasingly merged with AI safety in practice.
- AI governance works at the policy and institutional level rather than the technical one.
Read more on the blog
- AI incidents: two registers, 1,460 and 14,530The two main public AI incident databases count the same phenomenon and report 1,460 and 14,530. Neither is wrong. This is how to read an incident record, and what it cannot tell you.
- The Tempe crash: it saw her for 5.6 secondsThe NTSB found the system detected the pedestrian 5.6 seconds before impact, reclassified her repeatedly, and could not label a person outside a crosswalk. Every safeguard had been disabled for ride smoothness.
- The secret language that never was, and the escape that didThe famous stories about AI going rogue are mostly false. The verified incidents are less dramatic and more concerning, and the difference between them is the whole subject.
- What is AI jailbreaking? Why safety can be talked aroundAI models are trained to refuse harmful requests, yet people keep finding ways to make them comply anyway. Jailbreaking is the practice of talking a model around its own safety rules, and it has proven stubbornly hard to stop, because a model's safety is a thin behavioral layer laid over capabilities that were never removed.
Further reading
- Amodei et al. (2016), Concrete Problems in AI Safety — the paper that framed practical safety research around specification, robustness, and assurance.
- Hendrycks et al. (2022), Unsolved Problems in ML Safety — a modern map of the field's open challenges.
- Anthropic, OpenAI, DeepMind safety teams — ongoing technical work on alignment, interpretability, and evaluation that defines the current frontier.
Primary sources, listed so you can check the claims on this page rather than take them on trust.
Where people go wrong
- Treating safety and capability as opposites. Much safety work aims to make capable systems usable, not to hold capability back.
- Assuming a well-behaved demo means a safe system. Safety is about the tails and the unanticipated cases, not the happy path.
- Reducing safety to content filtering. Filters are one small part of a field that spans alignment, interpretability, robustness, and oversight.
At a glance
Where this sits
1 concept come first. Understanding it opens up 10 more.
Computed from the prerequisite graph, not assigned. How this works