Safety & Ethics

AI Safety

The umbrella field concerned with making AI systems reliably do what we intend and avoid causing harm — the parent discipline over alignment, interpretability, robustness, and fairness.

Reviewed July 16, 2026Stable
Reading level: Curious
Pick your depth ↓

When not to use it

  • As a synonym for AI ethics. Safety is about systems reliably doing what is intended; ethics is about what ought to be intended. They overlap but are not the same.
  • As a purely long-term concern. Framing safety only around future systems ignores the real present harms that make up most of the actual work.
  • As something a model alone provides. Safety is a property of the whole system and its deployment, not a checkbox inside the model.

Reach for something else instead

  • AI ethics addresses the normative questions of what values a system should serve, complementing safety's focus on reliable behaviour.
  • Reliability and security engineering cover overlapping ground for conventional software and are increasingly merged with AI safety in practice.
  • AI governance works at the policy and institutional level rather than the technical one.

Sources & further reading

  • Amodei et al. (2016), Concrete Problems in AI Safety — the paper that framed practical safety research around specification, robustness, and assurance.
  • Hendrycks et al. (2022), Unsolved Problems in ML Safety — a modern map of the field's open challenges.
  • Anthropic, OpenAI, DeepMind safety teams — ongoing technical work on alignment, interpretability, and evaluation that defines the current frontier.

Primary sources, listed so you can check the claims on this page rather than take them on trust.

Where people go wrong

  • Treating safety and capability as opposites. Much safety work aims to make capable systems usable, not to hold capability back.
  • Assuming a well-behaved demo means a safe system. Safety is about the tails and the unanticipated cases, not the happy path.
  • Reducing safety to content filtering. Filters are one small part of a field that spans alignment, interpretability, robustness, and oversight.

At a glance

FieldSafety & Ethics
Scopenear-term harms to long-term control
Core toolsalignment, interpretability, evaluation, robustness
Key tensioncapability outpacing assurance
DifficultyIntermediate → Advanced
Flashcards for this concept · study, save or share them →
Question
Answer
1 / 3