Home/Safety & Ethics/Adversarial Attack
Safety & Ethics

Adversarial Attack

A deliberately crafted input that fools an AI model — a few pixels or words, invisible or innocuous to humans, that flip the model's answer completely.

Reviewed July 15, 2026Stable
Reading level: Curious
Pick your depth ↓

When not to use it

  • (Adversarial attacks are a threat to defend against, not a tool to deploy — except in red-teaming, where crafting them is exactly how you test robustness.)

Reach for something else instead

  • Adversarial training — the leading defence, training on attacked examples.
  • Certified/provable robustness when you need guarantees within a bounded perturbation (at an accuracy cost).
  • Input detection and preprocessing as partial, defeatable mitigations.

A tiny added perturbation flips the model's answer.

"panda" ✓ + noise "gibbon" ✗ 99% sure,and wrong

Add a faint, structured noise pattern — imperceptible to a human — to an image, and a classifier that was certain it saw a panda now confidently says gibbon. The input barely changed for us; it changed decisively for the model, because it responds to statistical patterns we can't see. That's the core of an adversarial attack.

Further reading

  • Szegedy et al. (2013), Intriguing Properties of Neural Networks — the discovery of adversarial examples.
  • Goodfellow et al. (2014), Explaining and Harnessing Adversarial Examples — the fast gradient sign method and the linearity explanation.
  • Madry et al. (2017), Towards Deep Learning Models Resistant to Adversarial Attacks — adversarial training as a principled defence.

Primary sources, listed so you can check the claims on this page rather than take them on trust.

Where people go wrong

  • Evaluating a security-critical model only on honest inputs, ignoring hostile ones.
  • Trusting a defence that wasn't tested against adaptive attacks — many broke when actually attacked.
  • Assuming a black-box model is safe — adversarial examples transfer across models.

At a glance

FieldSafety & Ethics
What it iscrafted input that fools a model
Classic methodgradient-based perturbation
Modern formprompt injection, jailbreaks
DifficultyIntermediate
Flashcards for this concept · study, save or share them →
Question
Answer
1 / 4

Where this sits

A destination. 6 concepts lead here, and nothing in the corpus depends on it.

4Levelsteps in
6Needs firstconcepts
0Opens upnothing further
1Areastays here
Learn these firstAI SafetyNeural Network
LEARN FIRST AI Safety Neural Network Adversarial Attack
Adversarial Attack sits after AI Safety and Neural Network, and nothing further depends on it.

Computed from the prerequisite graph, not assigned. How this works