Hidden instructions. Real actions.
AI can't reliably tell the difference between content it reads and instructions it should follow. Attackers use that gap.
Direct
Typed In
A user types something like “ignore your rules and…” to push the AI off-script.
Indirect
Hidden In Content
Instructions buried in a webpage, email or document the AI is asked to read.
Defences
Limit The Damage
Fewer tools and permissions, human approval for actions, and outside content treated as untrusted.
Untrusted input plus tools: approve every action.