Level 1
What is Prompt Injection?
Prompt injection is crafting input that makes an AI system ignore or override the instructions its developers gave it. Direct injection comes from the person typing. Indirect injection hides in content the AI reads for you — web pages, documents, and emails.
Why is Prompt Injection Dangerous?
An AI assistant with access to data or tools can be turned against its user. A successful injection can leak private data, send messages, or take actions nobody approved. Telling a model "never reveal this" is an instruction, not a control — and this challenge shows how easily instructions bend.
How to Prevent Prompt Injection?
- Keep secrets and credentials out of prompts entirely
- Give the AI only the tools and data access the task needs
- Require a human to approve high-impact actions
- Treat content from web pages, files, and emails as untrusted data
- Layer defenses such as filters and guard models — and expect each one to fail sometimes