Research · AI Security
Securing AI Systems: Prompt Injection and Data Poisoning
AI systems can be attacked through their inputs, their training data and the tools they can reach. Securing them means extending familiar security practice to new attack paths such as prompt injection and data poisoning.
Key threats
- Prompt injection: instructions hidden in user input, documents, emails or web pages that cause a model to ignore its rules, leak data or misuse tools.
- Data poisoning: manipulated training or retrieval data that corrupts model behavior or plants hidden triggers.
- Sensitive data disclosure: models revealing confidential information from prompts, context or training data.
- Excessive agency: models or agents with more permissions than their task requires.
- Model theft and abuse: extraction of models or misuse of AI endpoints at scale.
Useful references
The OWASP Top 10 for Large Language Model Applications describes the most common risks in LLM-based applications. MITRE ATLAS catalogs adversary tactics and techniques against AI systems, modeled on the MITRE ATT&CK approach. NIST has also published guidance on adversarial machine learning.
Controls that help
- Treat all model inputs, including retrieved documents, as untrusted.
- Separate system instructions from user and external content where possible.
- Grant models and agents least-privilege access and require approval for sensitive actions.
- Filter and monitor outputs for sensitive data and policy violations.
- Validate and track the provenance of training and retrieval data.
- Red team AI applications before launch and after significant changes.
Common pitfalls
- Assuming the model provider handles all AI security.
- Relying on prompt wording alone to enforce security rules.
- Connecting AI to internal systems without threat modeling.
How to get started
- Inventory AI applications and what data and tools each can reach.
- Add AI-specific threats to security design reviews.
- Test high-risk applications for prompt injection.
- Log prompts, tool calls and outputs for investigation.
Questions leaders should ask
- Have our AI applications been tested for prompt injection?
- What is the worst action an attacker could make our AI agent take?
- Can our AI applications leak data from other users or systems?
- Do we log enough to investigate an AI security incident?