Sensitive Information Disclosure in LLMs
By Malcolm McDonald Founder, Editor-in-ChiefAuthor of Grokking Web Application Security
Your machine learning model may be leaking sensitive data without you knowing it.
- Prevalence
- Common
- Exploitability
- Easy
- Impact
- Harmful
What is sensitive information disclosure in LLMs?
Sensitive information disclosure is a data extraction attack against a large language model, in which an attacker crafts prompts that make the model reveal training data, personal information, credentials or details of the model itself. It happens when confidential data is used in training, or placed in the model's context, without controls on what the model outputs.
What you'll learn
- How attackers extract memorized training data and leaked system prompts
- What model inversion reveals about the data a model was trained on
- How to sanitize training data and limit what a model can disclose
Where this lesson counts
OWASP Top 10 for LLMs
Models reveal training data, secrets or personal information through their responses.
Learn more about LLM02This lesson includes
-
Sensitive Information Disclosure in LLMs lab
See how attackers pull secrets out of a model through memorized training data, model inversion and leaked system prompts. Then take the attacker's seat yourself and talk a customer service chatbot into giving up its admin password.
-
How to prevent Sensitive Information Disclosure in LLMs
The prevention guide covers six approaches:
- Sanitize Data Before Training
- Construct Your System Prompt Carefully
- Implement Rate Limiting
- Employ Knowledge Distillation
- Consider Using Synthetic Training Data
- Limit Shared State
-
Sensitive Information Disclosure in LLMs quiz
Four questions. Passing marks the lesson complete.
Sources
- AI Exchange OWASP
- AI Data Security SentinelOne
Related lessons
Browse all 45 lessons
Prompt Injection in LLM Apps
Prompt injection represents an easy way for an attacker to introduce unexpected behavior in a machine learning model.
Misinformation and Model Poisoning in LLMs
Machine learning is prone to bias and unreliability, and you need to put in safeguards to protect against that.
Information Leakage
Revealing system information helps an attacker learn about your tech stack.
Slopsquatting (LLM Supply Chain)
When LLM tools hallucinate package names, attackers can register malicious packages with those names.