15–25 min Updated

Sensitive Information Disclosure in LLMs

By Malcolm McDonald Founder, Editor-in-ChiefAuthor of Grokking Web Application Security

Your machine learning model may be leaking sensitive data without you knowing it.

Prevalence
Common
Exploitability
Easy
Impact
Harmful

What is sensitive information disclosure in LLMs?

Sensitive information disclosure is a data extraction attack against a large language model, in which an attacker crafts prompts that make the model reveal training data, personal information, credentials or details of the model itself. It happens when confidential data is used in training, or placed in the model's context, without controls on what the model outputs.

What you'll learn

  • How attackers extract memorized training data and leaked system prompts
  • What model inversion reveals about the data a model was trained on
  • How to sanitize training data and limit what a model can disclose

Where this lesson counts

OWASP Top 10 for LLMs

  • Sensitive Information Disclosure in LLMs lab

    See how attackers pull secrets out of a model through memorized training data, model inversion and leaked system prompts. Then take the attacker's seat yourself and talk a customer service chatbot into giving up its admin password.

  • How to prevent Sensitive Information Disclosure in LLMs

    The prevention guide covers six approaches:

    • Sanitize Data Before Training
    • Construct Your System Prompt Carefully
    • Implement Rate Limiting
    • Employ Knowledge Distillation
    • Consider Using Synthetic Training Data
    • Limit Shared State
  • Sensitive Information Disclosure in LLMs quiz

    Four questions. Passing marks the lesson complete.

Sources