15–25 min Updated

Misinformation and Model Poisoning in LLMs

By Malcolm McDonald Founder, Editor-in-ChiefAuthor of Grokking Web Application Security

Machine learning is prone to bias and unreliability, and you need to put in safeguards to protect against that.

Prevalence
Common
Exploitability
Easy
Impact
Harmful

What are misinformation and model poisoning in LLMs?

Misinformation and model poisoning are threats to the integrity of large language model output. Misinformation is when a model presents false or biased output as fact, often through hallucination. Model poisoning is a data poisoning attack in which an attacker tampers with training or fine-tuning data so that the model behaves in ways they control.

What you'll learn

  • How bias, poisoned training data and hallucinations corrupt model output
  • Why unvetted models from public hubs are a supply chain risk
  • How to reduce hallucinations and put operational safeguards around an LLM

Where this lesson counts

OWASP Top 10 for LLMs

  • Misinformation and Model Poisoning in LLMs lab

    Walk through the ways machine learning models go wrong: bias, poisoned training data, misrepresented expertise, compromised models from public hubs, unsafe code suggestions and hallucinations. Most stages show a documented case, including a chatbot recommending that you eat rocks and the airline taken to court over its chatbot's advice.

  • How to prevent Misinformation and Model Poisoning in LLMs

    The prevention guide covers six approaches:

    • Address Potential Bias
    • Protect Your Training Data
    • Secure Your Supply Chain
    • Reduce Hallucinations
    • Introduce Operational Safeguards
    • Put a Governance Framework in Place
  • Misinformation and Model Poisoning in LLMs quiz

    Four questions. Passing marks the lesson complete.

Sources