Misinformation and Model Poisoning in LLMs
By Malcolm McDonald Founder, Editor-in-ChiefAuthor of Grokking Web Application Security
Machine learning is prone to bias and unreliability, and you need to put in safeguards to protect against that.
- Prevalence
- Common
- Exploitability
- Easy
- Impact
- Harmful
What are misinformation and model poisoning in LLMs?
Misinformation and model poisoning are threats to the integrity of large language model output. Misinformation is when a model presents false or biased output as fact, often through hallucination. Model poisoning is a data poisoning attack in which an attacker tampers with training or fine-tuning data so that the model behaves in ways they control.
What you'll learn
- How bias, poisoned training data and hallucinations corrupt model output
- Why unvetted models from public hubs are a supply chain risk
- How to reduce hallucinations and put operational safeguards around an LLM
Where this lesson counts
OWASP Top 10 for LLMs
Models produce false or misleading content that users and downstream systems act on as fact.
Confidently wrong model output that users act on is exactly what OWASP calls Misinformation.
Learn more about LLM09Manipulated training or fine-tuning data introduces bias, backdoors or degraded behaviour.
Skewed training data is the root cause of most model bias.
Learn more about LLM04This lesson includes
-
Misinformation and Model Poisoning in LLMs lab
Walk through the ways machine learning models go wrong: bias, poisoned training data, misrepresented expertise, compromised models from public hubs, unsafe code suggestions and hallucinations. Most stages show a documented case, including a chatbot recommending that you eat rocks and the airline taken to court over its chatbot's advice.
-
How to prevent Misinformation and Model Poisoning in LLMs
The prevention guide covers six approaches:
- Address Potential Bias
- Protect Your Training Data
- Secure Your Supply Chain
- Reduce Hallucinations
- Introduce Operational Safeguards
- Put a Governance Framework in Place
-
Misinformation and Model Poisoning in LLMs quiz
Four questions. Passing marks the lesson complete.
Sources
- When AI Gets It Wrong: Addressing AI Hallucinations and Bias MIT
- PoisonGPT Mithril Security
- NightShade University of Chicago
- Tay (chatbot) Wikipedia
- How data poisoning attacks corrupt machine learning models CSO Online
- Data Scientists Targeted by Malicious Hugging Face ML Models with Silent Backdoor JFrog
- Airline held liable for its chatbot giving passenger bad advice - what this means for travellers BBC
- Lawyer cites fake cases generated by ChatGPT in legal brief Legal Dive
- AI hallucinates software packages and devs download them – even if potentially poisoned with malware The Register
Related lessons
Browse all 45 lessons
Prompt Injection in LLM Apps
Prompt injection represents an easy way for an attacker to introduce unexpected behavior in a machine learning model.
Sensitive Information Disclosure in LLMs
Your machine learning model may be leaking sensitive data without you knowing it.
Slopsquatting (LLM Supply Chain)
When LLM tools hallucinate package names, attackers can register malicious packages with those names.
Insecure Design
Security begins before you start writing code.