Social Science Research Council Research AMP Just Tech
Citation

Understanding and Mitigating Data Poisoning in LLMs

Author:
Shah, Atharva; Mathur, Ishani; Kanakia, Harshil
Year:
2025

Large Language Models (LLMs) have demonstrated remarkable performance across a wide array of natural language tasks. Still, their reliance on large training datasets makes them vulnerable to data poisoning attacks, in which attackers purposefully introduce biased or tainted data to manipulate model behavior. This paper gives an overview of data poisoning threats to LLMs, including training-time and post-training attacks. We examine real-world cases of LLM poisoning, consider potential societal impacts, and evaluate the effectiveness of various detection and prevention strategies. Our study identifies that in order to ensure the dependability and security of such powerful AI systems, the entire LLM supply chain-from data collection to model deployment-must be protected.