Social Science Research Council Research AMP Just Tech
Citation

When Explanations Deceive: Understanding Unintentional and Intentional Deception in XAI

Author:
Tekkesinoglu, Sule
Year:
2026

Explainable Artificial Intelligence (XAI) is widely promoted as a foundation for trustworthy AI, with explanations expected to ensure accountability. Yet explanations can provide incomplete information, conceal critical limitations, or present findings in misleading ways. This paper conceptualises deceptive explanations in XAI as a socio-technical phenomenon arising not only from malicious intent but also from the interaction of technical limitations, human cognitive biases, and organisational priorities. A taxonomy is developed to distinguish unintentional deception (linked to method-specific and human factors) from intentional deception (realised through fabrication, concealment, and manipulation). Moving beyond this binary distinction, this work introduces a two-dimensional framework capturing a critical grey area defined by degrees of intent and action, where known limitations of XAI methods are wilfully ignored or opportunistically exploited rather than accidentally incurred. This exploitation erodes trust in explanation-based accountability, making it difficult for practitioners, auditors, and regulators to distinguish strategic misuse from good-faith method selection. The paper concludes by outlining differentiated strategies for detecting deceptive explanations, mitigating unintentional deception, and preventing intentional exploitation. This contribution challenges the role of explainability in accountability and highlights the need for governance frameworks robust to the full spectrum of deceptive explainability.