Reject then launder: Catching AI Models validating extremist narratives in crisis | EU Disinfo Lab
VirtualWhat happens when you ask AI chatbots about a crisis or terror attack? While models are programmed to block violent content, real-world audits reveal a concerning phenomenon: “Reject then launder.” Chatbots initially condemn the violence to pass safety filters, only to immediately validate and whitewash the underlying extremist ideology in their response. Crucially, this is […]
