Social Science Research Council Research AMP Just Tech
Citation

Safety Theater in Anthropomorphic AI: Why We Need to Shift from Safety to Accountability Narratives

Author:
Maeda, Takuya; Buening, Rebecca
Year:
2026

Despite decades of research on anthropomorphic AI’s unique ability to create relational (interaction) harms, few legitimately effective interventions exist. Taking this as an epistemic problem, this paper explores the limitations of the “AI safety” narrative that dominates “responsible AI” discourse—a narrative that frames interaction harms as malfunctions/user errors and harm mitigation as technical optimization/disclosure. We compare this to “AI accountability” narratives that frame interaction harms as structural issues and mitigation as external regulatory oversight. Adopting an accountability narrative to chart the sociotechnical and political economic dimensions of anthropomorphic AI, we reveal what a safety narrative conceals: that anthropomorphic features are both social cues that instigate parasocial dynamics and engagement capture mechanisms that enable companies to extract more value from interactions. In this view, interaction harms are not discrete malfunctions but processual intensifications of target use cases that compound with increasing attachment (what we call “harm creep”). Because safety mitigation efforts reinforce “value-aligned” personality design, which in turn supports parasocial attachment, we suggest that safety interventions may reinforce interaction harms rather than mitigating them while externalizing responsibility to users (a context collapse condition we describe as “safety theater,” which perpetuates the existing accountability vacuum). A more appropriate solution would involve safety interventions that account for harm creep and accountability interventions that constrain extraction incentives through external monitoring and power redistribution.