Content moderation is critical to protecting users against harmful and illegal online content, although it may also restrict free expression. Over the past two decades, platforms have increasingly turned to automation to perform content moderation at scale. First through hashing algorithms and later through predictive machine learning methods, many large platforms have automated over 90% of content moderation actions. Today, online platforms are increasingly exploring using large language models (LLMs), as a new component of the content moderation technology stack. With online platforms conducting mass layoffs of trust and safety teams, while also increasing investments in their AI development, this convening will explore how LLMs fit in the content moderation landscape and what that means for platforms, policymakers, and users.
This convening brings together researchers, technologists, and practitioners to examine what this third wave of automated content moderation means in practice. We will explore how LLMs are being deployed by platforms today, how they compare to prior automated systems, and what the emerging research in this space tells us. We will hear from experts who are building their own custom systems using LLMs for content moderation and researchers who study their implications for information integrity. The conversation will also examine the policy environment shaping adoption: regulatory frameworks in the EU, Brazil and India, are setting new expectations for the speed and transparency of content moderation decisions, incentivizing further automation by online platforms. We will share a pre-read working paper on this topic in advance of this event. [Register at this link]