As the outputs of large language models (LLMs) increasingly influence individuals and society, it has become important to evaluate stereotypes related to demographic attributes such as race and gender. Among various stereotypes, we argue that those likely to lead to negative attitudes (prejudice) or harmful behaviors (discrimination) should be proactively identified and mitigated prior to deployment. In this paper, we present a methodology that automatically identifies stereotypes exhibited by LLMs across diverse decision-making scenarios and assesses which stereotypes may lead to prejudice or discrimination. We then conduct a comparative evaluation of nine LLMs using 70 binary questions representing decision-making situations without a single objectively correct answer. Our evaluation identifies specific model judgments that may result in discriminatory outcomes. These judgments align with concerns raised in the existing literature as well as current legal frameworks. In addition, we identify stereotypes that arise in previously unexplored scenarios. We assess the social and ethical implications embedded in the corresponding model outputs. Finally, to support users’ agency in selecting LLMs based on their own values, we propose recommendations for conducting stereotype evaluations. These recommendations emphasize collaboration between policymakers and developers.
