arXiv · 2606.24022
Do Language Models Pass the Bechdel Test? Auditing Gender Biases in LLM-Generated Screenplays
Abstract
As large language models (LLMs) are increasingly used in media production from journalistm to filmmaking, what impact do they have on the stories being told? Prior work has shown LLMs to perpetuate social biases, including those related to gender. We complement existing literature on gender bias in LLM outputs by auditing the network structure of LLM-generated movie screenplays through automating the Bechdel test, a popular measure of women's representation in literary and film works. We also introduce the use of social network analysis measures to further analyze representational bias in LLM-generated scripts. We evaluate screenplays generated by three state-of-the-art LLMs (GPT-5, Gemini 3 Pro, and Claude Sonnet 4.5) against 768 corresponding human-written screenplays, finding that human-written scripts are more likely to pass the Bechdel test. However, other network analyses, like centrality, homophily, and triadic relationships demonstrate that in some cases LLM-scripts have less bias, although all script types demonstrate some representational bias under most measures. We conclude by discussing the continued need for further quantitative assessments of media representations and AI-generated content.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Megha N. Govindu, Stephanie T. Wang, Sorelle A. Friedler, Danaé Metaxa. 2026-06-23. Do Language Models Pass the Bechdel Test? Auditing Gender Biases in LLM-Generated Screenplays. https://doi.org/10.1145/3805689.3812208
Cite the original work for its findings. Save a collection to share your selection of sources.