Occupational Gender Bias in Large Language Models: Reproduction of Stereotypes, Paradoxical Analysis, and a Path Towards Algorithmic Justice
Abstract
Large language models (LLMs) are now used in many critical social decisions, including hiring. The social biases they carry pose a serious challenge to fairness. This paper aims to examine occupational gender bias in LLMs, specifically the paradoxical coexistence of "oversaturation of female roles" and deep-seated stereotypes. By analyzing the complexity of this bias, it aims to provide theoretical and empirical support for building effective paths to algorithmic justice governance. This research employed literature and data analysis, examining quantitative test results from the GenderBench evaluation suite and integrating recent cutting-edge academic findings on model auditing and bias mechanisms. Large-scale language models replicate and amplify occupational gender stereotypes, while the apparent increase in female roles masks the entrenchment of structural biases. To address such issues, effective governance must go beyond simple technical tuning. This article proposes building a collaborative governance framework that includes standardized audits, mandatory manual reviews, and multi-party participation to achieve true algorithmic justice.
References
- Mirza, I., Jafari, A. A., Ozcinar, C., & Anbarjafari, G. (2025). Quantifying Gender Bias in Large Language Models Using Information-Theoretic and Statistical Analysis. Information, 16(5), 358.
- Chen, E., Zhan, R.-J., Lin, Y.-B., & Chen, H.-H. (2025). More Women, Same Stereotypes: Unpacking the Gender Bias Paradox in Large Language Models. arXiv preprint arXiv: 2503.15904.
- Kotek, H., Dockum, R., & Sun, D. Q. (2023). Gender bias and stereotypes in Large Language Models. In Proceedings of the 2023 ACM Collective Intelligence Conference (CI '23). Association for Computing Machinery.
- Salinas, A., Haim, A., & Nyarko, J. (2025). What's in a Name? Auditing Large Language Models for Race and Gender Bias. arXiv preprint arXiv: 2402.14875.
- Pikuliak, M. (2025). Gender Bench: Evaluation Suite for Gender Biases in LLMs. arXiv preprint arXiv: 2505.12054.
- Chaturvedi, S., & Chaturvedi, R. (2025). Who Gets the Callback? Generative AI and Gender Bias. arXiv: 2504.21400.
- Kong, H., Lee, S., Ahn, Y., & Maeng, Y. (2024). Gender Bias in LLM-generated Interview Responses. arXiv preprint arXiv: 2410.20739.
- Wilson, K., & Caliskan, A. (2024). Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval. arXiv: 2407.20371.
- Xu, J., Li, G., & Jiang, J. Y. (2025). AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights. arXiv: 2509.00462.
- Cyberspace Administration of China, et al. (2023). Interim Measures for the Management of Generative Artificial Intelligence Services. Available: https: //www.gov.cn/zhengce/zhengceku/202307/content_6891752.htm
- Panarese, P. (2025). Algorithmic bias, fairness, and inclusivity: a multilevel framework for justice-oriented AI. AI & SOCIETY. doi: 10.1007/s00146-025-02451-2.
- Klein, L., & D'Ignazio, C. (2024). Data feminism for AI. In Extended Abstracts of the 2024 CHI Conference on Human Factors in Computing Systems. ACM. doi: 10.1145/3630106.3658543.
- Wood, A. (2022). Disambiguating algorithmic bias: From neutrality to justice. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society. ACM. doi: 10.1145/3600211.3604695.
- Markelius, A. (2024). An Empirical Design Justice Approach to Identifying Ethical Considerations in the Intersection of Large Language Models and Social Robotics. arXiv: 2406.06400.