arXiv · 2411.08884
Quantifying Risk Propensities of Large Language Models: Ethical Focus and Bias Detection through Role-Play
Abstract
As Large Language Models (LLMs) become more prevalent, concerns about their safety, ethics, and potential biases have risen. Systematically evaluating LLMs' risk decision-making tendencies and attitudes, particularly in the ethical domain, has become crucial. This study innovatively applies the Domain-Specific Risk-Taking (DOSPERT) scale from cognitive science to LLMs and proposes a novel Ethical Decision-Making Risk Attitude Scale (EDRAS) to assess LLMs' ethical risk attitudes in depth. We further propose a novel approach integrating risk scales and role-playing to quantitatively evaluate systematic biases in LLMs. Through systematic evaluation and analysis of multiple mainstream LLMs, we assessed the "risk personalities" of LLMs across multiple domains, with a particular focus on the ethical domain, and revealed and quantified LLMs' systematic biases towards different groups. This research helps understand LLMs' risk decision-making and ensure their safe and reliable application. Our approach provides a tool for identifying and mitigating biases, contributing to fairer and more trustworthy AI systems. The code and data are available.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yifan Zeng, Liang Kairong, Fangzhou Dong, Peijia Zheng. 2024-10-26. Quantifying Risk Propensities of Large Language Models: Ethical Focus and Bias Detection through Role-Play. https://arxiv.org/abs/2411.08884
Cite the original work for its findings. Save a collection to share your selection of sources.