Intersectional Fairness in Large Language Models
Large Language Models (LLMs) are increasingly deployed in socially sensitive settings, raising concerns about fairness and bias, particularly when multiple sensitive attributes intersect. We systematically evaluate intersectional fairness in six LLMs using two datasets from the Bias Benchmark for Question Answering (BBQ), combining race with gender and socioeconomic status. We assess bias, subgroup fairness, accuracy, and consistency across contexts and question polarities. While models perform well in ambiguous contexts, sparse non-unknown predictions limit fairness evaluation. In disambiguated contexts, stereotype alignment affects accuracy differently across datasets: models favor stereotype-reinforcing items in Race-SES but counter-stereotype items in Race-Gender for five of six models. Subgroup fairness metrics reveal uneven outcome distributions despite low observed disparities in some cases, while repeated runs show variable response consistency. Overall, no model consistently achieves both fairness and reliability, highlighting the need to evaluate bias, subgroup fairness, and consistency across intersectional contexts.