arXiv · 2609.23999
Misaligned Clinical Risk Classification and Cost Asymmetry in Open-Weight Large Language Models
Abstract
How large language models (LLMs) integrate patient risk with clinical cost tradeoffs remains poorly understood. We investigated how four open-weight LLMs (Qwen-2.5-7B/32B and Llama-3.1-8B/70B) internally represent cost tradeoffs, how these representations relate to clinical predictions, and whether decisions shift as predicted by the specified cost direction and magnitude. Using a public diabetes dataset, we varied 11 false-negative (FN) to false-positive (FP) cost ratios across three phrasings and examined representations and behavioral outputs. Patient risk was linearly recoverable on par with conventional classifiers (AUC $\approx 0.83$), and cost direction was recoverable in every model. However, representational shifts in cost direction tracked output changes only in the two larger models, and responses to cost magnitude were predominantly direction-agnostic. Only 2 of 12 model-phrasings showed both opposing responses to increasing FN versus FP costs and cost-correct ordering. Representationally, a direction fitted on one cost side did not invert when transferred to the other, as expected under mirror-symmetric encoding. These findings suggest that LLMs encode risk and cost information but do not reliably integrate them into cost-correct decisions. Clinical evaluations should therefore include tradeoff tests, phrasing sensitivity, and default operating points alongside predictive performance.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Star S. D. Liu, Xiyu Ding, Robert B. Barrett, Alberto Santamaria-Pang, Nic Dobbins, Harold P. Lehmann. 2026-09-21. Misaligned Clinical Risk Classification and Cost Asymmetry in Open-Weight Large Language Models. https://arxiv.org/abs/2609.23999
Cite the original work for its findings. Save a collection to share your selection of sources.