arXiv · 2609.04373
Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets
Abstract
Large language models (LLMs) are being deployed at scale in consequential real-world systems, from financial markets to content moderation to hiring. We show that improving individual model capability can degrade rather than improve system-level outcomes. We hypothesize that shared training and architectures can lead more capable LLMs to behave more similarly, creating correlated actions that do not diversify away. We develop a general framework showing how this correlation creates a non-diversifiable risk floor and test its predictions in financial markets using an agent-based simulation with LLM traders of varying general-purpose capability. We find that: (1) frontier LLMs exhibit significantly correlated behavior that increases with capability; (2) when their shared reasoning is accurate, increasing agent participation reduces market-level risk; and (3) when agents share a common misinformation environment, the same correlated behavior becomes a liability. Together, these results identify a capability paradox: improving individual models does not necessarily produce better system-level outcomes. Whether the same dynamics arise in other domains is an open empirical question.
Explore related subjects
Keep this discovery
Jillian Ross, Eric So, Zoe De Simone, Charles Pozniak, Andrew W. Lo. 2026-09-03. Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets. https://arxiv.org/abs/2609.04373
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.