arXiv ScienceSearch

arXiv · 2609.01660

How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Making

Abstract

Production deployments of large language model (LLM) agents remain unreliable on long, multi-step workflows even as benchmark success rates climb steadily. We argue this gap is largely an artifact of task horizon: benchmarks are dominated by short-to-medium horizons where success remains high, while production workloads demand an order of magnitude more dependent steps. We measure the effect directly, characterizing the shape of agent degradation and disentangling its cause across a large controlled study spanning nine models, six open models from 1.2B to 671B parameters, and three deployed proprietary systems; four task families, including a genuinely agentic tool-use loop; five horizons; and three context regimes. Task success follows a geometric law governed by a single per-step reliability parameter, which rises with model scale but saturates well below 1 even for the strongest models, guaranteeing eventual collapse at sufficiently long horizons. The effect is sharpest on the agentic task, where every model tested, including widely deployed systems, falls from near-perfect success to near zero within sixteen steps of (n=10,664 analyzed trajectories. Degradation is driven by step count rather than context length: bounding the context window steepens decay rather than easing it (logit slope -0.69 vs. -0.44), p=3x10-6), contradicting a lost-in-the-middle explanation and warning against a common production shortcut. Projecting measured reliability onto representative benchmark horizons quantifies a substantial gap between benchmark and production conditions, from 0.42 at GAIA-length horizons to 0.24 at hundred-step production horizons. For teams responsible for agent orchestration and reliability at scale, these results argue for horizon-aware evaluation and reliability budgeting in place of aggregate pass-rate metrics. Code, prompts, seeds, and raw trajectories are released.

Explore related subjects

Keep this discovery

BibTeXRIS

Shubhra Mittal. 2026-08-31. How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Making. https://arxiv.org/abs/2609.01660

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

Rock, Paper, Scissors, ... Dynamite - A Model of Disruption from New Technologies

We seek to understand the effect of adding disruptive highly-capable new technologies to competitions by assessing the addition of Dynamite to Rock-Paper-Scissors. We find that providing a versatile Dynamite move to only one player provides limited value (win probability increases from 50% to 55.5%) and is played rarely. That value decreases further if the game is expanded beyond just the original three moves. We also observe several mechanisms by which prior moves can become strategically unplayable, or obsolete. We hope that this model illustrates some non-intuitive aspects of developing new versatile technologies. We also hope that it illustrates some pitfalls for developers and integrators to avoid in order to create value rather than merely capability.

physics.soc-ph

Toward a social psychology of AI: language-model agents reproduce human-like minimal-group bias

Language-model agents now interact in groups, but evaluations that probe memorised stereotype content or use models to simulate people leave this social behaviour unmeasured. We adapt the minimal-group paradigm---social psychology's classic test of intergroup bias---into a controlled probe: an agent distributes points among anonymous peers bearing only an arbitrary group label. Across four reasoning models, mere categorisation into meaningless groups elicited in-group favouritism that vanished under a group-blind control and was concentrated in the numerical minority: minority deciders over-allocated to their own group relative to their numbers, majority deciders allocated close to proportionally, and the asymmetry closed at equal group sizes. Disabling reasoning in one model did not remove the disposition---if anything it grew---but nearly erased the minority-majority asymmetry, implicating deliberation in where bias concentrates rather than whether it appears. These open-weight reasoning models reproduce the behavioural signature of human intergroup discrimination, independent of stereotype content, and social psychology's theories and methods offer a paradigm for measuring and governing AI's social behaviour.

physics.soc-ph

Small worlds and clustering in spatial networks

Networks with underlying metric spaces attract increasing research attention in network science, statistical physics, applied mathematics, computer science, sociology, and other fields. This attention is further amplified by the current surge of activity in graph embedding. In the vast realm of spatial network models, only a few reproduce even the most basic properties of real-world networks. Here, we focus on three such properties--sparsity, small worldness, and clustering--and identify the general subclass of spatial homogeneous and heterogeneous network models that are sparse small worlds and that have nonzero clustering in the thermodynamic limit. We rely on the maximum entropy approach where network links correspond to noninteracting fermions whose energy dependence on spatial distances determines network small worldness and clustering.

physics.soc-ph