arXiv Science⌕ Search

arXiv · 2610.02503

Compound AI System Reliability: A Failure Taxonomy and Resilience Pattern Catalog from 150 Production Incidents

Abstract

Deploying compound AI systems reliably and safely requires understanding failure modes that emerge at component boundaries, not within individual models. Cascading errors propagate across component boundaries, silent quality degradation evades standard monitoring, and coordination failures yield incorrect collective behavior from individually correct parts. We analyze 150 production incident reports from open-source compound AI projects and anonymized enterprise deployments to construct a taxonomy of 23 failure modes organized into five categories: retrieval failures, generation failures, tool failures, orchestration failures, and integration failures. For each category, we propose resilience patterns with measured effectiveness from controlled fault injection experiments. Circuit breakers reduce cascade propagation by 89%, output quality gates catch 73% of silent degradation before user impact, and component isolation reduces blast radius by 64%. Systems implementing three or more resilience patterns from our catalog reduce mean-time-to-recovery (MTTR) by 71% compared to unstructured monitoring baselines. We release the incident taxonomy and pattern catalog as a practitioner resource.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rudrendu Kumar Paul, Sourav Nandy. 2026-10-01. Compound AI System Reliability: A Failure Taxonomy and Resilience Pattern Catalog from 150 Production Incidents. https://arxiv.org/abs/2610.02503

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

LogLLM: Log-based Anomaly Detection Using Large Language Models

Software systems often record important runtime information in logs to help with troubleshooting. Log-based anomaly detection has become a key research area that aims to identify system issues through log data, ultimately enhancing the reliability of software systems. Existing methods often fall short in capturing the semantic information, typically expressed in natural language, or the sequential dependencies inherent in log sequences. In this paper, we propose LogLLM, a framework that enables collaboration between heterogeneous LLMs for log-based anomaly detection. LogLLM exploits the complementary capabilities of different LLM architectures: a Transformer encoder-based LLM is employed to extract fine-grained semantic vectors from individual log messages, while a Transformer decoder-based LLM is utilized to model sequential dependencies and generate anomaly detection decisions. To enable effective collaboration between these heterogeneous LLMs, we introduce a learnable projector to align their vector representation spaces. Furthermore, we design a progressive three-stage training strategy to optimize the collaboration between heterogeneous LLMs by gradually aligning their representations and adapting them to log anomaly detection. Unlike conventional methods that require log parsers to extract templates, LogLLM preprocesses log messages with regular expressions, streamlining the entire process. Experimental results on four public real-world datasets demonstrate that LogLLM outperforms state-of-the-art methods, achieving an average F$_1$-score improvement of 6.6% over the strongest existing approach. Further analyses provide insights into the effectiveness of the key components of the model architecture and progressive three-stage training strategy.

cs.SE↗

The EmpathiSEr: Development and Validation of Software Engineering Oriented Empathy Scales

Empathy plays a critical role in software engineering (SE), for instance in situations where developers interpret non-technical users' frustration with system usability or where product owners account for the technical constraints experienced by engineers during implementation. Such interactions shape collaboration, communication, and user-centred design outcomes. Although SE research has increasingly recognised empathy as a key human aspect, there remains no validated instrument specifically designed to measure it within the unique socio-technical contexts of SE. Existing generic empathy scales, while well-established in psychology and healthcare, often rely on language, scenarios, and assumptions that are not meaningful or interpretable for software practitioners. These scales fail to account for the diverse, role-specific, and domain-bound expressions of empathy in SE, such as understanding a non-technical user's frustrations or another practitioner's technical constraints, which differ substantially from empathy in clinical or everyday contexts. To address this gap, we developed and validated two domain-specific empathy scales: EmpathiSEr-P, assessing empathy among practitioners, and EmpathiSEr-U, capturing practitioner empathy towards users. Grounded in a practitioner-informed conceptual framework, the scales encompass three dimensions of empathy: cognitive empathy, affective empathy, and empathic responses. We followed a rigorous, multi-phase methodology, including expert evaluation, cognitive interviews, and two practitioner surveys. The resulting instruments represent the first psychometrically validated empathy scales tailored to SE, offering researchers and practitioners a tool for assessing empathy and designing empathy-enhancing interventions in software teams and user interactions.

cs.SE↗

Verify Before You Fix: Agentic Execution Grounding for Trustworthy Cross-Language Code Analysis

Learned classifiers deployed in agentic pipelines face a fundamental reliability problem: predictions are probabilistic inferences, not verified conclusions, and acting on them without grounding in observable evidence leads to compounding failures across downstream stages. Software vulnerability analysis makes this cost concrete and measurable. We address this through a unified cross-language vulnerability lifecycle framework built around three LLM-driven reasoning stages-hybrid structural-semantic detection, execution-grounded agentic validation, and validation-aware iterative repair-governed by a strict invariant: no repair action is taken without execution-based confirmation of exploitability. Cross-language generalization is achieved via a Universal Abstract Syntax Tree (uAST) normalizing Java, Python, and C++ into a shared structural schema, combined with a hybrid fusion of GraphSAGE and Qwen2.5-Coder-1.5B embeddings through learned two-way gating, whose per-sample weights provide intrinsic explainability at no additional cost. The framework achieves 89.84-92.02% intra-language detection accuracy and 74.43-80.12% zero-shot cross-language F1, resolving 69.74% of vulnerabilities end-to-end at a 12.27% total failure rate. Ablations establish necessity: removing uAST degrades cross-language F1 by 23.42%, while disabling validation increases unnecessary repairs by 131.7%. These results demonstrate that execution-grounded closed-loop reasoning is a principled and practically deployable mechanism for trustworthy LLM-driven agentic AI.

cs.SE↗