arXiv Science⌕ Search

arXiv · 2610.08806

Agentic AI Systems and Financial Stability, From Model Risk to Systemic Risk

Abstract

Financial stability rests on the premise that distress is largely idiosyncratic and therefore diversifiable: when one institution errs, the rest of the system absorbs the shock. That premise fails when many decision-makers depend on the same infrastructure. Agentic AI systems, which take consequential actions rather than only emitting predictions, are becoming such a substrate, and the binding concern shifts from the model risk of a single deployment to the systemic risk of the population. We develop this account in six mathematical settings. Chapter 1 recasts the PD/LGD/EAD decomposition of expected loss as an expected-harm identity and introduces a set-valued containment-risk measure; under a lever-assignment axiom that review cannot intercept an irreversible action, no non-preventive control satisfies a coherent tail constraint. Chapter 2 lifts this to a fleet sharing a foundation model: the shared model is a non-diversifiable common exposure whose expected-shortfall floor holds however large the fleet grows, and a percolation threshold governs contagion. Chapter 3 recasts these dynamics as a marked, Hawkes-excited jump diffusion, identifying the robust stress problem with a time-consistent entropic risk measure. Chapter 4 treats runtime guardrails as partially observed stochastic control, giving an observability trichotomy, a detection floor, and, under commit exogeneity, a runtime-impossibility corollary. Chapter 5 makes the adversary a player, yielding an underinvestment wedge proportional to systemic reach. Chapter 6 makes capability a state variable in an arms race. One claim runs through all six: the systematic component of agentic risk cannot be diversified, detected away, or reversed, and only ex-ante structural prevention moves it.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sriram Nagaraj, Seung Jung Lee. 2026-08-23. Agentic AI Systems and Financial Stability, From Model Risk to Systemic Risk. https://arxiv.org/abs/2610.08806

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Implementability in Insurance Markets with Adverse Selection

We consider an insurance market with hidden information, where the agent's type is private information and is drawn from an arbitrary type space. We study implementability of a collection of retention functions, namely, how to select premium schedules so that the resulting menu of contracts is incentive compatible, or truthful. Specifically, for general type spaces, implementability is equivalent to cyclical monotonicity of the collection of retention functions. For compact interval type spaces, we show that submodularity is a sufficient condition for implementability, under suitable type ordering assumptions. Moreover, for any implementable collection of retention functions, we characterize all corresponding premium schedules, up to a common additive constant. Finally, we apply our results to several standard classes of insurance contracts, for which the general implementability conditions admit simpler characterizations, and we provide several numerical illustrations.

q-fin.RM↗

A Functional Representation of Credit Behavior for Probability of Default Modeling

This paper proposes a framework for modeling probability of default via functional data analysis. By representing a series of credit variables as functions, we investigate whether intra-monthly information improves default predictability in linear models. We further show how a range of widely used variables, among them available funds, utilization rate, and overdraft, can all be derived from three quantities observed over time. Namely, 1) the type of each account, 2) the balance of the account, and 3) the size of the credit limit on the account. Retaining these quantities as continuous-time processes, rather than reducing them to monthly aggregated values, yields a continuous faithful representation of the borrower. When analyzing these three processes, we discovered that recurring events associated with the ordinal position among banking days created strong cyclical patterns. We therefore develop a relative time framework that aligns the recurring events across borrowers, ensuring that borrowers possess the same cyclical pattern, regardless of real time. We assess the framework using functional logistic regression. This approach accommodates the continuous representation while remaining closely related to a logistic regression model commonly used in credit risk practice, due to strict regulatory constraints. We show that, when equipped with an effective functional representation of transactional trajectories, the proposed model attains predictive performance on par with XGBoost while consistently outperforming logistic regression. Importantly, the model balances predictive performance of a machine learning model with the interpretability of linear default models, potentially enabling financial institutions to use the model, even under strict regulation.

q-fin.RM↗

Learned Monotone Recurrent Features in Governed Credit Scoring: The Price of the Frame and the Necessity of Macro Conditioning

Regulated credit scoring requires scores monotone non-decreasing in every exposure input. Deployed pipelines -- hand-crafted monotone aggregates feeding sign-constrained gradient boosting -- already meet this by composition; the open question is what learned temporal aggregation is worth inside one. We answer on five production-scale credit datasets at matched admissibility (one priced baseline convention excepted), with a monotone recurrent architecture whose per-input guarantee we extend, with proofs, to vector-valued inputs and to exogenously macro-conditioned decay gates, severities, thresholds, and peak memory. Two findings result. First, a strictness ladder: the value of learned monotone features rises with governance-frame strictness -- zero on unconstrained engineered panels, maximal in summaries-only frames -- replicated across two datasets and an official temporal-stability metric, though unconditioned features degrade on externally adjudicated later weeks. Second, a conditioning-delivery asymmetry under regime shift. On a train-on-boom, test-on-crisis mortgage design, two public macroeconomic series hurt as input columns, yet conditioning the recurrence on them delivers the paper's only learned-block crisis-cohort uplifts. The confirmed effect: +0.006 to +0.013 AUC on an internally pre-registered Freddie Mac replication, at all five held-out seeds. The discovery estimate: +0.015 to +0.021 on Fannie Mae (three of five seeds post hoc), worth 10-27 basis points of defaulted balance at an 80% approval cutoff, and grows with early-prepaid loans excluded. A state-level test identifies the mechanism: between-cohort calibration transfer. A pandemic-band episode bounds scope: under forbearance-distorted labels the gain generalizes at a quarter to a third of crisis size on Fannie Mae, on Freddie Mac only against the capacity control.

q-fin.RM↗