arXiv Science⌕ Search

arXiv · 2610.08869

Learned Monotone Recurrent Features in Governed Credit Scoring: The Price of the Frame and the Necessity of Macro Conditioning

Abstract

Regulated credit scoring requires scores monotone non-decreasing in every exposure input. Deployed pipelines -- hand-crafted monotone aggregates feeding sign-constrained gradient boosting -- already meet this by composition; the open question is what learned temporal aggregation is worth inside one. We answer on five production-scale credit datasets at matched admissibility (one priced baseline convention excepted), with a monotone recurrent architecture whose per-input guarantee we extend, with proofs, to vector-valued inputs and to exogenously macro-conditioned decay gates, severities, thresholds, and peak memory. Two findings result. First, a strictness ladder: the value of learned monotone features rises with governance-frame strictness -- zero on unconstrained engineered panels, maximal in summaries-only frames -- replicated across two datasets and an official temporal-stability metric, though unconditioned features degrade on externally adjudicated later weeks. Second, a conditioning-delivery asymmetry under regime shift. On a train-on-boom, test-on-crisis mortgage design, two public macroeconomic series hurt as input columns, yet conditioning the recurrence on them delivers the paper's only learned-block crisis-cohort uplifts. The confirmed effect: +0.006 to +0.013 AUC on an internally pre-registered Freddie Mac replication, at all five held-out seeds. The discovery estimate: +0.015 to +0.021 on Fannie Mae (three of five seeds post hoc), worth 10-27 basis points of defaulted balance at an 80% approval cutoff, and grows with early-prepaid loans excluded. A state-level test identifies the mechanism: between-cohort calibration transfer. A pandemic-band episode bounds scope: under forbearance-distorted labels the gain generalizes at a quarter to a third of crisis size on Fannie Mae, on Freddie Mac only against the capacity control.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yew Lee Tan. 2026-10-06. Learned Monotone Recurrent Features in Governed Credit Scoring: The Price of the Frame and the Necessity of Macro Conditioning. https://arxiv.org/abs/2610.08869

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Implementability in Insurance Markets with Adverse Selection

We consider an insurance market with hidden information, where the agent's type is private information and is drawn from an arbitrary type space. We study implementability of a collection of retention functions, namely, how to select premium schedules so that the resulting menu of contracts is incentive compatible, or truthful. Specifically, for general type spaces, implementability is equivalent to cyclical monotonicity of the collection of retention functions. For compact interval type spaces, we show that submodularity is a sufficient condition for implementability, under suitable type ordering assumptions. Moreover, for any implementable collection of retention functions, we characterize all corresponding premium schedules, up to a common additive constant. Finally, we apply our results to several standard classes of insurance contracts, for which the general implementability conditions admit simpler characterizations, and we provide several numerical illustrations.

q-fin.RM↗

A Functional Representation of Credit Behavior for Probability of Default Modeling

This paper proposes a framework for modeling probability of default via functional data analysis. By representing a series of credit variables as functions, we investigate whether intra-monthly information improves default predictability in linear models. We further show how a range of widely used variables, among them available funds, utilization rate, and overdraft, can all be derived from three quantities observed over time. Namely, 1) the type of each account, 2) the balance of the account, and 3) the size of the credit limit on the account. Retaining these quantities as continuous-time processes, rather than reducing them to monthly aggregated values, yields a continuous faithful representation of the borrower. When analyzing these three processes, we discovered that recurring events associated with the ordinal position among banking days created strong cyclical patterns. We therefore develop a relative time framework that aligns the recurring events across borrowers, ensuring that borrowers possess the same cyclical pattern, regardless of real time. We assess the framework using functional logistic regression. This approach accommodates the continuous representation while remaining closely related to a logistic regression model commonly used in credit risk practice, due to strict regulatory constraints. We show that, when equipped with an effective functional representation of transactional trajectories, the proposed model attains predictive performance on par with XGBoost while consistently outperforming logistic regression. Importantly, the model balances predictive performance of a machine learning model with the interpretability of linear default models, potentially enabling financial institutions to use the model, even under strict regulation.

q-fin.RM↗

When Is the Gini Loading More Prudent? Tail Structure and the Ordering of the Standard Deviation and the Gini Mean Difference

The standard deviation (SD) and the Gini mean difference (GMD) are the two canonical measures of variability used to load premiums, set risk margins and allocate capital, yet no universal ordering between them exists. We show that the comparison is \emph{equivalent} to asking whether the coefficient of variation of the spacing $|X-X'|$ generated by two independent copies of the risk exceeds unity, so that the exponential law -- whose spacing is again exponential -- is the universal knife-edge separating the two regimes. Reading the GMD as twice the maxiance, that is, as a second-order \emph{dual} moment in the sense of Yaari's dual theory, the problem becomes an explicit comparison of primal and dual second-order variability. We derive a closed-form representation of the mean excess function of the spacing in terms of the hazard and reverse hazard rates of $X$, and use it to prove that heavy-tailed behavior -- a decreasing hazard rate or an increasing reverse hazard rate -- yields SD dominance, whereas two-sided light tails yield GMD dominance; when both the hazard and the reverse hazard rate are monotone, equality characterizes the exponential law. The sufficient conditions for SD dominance are stable under mixing, those for GMD dominance under convolution, and both under tail truncation, which makes them operational in frailty and collective risk models. We classify the severity, lifetime and frequency distributions of actuarial practice accordingly, quantify the consequences for SD- and Gini-loaded premium principles and for Gini-type tail risk measures, and show that the sign of $\SD-\GMD$ across thresholds furnishes a simple diagnostic for tail aging.

q-fin.RM↗