arXiv ScienceSearch

arXiv subjects

Arthur Charpentier

Publications and source records attributed to Arthur Charpentier.

At least 19 recordsLinked to original sources

Quantifying Portfolio Demutualization: A Benchmark-Relative Pooling--Profiling Scale

Insurance pricing combines pooling with differentiation: a tariff may leave benchmark differences in expected loss partly mutualized or translate them into policy-level premium differences. We propose a benchmark-relative pooling--profiling scale with two complementary coordinates. The coupled $L^p$ coordinate measures policy-level alignment between an evaluated tariff and a stated benchmark pure premium, whereas the marginal Wasserstein coordinate compares their exposure-weighted premium distributions. The difference between their residual $p$-costs defines an allocation mismatch. Under portfolio balance, the coupled $L^1$ coordinate has an exact actuarial interpretation: it is the fraction of the transfer volume induced by full pooling that the tariff removes. Synthetic and motor-insurance applications show that broad classes, proxies, shrinkage and tail caps can affect marginal differentiation, policy-level allocation and transfers differently. A barycentric group-parity intervention further shows that conditional premium disparities can fall mainly through reallocation and restored benchmark-relative transfers, with little change in marginal differentiation.

stat.AP

What Does a Benford Test Actually Test? Marginal Conformity, Sampling Structure, and Forensic Inference

Benford's law specifies a marginal distribution for significant digits, whereas the usual first-digit Pearson $p$-value is calibrated under an independent multinomial sampling model. We separate these statements with four constructions that share the same one-time or pooled Benford target but have different joint structures. Under a wrapped-Gaussian circular Markov construction, the nominal 5% Pearson test rejects 21.6% of samples at a fixed persistence level even though every one-time marginal is exactly Benford; a sequence-level second-order moment-matching calibration reduces the rejection rate to 4.8%. The same strategy performs well in the randomized-rotation and random-composition designs. The effect persists across changes in sample size and sequence length, and a grid approximation to a continuous log-significand Cramér-von Mises discrepancy shows the same design dependence. In the wrapped-Gaussian design examined here, corrected procedures retain rejection probabilities that rise with the size of a smooth marginal departure. Finally, we distinguish calibration from an information boundary: significand-only methods cannot detect changes that leave the complete significand process unchanged, but they readily detect perturbations that alter it. Benford conformity is a marginal statement; a Benford $p$-value is valid only relative to a specified statistic, sampling law, and calibration procedure; an integrity claim requires substantive competing models.

stat.ME

From Risk Prediction to Risk Mechanisms: A Multi-Resolution Causal Representation for Road Safety and Motor Insurance

Operational risk models can estimate event frequencies precisely while leaving the underlying mechanisms weakly resolved. We study this resolution mismatch using motor insurance and road safety. A revisable DAG represents trip-level crash generation; a separate predictive layer links annual rating information to latent driving states; and an observation process maps crashes into recorded liability claims. Compatibility sets collect the structural and crash-to-claim laws that reproduce an observed annual contrast under stated restrictions. External studies enter only through explicit bridge assumptions and sensitivity bounds. The framework therefore distinguishes sampling uncertainty from uncertainty about structure, observation, and study-to-target correspondence. Two limited examples illustrate the gap. A sublinear mileage relation constrains an aggregate accident rate per unit distance but not its mechanism. In French motor-liability data, the 18-20 versus 40-49 claim-frequency relativity is 3.388 in a model including vehicle and geographic variables and 1.235 when the same model also conditions on a medium-resolution bonus-malus score. These are different predictive functionals, not successive causal adjustments. A Spanish culpability estimate is used only to show how a strong cross-study bridge would restrict a toy bookkeeping region. The framework does not estimate the full DAG; it makes explicit which assumptions are needed before an annual predictive contrast can support a mechanism-specific risk statement.

stat.AP

Beyond Zipf's Law: Equifinality and Mechanistic Inference from Scaling Laws

Scaling laws summarize complex systems through low-dimensional regularities, but the same marginal law can arise from different stochastic dynamics. We examine this ambiguity for Zipf rank--frequency scaling. An i.i.d. finite-Zipf process, a persistent Markov chain, and canonical sample-space reduction (SSR) are constructed to have exactly the same stationary marginal, $p_j=(jH_V)^{-1}$, while a latent-scale mixture produces a similar marginal through aggregation. The first three models therefore hold the population rank distribution fixed while changing temporal organization. Markov dependence shifts the finite-sample distribution of fitted exponents; at moderate persistence, a block-adjusted effective sample size reproduces most of this shift. Excess lag-1 mutual information detects serial dependence relative to a shuffle null, whereas transition direction separates reversible Markov persistence from the directional contraction of SSR. Conditioning on latent scale reveals the aggregation mechanism. Fit-window, alphabet, and sequence-boundary analyses quantify sensitivity to the observation design. These examples separate the population marginal, the finite-sample behavior of a fitted exponent, and temporal structure. Matching a scaling law is therefore a compatibility condition, not a mechanism identifier: discrimination requires observables on which candidate processes make different predictions.

stat.AP

Direct and Indirect Discrimination in Generalized Linear Models

Generalized linear models are central to actuarial modelling of binary risk, claim frequency, utilization, and cost-related outcomes. Yet fairness diagnostics often rely on linear-model intuitions, although GLM predictions are obtained by transporting a latent score through a nonlinear inverse link. We develop a moment-based decomposition framework for diagnosing group disparities in fitted GLM predictions. In an exact linear-Gaussian benchmark, the Wasserstein barycentric criterion for distributional demographic-parity violation reduces to a two-moment criterion and decomposes into direct mean, indirect mean, interaction, and structural components. For GLMs, we distinguish the empirical output-scale criterion $U_2(f)$, a within-group proxy $\widetilde U_2(f)$, and a leading decomposition $D_1(f)$. This leading term preserves the four linear channels and adds two curvature components induced by the inverse link: curvature coupling and curvature amplification. We derive explicit formulas for logistic, Poisson, and Tweedie specifications and illustrate the diagnostic on medical-expenditure survey data. The framework is not a legal test of discrimination, nor a full characterization of distributional parity outside the linear-Gaussian case. It is a tractable actuarial diagnostic for identifying whether fitted prediction disparities arise from explicit sensitive effects, proxy-mediated covariate profiles, covariance-structure differences, or nonlinear link effects.

stat.ME

Beyond Procedure: Substantive Fairness in Conformal Prediction

Conformal prediction (CP) offers distribution-free uncertainty quantification for machine learning models, yet its interplay with fairness in downstream decision-making remains underexplored. Moving beyond CP as a standalone operation (procedural fairness), we analyze the holistic decision-making pipeline to evaluate substantive fairness-the equity of downstream outcomes. Theoretically, we derive an upper bound that decomposes prediction-set size disparity into interpretable components, clarifying how label-clustered CP helps control method-driven contributions to unfairness. To facilitate scalable empirical analysis, we introduce an LLM-in-the-loop evaluator that approximates human assessment of substantive fairness across diverse modalities. Our experiments show that label-clustered CP often provides a favorable balance between utility and substantive fairness, while reducing set-size disparities in line with our theory. Finally, we empirically show that equalized set sizes, rather than coverage, strongly correlate with improved substantive fairness, enabling practitioners to design more fair CP systems. Our code is available at https://github.com/layer6ai-labs/llm-in-the-loop-conformal-fairness.

stat.ML

Fair regression under localized demographic parity constraints

Demographic parity (DP) is a widely used group fairness criterion requiring predictive distributions to be invariant across sensitive groups. While natural in classification, full distributional DP is often overly restrictive in regression and can lead to substantial accuracy loss. We propose a relaxation of DP tailored to regression, enforcing parity only at a finite set of quantile levels and/or score thresholds. Concretely, we introduce a novel (${\ell}$, Z)-fair predictor, which imposes groupwise CDF constraints of the form F f |S=s (z m ) = ${\ell}$ m for prescribed pairs (${\ell}$ m , z m ). For this setting, we derive closed-form characterizations of the optimal fair discretized predictor via a Lagrangian dual formulation and quantify the discretization cost, showing that the risk gap to the continuous optimum vanishes as the grid is refined. We further develop a model-agnostic post-processing algorithm based on two samples (labeled for learning a base regressor and unlabeled for calibration), and establish finite-sample guarantees on constraint violation and excess penalized risk. In addition, we introduce two alternative frameworks where we match group and marginal CDF values at selected score thresholds. In both settings, we provide closed-form solutions for the optimal fair discretized predictor. Experiments on synthetic and real datasets illustrate an interpretable fairness-accuracy trade-off, enabling targeted corrections at decision-relevant quantiles or thresholds while preserving predictive performance.

stat.ML

Sequential Transport for Causal Mediation Analysis

We propose sequential transport (ST), a distributional framework for mediation analysis that combines optimal transport (OT) with a mediator directed acyclic graph (DAG). Instead of relying on cross-world counterfactual assumptions, ST constructs unit-level mediator counterfactuals by minimally transporting each mediator, either marginally or conditionally, toward its distribution under an alternative treatment while preserving the causal dependencies encoded by the DAG. For numerical mediators, ST uses monotone (conditional) OT maps based on conditional CDF/quantile estimators; for categorical mediators, it extends naturally via simplex-based transport. We establish consistency of the estimated transport maps and of the induced unit-level decompositions into mutatis mutandis direct and indirect effects under standard regularity and support conditions. When the treatment is randomized or ignorable (possibly conditional on covariates), these decompositions admit a causal interpretation; otherwise, they provide a principled distributional attribution of differences between groups aligned with the mediator structure. Gaussian examples show that ST recovers classical mediation formulas, while additional simulations confirm good performance in nonlinear and mixed-type settings. An application to the COMPAS dataset illustrates how ST yields deterministic, DAG-consistent counterfactual mediators and a fine-grained mediator-level attribution of disparities.

stat.ME

Decomposing Probabilistic Scores: Reliability, Information Loss and Uncertainty

Calibration is a conditional property that depends on the information retained by a predictor. We develop decomposition identities for arbitrary proper losses that make this dependence explicit. At any information level $\mathcal A$, the expected loss of an $\mathcal A$-measurable predictor splits into a proper-regret (reliability) term and a conditional entropy (residual uncertainty) term. For nested levels $\mathcal A\subseteq\mathcal B$, a chain decomposition quantifies the information gain from $\mathcal A$ to $\mathcal B$. Applied to classification with features $\boldsymbol{X}$ and score $S=s(\boldsymbol{X})$, this yields a three-term identity: miscalibration, a {\em grouping} term measuring information loss from $\boldsymbol{X}$ to $S$, and irreducible uncertainty at the feature level. We leverage the framework to analyze post-hoc recalibration, aggregation of calibrated models, and stagewise/boosting constructions, with explicit forms for Brier and log-loss.

cs.LG

Federated Measurement of Demographic Disparities from Quantile Sketches

Many fairness goals are defined at a population level that misaligns with siloed data collection, which remains unsharable due to privacy regulations. Horizontal federated learning (FL) enables collaborative modeling across clients with aligned features without sharing raw data. We study federated auditing of demographic parity through score distributions, measuring disparity as a Wasserstein--Frechet variance between sensitive-group score laws, and expressing the population metric in federated form that makes explicit how silo-specific selection drives local-global mismatch. For the squared Wasserstein distance, we prove an ANOVA-style decomposition that separates (i) selection-induced mixture effects from (ii) cross-silo heterogeneity, yielding tight bounds linking local and global metrics. We then propose a one-shot, communication-efficient protocol in which each silo shares only group counts and a quantile summary of its local score distributions, enabling the server to estimate global disparity and its decomposition, with $O(1/k)$ discretization bias ($k$ quantiles) and finite-sample guarantees. Experiments on synthetic data and COMPAS show that a few dozen quantiles suffice to recover global disparity and diagnose its sources.

stat.ML

Perceived Fairness in Networks

The usual definitions of algorithmic fairness focus on population-level statistics, such as demographic parity or equal opportunity. However, in many social or economic contexts, fairness is not perceived globally, but locally, through an individual's peer network and comparisons. We propose a theoretical model of perceived fairness networks, in which each individual's sense of discrimination depends on the local topology of interactions. We show that even if a decision rule satisfies standard criteria of fairness, perceived discrimination can persist or even increase in the presence of homophily or assortative mixing. We propose a formalism for the concept of fairness perception, linking network structure, local observation, and social perception. Analytical and simulation results highlight how network topology affects the divergence between objective fairness and perceived fairness, with implications for algorithmic governance and applications in finance and collaborative insurance.

econ.TH

Decomposing Direct and Indirect Biases in Linear Models under Demographic Parity Constraint

Linear models are widely used in high-stakes decision-making due to their simplicity and interpretability. Yet when fairness constraints such as demographic parity are introduced, their effects on model coefficients, and thus on how predictive bias is distributed across features, remain opaque. Existing approaches on linear models often rely on strong and unrealistic assumptions, or overlook the explicit role of the sensitive attribute, limiting their practical utility for fairness assessment. We extend the work of (Chzhen and Schreuder, 2022) and (Fukuchi and Sakuma, 2023) by proposing a post-processing framework that can be applied on top of any linear model to decompose the resulting bias into direct (sensitive-attribute) and indirect (correlated-features) components. Our method analytically characterizes how demographic parity reshapes each model coefficient, including those of both sensitive and non-sensitive features. This enables a transparent, feature-level interpretation of fairness interventions and reveals how bias may persist or shift through correlated variables. Our framework requires no retraining and provides actionable insights for model auditing and mitigation. Experiments on both synthetic and real-world datasets demonstrate that our method captures fairness dynamics missed by prior work, offering a practical and interpretable tool for responsible deployment of linear models.

stat.ML

Functional Analysis of Loss-development Patterns in P&C Insurance

We analyze loss development in NAIC Schedule P loss triangles using functional data analysis methods. Adopting the functional viewpoint, our dataset comprises 3300+ curves of incremental loss ratios (ILR) of workers' compensation lines over 24 accident years. Relying on functional data depth, we first study similarities and differences in development patterns based on company-specific covariates, as well as identify anomalous ILR curves. The exploratory findings motivate the probabilistic forecasting framework developed in the second half of the paper. We propose a functional model to complete partially developed ILR curves based on partial least squares regression of PCA scores. Coupling the above with functional bootstrapping allows us to quantify future ILR uncertainty jointly across all future lags. We demonstrate that our method has much better probabilistic scores relative to Chain Ladder and in particular can provide accurate functional predictive intervals.

stat.AP

Linear Risk Sharing on Networks

Over the past decade alternatives to traditional insurance and banking have grown in popularity. The desire to encourage local participation has lead products such as peer-to-peer insurance, reciprocal contracts, and decentralized finance platforms to increasingly rely on network structures to redistribute risk among participants. In this paper, we develop a comprehensive framework for linear risk sharing (LRS), where random losses are reallocated through nonnegative linear operators which can accommodate a wide range of networks. Building on the theory of stochastic and doubly stochastic matrices, we establish conditions under which constraints such as budget balance, fairness, and diversification are guaranteed. The convex order framework allows us to compare different allocations rigorously, highlighting variance reduction and majorization as natural consequences of doubly stochastic mixing. We then extend the analysis to network-based sharing, showing how their topology shapes risk outcomes in complete, star, ring, random, and scale-free graphs. A second layer of randomness, where the sharing matrix itself is random, is introduced via Erdős--Rényi and preferential-attachment networks, connecting risk-sharing properties to degree distributions. Finally, we study convex combinations of identity and network-induced operators, capturing the trade-off between self-retention and diversification. Our results provide design principles for fair and efficient peer-to-peer insurance and network-based risk pooling, combining mathematical soundness with economic interpretability.

econ.TH

KNN and K-means in Gini Prametric Spaces

This paper introduces enhancements to the K-means and K-nearest neighbors (KNN) algorithms based on the concept of Gini prametric spaces, instead of traditional metric spaces. Unlike standard distance metrics, Gini prametrics incorporate both value-based and rank-based measures, offering robustness to noise and outliers. The main contributions include: (1) a Gini prametric that captures rank information alongside value distances; (2) a Gini K-means algorithm that is provably convergent and resilient to noisy data; and (3) a Gini KNN method that performs competitively with state-of-the-art approaches like Hassanat's distance in noisy environments. Experimental evaluations on 16 UCI datasets demonstrate the superior performance and efficiency of the Gini-based algorithms in clustering and classification tasks. This work opens new directions for rank-based prametrics in machine learning and statistical analysis.

cs.LG

Disentangled Deep Smoothed Bootstrap for Fair Imbalanced Regression

Imbalanced distribution learning is a common and significant challenge in predictive modeling, often reducing the performance of standard algorithms. Although various approaches address this issue, most are tailored to classification problems, with a limited focus on regression. This paper introduces a novel method to improve learning on tabular data within the Imbalanced Regression (IR) framework, which is a critical problem. We propose using Variational Autoencoders (VAEs) to model and define a latent representation of data distributions. However, VAEs can be inefficient with imbalanced data like other standard approaches. To address this, we develop an innovative data generation method that combines a disentangled VAE with a Smoothed Bootstrap applied in the latent space. We evaluate the efficiency of this method through numerical comparisons with competitors on benchmark datasets for IR.

cs.LG

When Numbers Mislead Us

The belief that numbers offer a single, objective description of reality overlooks a crucial truth: data does not speak for itself. Every dataset results from choices-what to measure, how, when, and with whom-which inevitably reflect implicit, and sometimes ideological, assumptions about what is worth quantifying. Moreover, in any analysis, what remains unmeasured can be just as significant as what is captured. When a key variable is omitted-whether by neglect, design, or ignorance-it can distort the observed relationships between other variables. This phenomenon, known as omitted variable bias, may produce misleading correlations or conceal genuine effects. In some cases, accounting for this hidden factor can completely overturn the conclusions drawn from a superficial analysis. This is precisely the mechanism behind Simpson's paradox.

stat.OT

Disaster Risk Financing through Taxation: A Framework for Regional Participation in Collective Risk-Sharing

We consider an economy composed of different risk profile regions wishing to be hedged against a disaster risk using multi-region catastrophe insurance. Such catastrophic events inherently have a systemic component; we consider situations where the insurer faces a non-zero probability of insolvency. To protect the regions against the risk of the insurer's default, we introduce a public-private partnership between the government and the insurer. When a disaster generates losses exceeding the total capital of the insurer, the central government intervenes by implementing a taxation system to share the residual claims. In this study, we propose a theoretical framework for regional participation in collective risk-sharing through tax revenues by accounting for their disaster risk profiles and their economic status.

econ.TH