arXiv ScienceSearch

arXiv · 2609.14576

Towards foundation models for insurance risk modelling

Abstract

Claim narratives, images and sensor data contain information about insured risks that is difficult to use through existing actuarial models. Foundation models learn patterns from large datasets before being adapted to particular tasks. By turning these high-dimensional sources into variables or numerical representations, they could help insurers use more of the information they already collect, potentially reducing the experience needed to develop each application. For example, a language model could identify a worsening injury in a new claim note, allowing a reserving model to recognise the change in expected cost before the payments reveal the deterioration. In this paper, we review language, vision, geospatial, time series, tabular and scientific models, explaining existing insurance applications and potential future uses. Scientific models extend this approach to future weather and climate conditions: their simulations can inform loss estimates once local hazards are linked to asset damage, repair costs and insurance coverage. We propose a process to connect these model outputs to actuarial calculations and to assess their predictive contribution, stability and compliance with rules on information use. Evaluating these applications is difficult when final claim costs become known only after long delays, large losses are rare or patterns learned elsewhere fail to transfer to the target portfolio. Richer data can reveal private information and support finer risk classification, which can change access to insurance. Reusing the same models across insurers also creates dependence on shared predictions and providers.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Christopher Blier-Wong. 2026-09-13. Towards foundation models for insurance risk modelling. https://arxiv.org/abs/2609.14576

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Quantifying the 2027 Solvency II Risk Margin Reform

The 2027 Solvency II reform recalibrates the Risk Margin by reducing the prescribed cost-of-capital rate from 6% to 4.75% and introducing a time-dependent attenuation of future Solvency Capital Requirements. This paper develops an analytical and numerical framework for characterizing the effect of the final regulatory calibration. By normalizing discounted projected capital requirements into a probability distribution over run-off time, we obtain an exact representation of the ratio between the revised and previous Risk Margins. The framework yields sharp universal and horizon-specific bounds, characterizes the effect of later capital timing through stochastic dominance, and shows that mean run-off time alone does not determine the reform effect when temporal dispersion varies. Additional bounds are derived conditional on mean run-off time and horizon. Conditional on a given projected SCR path, the mechanical reduction lies between 20.83% and 60.42%. For proportional Best Estimate projections, an exact covariance decomposition identifies how departures from proportionality affect the relative reform ratio. Holding the projected SCR path fixed, the revised formula has a lower direct interest-rate semi-elasticity than the previous calibration. A reduced-form stochastic extension further quantifies convexity effects arising from uncertainty in capital persistence. Numerical applications reconstruct Risk Margin calculations from published actuarial run-off profiles and complement them with controlled long-horizon, interest-rate, and persistence experiments. The results show that the reform effect is governed by the temporal structure of future capital and provide tractable tools for assessing projected capital profiles, Risk Margin simplifications, and direct discount-rate sensitivity.

q-fin.RM

Forecasting Liquidity Withdraw with Machine Learning Models

Liquidity withdrawal is a critical indicator of market fragility. In this project, I test a framework for forecasting liquidity withdrawal at the individual-stock level, ranging from less liquid stocks to highly liquid large-cap tickers, and evaluate the relative performance of competing model classes in predicting short-horizon order book stress. We introduce the Liquidity Withdrawal Index (LWI) -- defined as the ratio of order cancellations to the sum of standing depth and new additions at the best quotes -- as a bounded, interpretable measure of transient liquidity removal. Using Nasdaq market-by-order (MBO) data, we compare a spectrum of approaches: linear benchmarks (AR, HAR), and non-linear tree ensembles (XGBoost), across horizons ranging from 250\,ms to 5\,s. Beyond predictive accuracy, our results provide insights into order placement and cancellation dynamics, identify regimes where linear versus non-linear signals dominate, and highlight how early-warning indicators of liquidity withdrawal can inform both market surveillance and execution.

q-fin.RM

Simplifying Cyber Cat(astrophe)s with Cyber Kittens: Power Law Plausibility for Cyber Insurance Risks

Cyber insurance requires accurate modeling of worst-case catastrophic (cat) events, but the field lacks robust quantitative approaches for estimating upper-bound losses. Building on a recent dataset of 24 cyber cat events over 30 years, this work tests whether cyber economic losses follow a power law distribution. We analyze "cyber kittens" - sub-1B USD events distinguished from cat events (1B+ USD) only by magnitude - extracted via LLM from cyber insurance claims data (2020-2024). Using victim count (weighted by claim year) as a proxy for economic loss, we link kitten-sized events to known cat events to estimate losses. The kitten distribution proved consistent with the cat dataset, and power laws were statistically plausible: each order-of-magnitude increase in event size corresponds to a 5-7x drop in probability. Extrapolating, an event 100x the largest 2020-2024 cat event is expected roughly every 206 years, translating to 100-250B USD in losses - catastrophic, but not extraordinary relative to other insurance lines.

q-fin.RM