arXiv ScienceSearch

arXiv subjects

Finn Lindgren

Publications and source records attributed to Finn Lindgren.

At least 19 recordsLinked to original sources

Spatially continuous modelling of aggregated outcome data

This work develops a block aggregation approach to spatial estimation and prediction when the response is observed at a coarse spatial scale, for example as counts of events in administrative areas, or blocks, while covariates are available at a finer spatial resolution, typically as raster images. Our approach specifies a linear predictor at the finer resolution as a combination of covariate effects and a latent, spatially continuous Gaussian process. This linear predictor then determines the distribution of the response through an inverse link function and spatial integration. We use a simulation study to evaluate the performance of the proposed approach in comparison to two industry standard approaches: a traditional geostatistical model that associates each response with the centroid of its block; and a Markov random field (MRF) approach that aggregates covariate data to block-level. As expected, the differences in performance among the three approaches are small with respect to block-level prediction. The rationale for, and advantage of, the block aggregation approach lies in its delivery of reliable inferences at whatever spatial resolution is required in a particular application. We describe two applications: a linear Gaussian sampling model of wastewater virus concentrations in England, using population density as covariate; and log-linear Poisson model of cardiovascular hospitalisations in England using socio-demographic variables at fine-scale administrative units as covariates.

stat.ME

Areal Disaggregation: A Small Area Estimation Perspective

Producing reliable estimates of health and demographic indicators at fine areal scales is crucial for examining heterogeneity and supporting localized health policy. However, many surveys release outcomes only at coarser administrative levels, thereby limiting their relevance for decision-making. We propose a fully Bayesian, single-stage spatial modeling framework for area-level disaggregation that generates fine-scale estimates of indicators directly from coarsely aggregated survey data. By defining a latent spatial process at the target resolution and linking it to observed outcomes through an aggregation step, the framework adopts small-area estimation techniques while incorporating covariates and delivering coherent uncertainty quantification. The proposed methods are implemented with inlabru to achieve computational efficiency. We evaluate performance through a simulation study of general fertility rates in Kenya to demonstrate the models' ability to recover fine-scale variation across diverse data-generating scenarios. We further apply the framework to two national surveys to produce district-level fertility estimates from the 2022 Kenya Demographic and Health Survey and, more importantly, district-level indicators for unpaid care and domestic work and mass media usage from the 2021 Kenya Time Use Survey.

stat.ME

Influence of river incision on landslides triggered in Nepal by the Gorkha earthquake: Results from a pixel-based susceptibility model using inlabru

This study presents a comprehensive framework for modelling earthquake-induced landslides (EQILs) through a channel-based analysis of landslide centroid distributions. A key innovation is the incorporation of the normalised channel steepness index ($k_{sn}$) as a physically meaningful and novel covariate, inferring hillslope erosion and fluvial incision processes. Used within spatial point process models, $k_{sn}$ supports the generation of landslide susceptibility maps with quantified uncertainty. To address spatial data misalignment between covariates and landslide observations, we leverage the inlabru framework, which enables coherent integration through mesh-based disaggregation, thereby overcoming challenges associated with spatially misaligned data integration. Our modelling strategy explicitly prioritises prospective transferability to unseen geographical regions, provided that explanatory variable data are available. By modelling both landslide locations and sizes, we find that elevated $k_{sn}$ is strongly associated with increased landslide susceptibility but not with landslide magnitude. The best-fitting Bayesian model, validated through cross-validation, offers a scalable and interpretable solution for predicting earthquake-induced landslides in complex terrain.

stat.AP

Joint Modelling of Line and Point Data on Metric Graphs

Metric graphs are useful tools for describing spatial domains like road and river networks, where spatial dependence act along the network. We take advantage of recent developments for such Gaussian Random Fields (GRFs), and consider joint spatial modelling of observations with different spatial supports. Motivated by an application to traffic state modelling in Trondheim, Norway, we consider line-referenced data, which can be described by an integral of the GRF along a line segment on the metric graph, and point-referenced data. Through a simulation study inspired by the application, we investigate the number of replicates that are needed to estimate parameters and to predict unobserved locations. The former is assessed using bias and variability, and the latter is assessed through root mean square error (RMSE), continuous rank probability scores (CRPSs), and coverage. Joint modelling is contrasted with a simplified approach that treat line-referenced observations as point-referenced observations. The results suggest joint modelling leads to strong improvements. The application to Trondheim, Norway, combines point-referenced induction loop data and line-referenced public transportation data. To ensure positive speeds, we use a non-linear link function, which requires integrals of non-linear combinations of the linear predictor. This is made computationally feasible by a combination of the R packages inlabru and MetricGraph, and new code for processing geographical line data to work with existing graph representations and fmesher methods for dealing with line support in inlabru on objects from MetricGraph. We fit the model to two datasets where we expect different spatial dependency and compare the results.

stat.ME

Bayesian inference for case-control point pattern data in spatial epidemiology with the \texttt{inlabru} R package

The analysis of case-control point pattern data is an important problem in spatial epidemiology. The spatial variation of cases if often compared to that of a set of controls to assess spatial risk variation as well as the detection of risk factors and exposure to putative pollution sources using spatial regression models. The intensities of the point patterns of cases and controls are estimated using log-Gaussian Cox models, so that fixed and spatial random effects can be included. Bayesian inference is conducted via the integrated Nested Laplace approximation (INLA) method using the inlabru R package. In this way, potential risk factors can be assessed by including them as fixed effects while residual spatial variation is considered as a Gaussian process with Mat\'ern covariance. In addition, exposure to pollution sources is modeled using different smooth terms. The proposed methods have been applied to the Chorley-Ribble dataset, that records the locations of lung and larynx cancer cases as well as the location of an disused old incinerator in the area of Lancashire (England, United Kingdom). Taking the locations of lung cancer as controls, the spatial variation of both types of cases has been estimated and the increase of larynx cases in the vicinity of the incinerator has been assessed. The results are similar to those found in the literature. In a nutshell, a framework for Bayesian analysis of multivariate case-control point patterns within an epidemiological framework has been presented. Models to assess spatial variation and the effect of risk factors and pollution sources can be fit with ease with the inlabru R package.

stat.ME

Validating uncertainty propagation approaches for two-stage Bayesian spatial models using simulation-based calibration

This work tackles the problem of uncertainty propagation in two-stage Bayesian models, with a focus on spatial applications. A two-stage modeling framework has the advantage of being more computationally efficient than a fully Bayesian approach when the first-stage model is already complex in itself, and avoids the potential problem of unwanted feedback effects. Two ways of doing two-stage modeling are the crude plug-in method and the posterior sampling method. The former ignores the uncertainty in the first-stage model, while the latter can be computationally expensive. This paper validates the two aforementioned approaches and proposes a new approach to do uncertainty propagation, which we call the $\mathbf{Q}$ uncertainty method, implemented using the Integrated Nested Laplace Approximation (INLA). We validate the different approaches using the simulation-based calibration method, which tests the self-consistency property of Bayesian models. Results show that the crude plug-in method underestimates the true posterior uncertainty in the second-stage model parameters, while the resampling approach and the proposed method are correct. We illustrate the approaches in a real life data application which aims to link relative humidity and Dengue cases in the Philippines for August 2018.

stat.ME

Coherent Disaggregation and Uncertainty Quantification for Spatially Misaligned Data

Spatial misalignment arises when datasets are aggregated or collected at different spatial scales, leading to information loss. We develop a Bayesian disaggregation framework that links misaligned data to a continuous-domain model through an iteratively linearised integration scheme implemented with the Integrated Nested Laplace Approximation (INLA). The framework accommodates different ways of handling observations depending on the application, resulting in four variants: (i) \textit{Raster at Full Resolution}, (ii) \textit{Raster Aggregation}, (iii) \textit{Polygon Aggregation} (PolyAgg), and (iv) \textit{Point Values} (PointVal). The first three represent increasing levels of spatial averaging, while the last two address situations with incomplete covariate information. For PolyAgg and PointVal, we reconstruct the covariate field using three strategies -- \textit{Value Plugin}, \textit{Joint Uncertainty}, and \textit{Uncertainty Plugin} -- with the latter two propagating uncertainty. We illustrate the framework with an example motivated by landslide modelling, focusing on methodology rather than interpreting landslide processes. Simulations show that uncertainty-propagating approaches outperform \textit{Value Plugin} method and remain robust under model misspecification. Point-pattern observations and full-resolution covariates are therefore preferable, and when covariate fields are incomplete, uncertainty-aware methods are most reliable. The framework is well suited to landslide susceptibility modelling and other spatial mapping tasks, and integrates seamlessly with INLA-based tools.

stat.ME

A parameterization of anisotropic Gaussian fields with penalized complexity priors

Gaussian random fields (GFs) are fundamental tools in spatial modeling and can be represented flexibly and efficiently as solutions to stochastic partial differential equations (SPDEs). The SPDEs depend on specific parameters, which enforce various field behaviors and can be estimated using Bayesian inference. However, even under in-fill asymptotics, the likelihood only provides limited insights into the covariance structure. In response, it is essential to leverage priors to achieve appropriate, meaningful covariance structures in the posterior. This study introduces a smooth, invertible parameterization of the correlation length and diffusion matrix of an anisotropic GF and constructs penalized complexity (PC) priors for the model when the parameters are constant in space. The formulated prior is weakly informative, effectively penalizing complexity by pushing the correlation range toward infinity and the anisotropy to zero.

stat.ME

inlabru: software for fitting latent Gaussian models with non-linear predictors

The integrated nested Laplace approximation (INLA) method has become a popular approach for computationally efficient approximate Bayesian computation. In particular, by leveraging sparsity in random effect precision matrices, INLA is commonly used in spatial and spatio-temporal applications. However, the speed of INLA comes at the cost of restricting the user to the family of latent Gaussian models and the likelihoods currently implemented in {INLA}, the main software implementation of the INLA methodology. {inlabru} is a software package that extends the types of models that can be fitted using INLA by allowing the latent predictor to be non-linear in its parameters, moving beyond the additive linear predictor framework to allow more complex functional relationships. For inference it uses an approximate iterative method based on the first-order Taylor expansion of the non-linear predictor, fitting the model using INLA for each linearised model configuration. {inlabru} automates much of the workflow required to fit models using {R-INLA}, simplifying the process for users to specify, fit and predict from models. There is additional support for fitting joint likelihood models by building each likelihood individually. {inlabru} also supports the direct use of spatial data structures, such as those implemented in the {sf} and {terra} packages. In this paper we outline the statistical theory, model structure and basic syntax required for users to understand and develop their own models using {inlabru}. We evaluate the approximate inference method using a Bayesian method checking approach. We provide three examples modelling simulated spatial data that demonstrate the benefits of the additional flexibility provided by {inlabru}.

stat.ME

A Data Fusion Model for Meteorological Data using the INLA-SPDE method

This work aims to combine two primary meteorological data sources in the Philippines: data from a sparse network of weather stations and outcomes of a numerical weather prediction model. To this end, we propose a data fusion model which is primarily motivated by the problem of sparsity in the observational data and the use of a numerical prediction model as an additional data source in order to obtain better predictions for the variables of interest. The proposed data fusion model assumes that the different data sources are error-prone realizations of a common latent process. The outcomes from the weather stations follow the classical error model while the outcomes of the numerical weather prediction model involves a constant multiplicative bias parameter and an additive bias which is spatially-structured and time-varying. We use a Bayesian model averaging approach with the integrated nested Laplace approximation (INLA) for doing inference. The proposed data fusion model outperforms the stations-only model and the regression calibration approach, when assessed using leave-group-out cross-validation (LGOCV). We assess the benefits of data fusion and evaluate the accuracy of predictions and parameter estimation through a simulation study. The results show that the proposed data fusion model generally gives better predictions compared to the stations-only approach especially with sparse observational data.

stat.AP

Joint model for longitudinal and spatio-temporal survival data

In credit risk analysis, survival models with fixed and time-varying covariates are widely used to predict a borrower's time-to-event. When the time-varying drivers are endogenous, modelling jointly the evolution of the survival time and the endogenous covariates is the most appropriate approach, also known as the joint model for longitudinal and survival data. In addition to the temporal component, credit risk models can be enhanced when including borrowers' geographical information by considering spatial clustering and its variation over time. We propose the Spatio-Temporal Joint Model (STJM) to capture spatial and temporal effects and their interaction. This Bayesian hierarchical joint model reckons the survival effect of unobserved heterogeneity among borrowers located in the same region at a particular time. To estimate the STJM model for large datasets, we consider the Integrated Nested Laplace Approximation (INLA) methodology. We apply the STJM to predict the time to full prepayment on a large dataset of 57,258 US mortgage borrowers with more than 2.5 million observations. Empirical results indicate that including spatial effects consistently improves the performance of the joint model. However, the gains are less definitive when we additionally include spatio-temporal interactions.

q-fin.RM

Bayesian inference of grid cell firing patterns using Poisson point process models with latent oscillatory Gaussian random fields

Questions about information encoded by the brain demand statistical frameworks for inferring relationships between neural firing and features of the world. The landmark discovery of grid cells demonstrates that neurons can represent spatial information through regularly repeating firing fields. However, the influence of covariates may be masked in current statistical models of grid cell activity, which by employing approaches such as discretizing, aggregating and smoothing, are computationally inefficient and do not account for the continuous nature of the physical world. These limitations motivated us to develop likelihood-based procedures for modelling and estimating the firing activity of grid cells conditionally on biologically relevant covariates. Our approach models firing activity using Poisson point processes with latent Gaussian effects, which accommodate persistent inhomogeneous spatial-directional patterns and overdispersion. Inference is performed in a fully Bayesian manner, which allows us to quantify uncertainty. Applying these methods to experimental data, we provide evidence for temporal and local head direction effects on grid firing. Our approaches offer a novel and principled framework for analysis of neural representations of space.

stat.ME

Bayesian modelling of the temporal evolution of seismicity using the ETAS.inlabru R-package

The Epidemic Type Aftershock Sequence (ETAS) model is widely used to model seismic sequences and underpins Operational Earthquake Forecasting (OEF). However, it remains challenging to assess the reliability of inverted ETAS parameters for a range of reasons. The most common algorithms just return point estimates with little quantification of uncertainty, and Bayesian Markov Chain Monte Carlo implementations remain slow to run, do not scale well and few have been extended to include spatial structure. Here we present a new approach to ETAS modelling using an alternative Bayesian method, the Integrated Nested Laplace Approximation (INLA). We have implemented this model in a new R-Package called ETAS.inlabru, which builds on the R packages R-INLA and inlabru . Whilst we just present the temporal component here, the model scales to a spatio-temporal model and may include a variety of spatial covariates. Using a series of synthetic case studies, we explore the robustness of our ETAS inversion method. We demonstrate that reliable estimates of the model parameters require that the catalogue data contains periods of relative quiescence as well as triggered sequences. We explore the robustness under stochastic uncertainty in the training data and show that the method is robust to a wide range of starting conditions. We show how the inclusion of historic earthquakes prior to the modelled domain affects the quality of the inversion. Finally, we show that rate dependent incompleteness after large earthquakes has a significant and detrimental effect on the ETAS posteriors. We believe that the speed of the inlabru inversion, which include a rigorous estimation of uncertainty, will enable a deeper exploration of how to use ETAS robustly for seismicity modelling and operational earthquake forecasting.

stat.AP

Approximation of bayesian Hawkes process models with Inlabru

Hawkes process are very popular mathematical tools for modelling phenomena exhibiting a \textit{self-exciting} or \textit{self-correcting} behaviour. Typical examples are earthquakes occurrence, wild-fires, drought, capture-recapture, crime violence, trade exchange, and social network activity. The widespread use of Hawkes process in different fields calls for fast, reproducible, reliable, easy-to-code techniques to implement such models. We offer a technique to perform approximate Bayesian inference of Hawkes process parameters based on the use of the R-package \inlabru. The \inlabru R-package, in turn, relies on the INLA methodology to approximate the posterior of the parameters. Our Hawkes process approximation is based on a decomposition of the log-likelihood in three parts, which are linearly approximated separately. The linear approximation is performed with respect to the mode of the parameters' posterior distribution, which is determined with an iterative gradient-based method. The approximation of the posterior parameters is therefore deterministic, ensuring full reproducibility of the results. The proposed technique only requires the user to provide the functions to calculate the different parts of the decomposed likelihood, which are internally linearly approximated by the R-package \inlabru. We provide a comparison with the \bayesianETAS R-package which is based on an MCMC method. The two techniques provide similar results but our approach requires two to ten times less computational time to converge, depending on the amount of data.

stat.CO

The SPDE approach for Gaussian and non-Gaussian fields: 10 years and still running

Gaussian processes and random fields have a long history, covering multiple approaches to representing spatial and spatio-temporal dependence structures, such as covariance functions, spectral representations, reproducing kernel Hilbert spaces, and graph based models. This article describes how the stochastic partial differential equation approach to generalising Mat\'ern covariance models via Hilbert space projections connects with several of these approaches, with each connection being useful in different situations. In addition to an overview of the main ideas, some important extensions, theory, applications, and other recent developments are discussed. The methods include both Markovian and non-Markovian models, non-Gaussian random fields, non-stationary fields and space-time fields on arbitrary manifolds, and practical computational considerations.

stat.ME

Quantile based modelling of diurnal temperature range with the five-parameter lambda distribution

Diurnal temperature range is an important variable in climate science that can provide information regarding climate variability and climate change. Changes in diurnal temperature range can have implications for hydrology, human health and ecology, among others. Yet, the statistical literature on modelling diurnal temperature range is lacking. In this paper we propose to model the distribution of diurnal temperature range using the five-parameter lambda (FPL) distribution. Additionally, in order to model diurnal temperature range with explanatory variables, we propose a distributional quantile regression model that combines quantile regression with marginal modelling using the FPL distribution. Inference is performed using the method of quantiles. The models are fitted to 30 years of daily observations of diurnal temperature range from 112 weather stations in the southern part of Norway. The flexible FPL distribution shows great promise as a model for diurnal temperature range, and performs well against competing models. The distributional quantile regression model is fitted to diurnal temperature range data using geographic, orographic and climatological explanatory variables. It performs well and captures much of the spatial variation in the distribution of diurnal temperature range in Norway.

stat.AP

Ranking earthquake forecasts using proper scoring rules: Binary events in a low probability environment

Operational earthquake forecasting for risk management and communication during seismic sequences depends on our ability to select an optimal forecasting model. To do this, we need to compare the performance of competing models with each other in prospective forecasting mode, and to rank their performance using a fair, reproducible and reliable method. The Collaboratory for the Study of Earthquake Predictability (CSEP) conducts such prospective earthquake forecasting experiments around the globe. One metric that has been proposed to rank competing models is the Parimutuel Gambling score, which has the advantage of allowing alarm-based (categorical) forecasts to be compared with probabilistic ones. Here we examine the suitability of this score for ranking competing earthquake forecasts. First, we prove analytically that this score is in general improper, meaning that, on average, it does not prefer the model that generated the data. Even in the special case where it is proper, we show it can still be used in an improper way. Then, we compare its performance with two commonly-used proper scores (the Brier and logarithmic scores), taking into account the uncertainty around the observed average score. We estimate the confidence intervals for the expected score difference which allows us to define if and when a model can be preferred. Our findings suggest the Parimutuel Gambling score should not be used to distinguishing between multiple competing forecasts. They also enable a more rigorous approach to distinguish between the predictive skills of candidate forecasts in addition to their rankings.

stat.AP

A diffusion-based spatio-temporal extension of Gaussian Mat\'ern fields

Gaussian random fields with Mat\'ern covariance functions are popular models in spatial statistics and machine learning. In this work, we develop a spatio-temporal extension of the Gaussian Mat\'ern fields formulated as solutions to a stochastic partial differential equation. The spatially stationary subset of the models have marginal spatial Mat\'ern covariances, and the model also extends to Whittle-Mat\'ern fields on curved manifolds, and to more general non-stationary fields. In addition to the parameters of the spatial dependence (variance, smoothness, and practical correlation range) it additionally has parameters controlling the practical correlation range in time, the smoothness in time, and the type of non-separability of the spatio-temporal covariance. Through the separability parameter, the model also allows for separable covariance functions. We provide a sparse representation based on a finite element approximation, that is well suited for statistical inference and which is implemented in the R-INLA software. The flexibility of the model is illustrated in an application to spatio-temporal modeling of global temperature data.

stat.ME