arXiv ScienceSearch

arXiv subjects

Ben Swallow

Publications and source records attributed to Ben Swallow.

13 recordsLinked to original sources

Comparing Missing Data Methods for Estimating Average Treatment Effects Under Time-Varying Confounding: A Simulation Study

Missing data and confounding are common in real-world statistical applications, yet few studies have examined how imputation methods perform under time-varying confounding in binary variables, or how missingness mechanism, missing rate, missingness location and sample size jointly affect performance and the underlying identifiability conditions. We generated synthetic data and conducted a simulation study comparing missing data methods across scenarios varying these factors. Missingness was introduced in both treatment and outcome variables, and we applied stratified hot deck imputation, single mode imputation, multiple imputation with chained equations (MICE), and complete-case analysis. Average treatment effect (ATE) estimates were obtained using logistic regression with propensity score weighting, and we measured coverage, absolute bias and empirical standard errors across 48 scenarios with 500 replications each. Performance was primarily driven by the missingness mechanism and choice of method, with multiple imputation generally achieving better coverage and lower bias than other methods. Missingness location was also important, while missing rate and sample size primarily affected positivity violations, which were most pronounced under MNAR, high missingness and low sample sizes. Exchangeability violations from mild to moderate confounding were adequately controlled for by propensity score models, whereas strong confounding produced a modest decrease in coverage. Further research should examine additional ways identifiability conditions can be violated under missingness, using more advanced methods and more complex missingness scenarios.

stat.ME

Constructing Contact and Connectivity Matrices for Infectious Disease Modelling

Contact (or mixing, or more generally connectivity) matrices are a fundamental component of modelling and inference for infectious disease epidemiology. Their structure and parametrisation directly accounts for the frequency of interactions between different subpopulations of individuals, as well as having the potential to encode dynamic heterogeneity in these interactions across demographic axes, space and time. Considerable research has been devoted to the structure and estimation of (components of) these matrices to help inform outbreak control and forecast disease spread. In this paper, we review the existing literature on the data types used to construct contact matrices and the methods for incorporating uncertainties and heterogeneities into them. We also highlight remaining challenges and future directions in the use of these contact matrices for epidemiological research.

stat.AP

Discovering Causal Relationships Between Time Series With Spatial Structure

Causal discovery is the subfield of causal inference concerned with estimating the structure of cause-and-effect relationships in a system of interrelated variables, as opposed to quantifying the strength or describing the form of causal effects. As interest in causal discovery builds in fields such as ecology, public health, and environmental sciences where data are regularly collected with spatial and temporal structures, approaches must evolve to manage autocorrelation and complex confounding. As it stands, the few proposed causal discovery algorithms for spatiotemporal data require summarizing across locations, ignore spatial autocorrelation, and/or scale poorly to high dimensions. Here, we introduce our developing framework that extends time-series causal discovery to systems with spatial structure, building upon work on causal discovery across contexts and methods for handling spatial confounding in causal effect estimation. We close by outlining remaining gaps in the literature and directions for future research.

stat.ME

XGBoost meets INLA: a two-stage spatio-temporal forecasting of wildfires in Portugal

Wildfires pose a major threat to Portugal, with over 115,000 hectares burned annually on average during 1980-2024, and the country has faced devastating mega-fires such as those in 2017. Accurate forecasts of wildfire occurrence and burned area are therefore essential for firefighting resource allocation and emergency preparedness. In this study, we propose a novel two-stage ensemble that extends the widely used latent Gaussian modelling framework with integrated nested Laplace approximation (INLA) for spatio-temporal wildfire forecasting. Stage 1 applies a gradient boosting model (XGBoost) to environmental covariates and historical fire records to produce one-month-ahead point forecasts of fire counts and burned area. Stage 2 uses these predictions as external covariates in a latent Gaussian model with additional spatiotemporal random effects to generate probabilistic forecasts of monthly total fire counts and burned area at the council level. To capture both moderate and extreme events, we implement the extended generalised Pareto (eGP) likelihood (a sub-asymptotic distribution) within INLA, develop Penalised Complexity (PC) priors for its parameters, and compare the eGP likelihood with common alternatives (e.g., Gamma and Weibull). Our framework tackles the unavailability of future environmental covariates at prediction time and performs strongly for one-month-ahead forecasts.

stat.AP

Advances in Approximate Bayesian Inference for Models in Epidemiology

Bayesian inference methods are useful in infectious diseases modeling due to their capability to propagate uncertainty, manage sparse data, incorporate latent structures, and address high-dimensional parameter spaces. However, parameter inference through assimilation of observational data in these models remains challenging. While asymptotically exact Bayesian methods offer theoretical guarantees for accurate inference, they can be computationally demanding and impractical for real-time outbreak analysis. This review synthesizes recent advances in approximate Bayesian inference methods that aim to balance inferential accuracy with scalability. We focus on four prominent families: Approximate Bayesian Computation, Bayesian Synthetic Likelihood, Integrated Nested Laplace Approximation, and Variational Inference. For each method, we evaluate its relevance to epidemiological applications, emphasizing innovations that improve both computational efficiency and inference accuracy. We also offer practical guidance on method selection across a range of modeling scenarios. Finally, we identify hybrid exact approximate inference as a promising frontier that combines methodological rigor with the scalability needed for the response to outbreaks. This review provides epidemiologists with a conceptual framework to navigate the trade-off between statistical accuracy and computational feasibility in contemporary disease modeling.

stat.ME

Gaussian process modelling of infectious diseases using the Greta software package and GPUs

Gaussian process are a widely-used statistical tool for conducting non-parametric inference in applied sciences, with many computational packages available to fit to data and predict future observations. We study the use of the Greta software for Bayesian inference to apply Gaussian process regression to spatio-temporal data of infectious disease outbreaks and predict future disease spread. Greta builds on Tensorflow, making it comparatively easy to take advantage of the significant gain in speed offered by GPUs. In these complex spatio-temporal models, we show a reduction of up to 70\% in computational time relative to fitting the same models on CPUs. We show how the choice of covariance kernel impacts the ability to infer spread and extrapolate to unobserved spatial and temporal units. The inference pipeline is applied to weekly incidence data on tuberculosis in the East and West Midlands regions of England over a period of two years.

stat.CO

Gradient-Based Approximate Bayesian Inference with Entropy-Optimized Summary Statistics for Compartmental Models

Recent pandemics have highlighted the critical role of infectious disease models in guiding public health decision-making, driving demand for realistic models that can provide timely answers under uncertainty. Compartmental models are widely used to capture disease dynamics, and advances in data availability, computational resources, and epidemiological understanding have allowed the development of models that incorporate detailed representations of population structure, disease progression, and intervention effects. While these improvements improve model fidelity, they also increase model complexity, leading to high-dimensional parameter spaces, intractable likelihoods, and computational challenges for fitting models to limited surveillance data in real time. Existing likelihood-free methods, such as Approximate Bayesian Computation (ABC) and Bayesian Synthetic Likelihood (BSL), have developed largely independently, each with distinct strengths and limitations. We propose an integrated three-stage framework that synthesizes advances from both likelihood-based and likelihood-free method: (1) ABC-based entropy minimization to identify low-dimensional, approximately orthogonal summary statistics; (2) BSL inference using these optimized summaries to construct tractable Gaussian approximations; and (3) Hamiltonian Monte Carlo sampling for efficient posterior exploration. Through SEIR simulation study and application to the 1978 British boarding school influenza outbreak, we demonstrate that our framework achieves reliable parameter estimation and uncertainty quantification while maintaining computational efficiency.

stat.ME

A Bayesian multivariate extreme value mixture model

Impact assessment of natural hazards requires the consideration of both extreme and non-extreme events. Extensive research has been conducted on the joint modeling of bulk and tail in univariate settings; however, the corresponding body of research in the context of multivariate analysis is comparatively scant. This study extends the univariate joint modeling of bulk and tail to the multivariate framework. Specifically, it pertains to cases where multivariate observations exceed a high threshold in at least one component. We propose a multivariate mixture model that assumes a parametric model to capture the bulk of the distribution, which is in the max-domain of attraction (MDA) of a multivariate extreme value distribution (mGEVD). The tail is described by the multivariate generalized Pareto distribution, which is asymptotically justified to model multivariate threshold exceedances. We show that if all components exceed the threshold, our mixture model is in the MDA of an mGEVD. Bayesian inference based on multivariate random-walk Metropolis-Hastings and the automated factor slice sampler allows us to incorporate uncertainty from the threshold selection easily. Due to computational limitations, simulations and data applications are provided for dimension $d=2$, but a discussion is provided with views toward scalability based on pairwise likelihood.

stat.ME

Data Fusion in a Two-stage Spatio-Temporal Model using the INLA-SPDE Approach

This paper proposes a two-stage estimation approach for a spatial misalignment scenario that is motivated by the epidemiological problem of linking pollutant exposures and health outcomes. We use the integrated nested Laplace approximation method to estimate the parameters of a two-stage spatio-temporal model; the first stage models the exposures while the second stage links the health outcomes to exposures. The first stage is based on the Bayesian melding model, which assumes a common latent field for the geostatistical monitors data and a high-resolution data such as satellite data. The second stage fits a GLMM using the spatial averages of the estimated latent field, and additional spatial and temporal random effects. Uncertainty from the first stage is accounted for by simulating repeatedly from the posterior predictive distribution of the latent field. A simulation study was carried out to assess the impact of the sparsity of the data on the monitors, number of time points, and the specification of the priors in terms of the biases, RMSEs, and coverage probabilities of the parameters and the block-level exposure estimates. The results show that the parameters are generally estimated correctly but there is difficulty in estimating the latent field parameters. The method works very well in estimating block-level exposures and the effect of exposures on the health outcomes, which is the primary parameter of interest for spatial epidemiologists and health policy makers, even with the use of non-informative priors.

stat.ME

Bayesian inference for stochastic oscillatory systems using the phase-corrected Linear Noise Approximation

Likelihood-based inference in stochastic non-linear dynamical systems, such as those found in chemical reaction networks and biological clock systems, is inherently complex and has largely been limited to small and unrealistically simple systems. Recent advances in analytically tractable approximations to the underlying conditional probability distributions enable long-term dynamics to be accurately modelled, and make the large number of model evaluations required for exact Bayesian inference much more feasible. We propose a new methodology for inference in stochastic non-linear dynamical systems exhibiting oscillatory behaviour and show the parameters in these models can be realistically estimated from simulated data. Preliminary analyses based on the Fisher Information Matrix of the model can guide the implementation of Bayesian inference. We show that this parameter sensitivity analysis can predict which parameters are practically identifiable. Several Markov chain Monte Carlo algorithms are compared, with our results suggesting a parallel tempering algorithm consistently gives the best approach for these systems, which are shown to frequently exhibit multi-modal posterior distributions.

stat.CO

Visualization for Epidemiological Modelling: Challenges, Solutions, Reflections & Recommendations

We report on an ongoing collaboration between epidemiological modellers and visualization researchers by documenting and reflecting upon knowledge constructs -- a series of ideas, approaches and methods taken from existing visualization research and practice -- deployed and developed to support modelling of the COVID-19 pandemic. Structured independent commentary on these efforts is synthesized through iterative reflection to develop: evidence of the effectiveness and value of visualization in this context; open problems upon which the research communities may focus; guidance for future activity of this type; and recommendations to safeguard the achievements and promote, advance, secure and prepare for future collaborations of this kind. In describing and comparing a series of related projects that were undertaken in unprecedented conditions, our hope is that this unique report, and its rich interactive supplementary materials, will guide the scientific community in embracing visualization in its observation, analysis and modelling of data as well as in disseminating findings. Equally we hope to encourage the visualization community to engage with impactful science in addressing its emerging data challenges. If we are successful, this showcase of activity may stimulate mutually beneficial engagement between communities with complementary expertise to address problems of significance in epidemiology and beyond. https://ramp-vis.github.io/RAMPVIS-PhilTransA-Supplement/

cs.HC

Tracking the national and regional COVID-19 epidemic status in the UK using directed Principal Component Analysis

One of the difficulties in monitoring an ongoing pandemic is deciding on the metric that best describes its status when multiple intercorrelated measurements are available. Having a single measure, such as the effective reproduction number R, has been a simple and useful metric for tracking the epidemic and for imposing policy interventions to curb the increase when R >1. While R is easy to interpret in a fully susceptible population, it is more difficult to interpret for a population with heterogeneous prior immunity, e.g., from vaccination and prior infection. We propose an additional metric for tracking the UK epidemic which can capture the different spatial scales. These are the principal scores (PCs) from a weighted Principal Component Analysis. In this paper, we have used the methodology across the four UK nations and across the first two epidemic waves (January 2020-March 2021) to show that first principal score across nations and epidemic waves is a representative indicator of the state of the pandemic and are correlated with the trend in R. Hospitalisations are shown to be consistently representative, however, the precise dominant indicator, i.e. the principal loading(s) of the analysis, can vary geographically and across epidemic waves.

stat.AP

Parallel tempering as a mechanism for facilitating inference in hierarchical hidden Markov models

The study of animal behavioural states inferred through hidden Markov models and similar state switching models has seen a significant increase in popularity in recent years. The ability to account for varying levels of behavioural scale has become possible through hierarchical hidden Markov models, but additional levels lead to higher complexity and increased correlation between model components. Maximum likelihood approaches to inference using the EM algorithm and direct optimisation of likelihoods are more frequently used, with Bayesian approaches being less favoured due to computational demands. Given these demands, it is vital that efficient estimation algorithms are developed when Bayesian methods are preferred. We study the use of various approaches to improve convergence times and mixing in Markov chain Monte Carlo methods applied to hierarchical hidden Markov models, including parallel tempering as an inference facilitation mechanism. The method shows promise for analysing complex stochastic models with high levels of correlation between components, but our results show that it requires careful tuning in order to maximise that potential.

stat.CO