arXiv ScienceSearch

arXiv subjects

Alex Stringer

Publications and source records attributed to Alex Stringer.

18 recordsLinked to original sources

Fast Computation of Nested Cross-Validation for Penalized Regression

Cross-validation is a resampling procedure that provides a point estimate of generalization error for any predictive model. Cross-validation is widely used for model selection and evaluation. Uncertainty in the cross-validation estimate is challenging to quantify, and estimation of its variance is known to require multiple runs of the entire resampling procedure. Nested cross-validation computes a prediction interval for the generalization error for a given model and training dataset by resampling the entire cross-validation procedure but incurs extraordinary computational cost. We provide an efficient method for computing the nested cross-validation prediction interval using only a single model fit for some penalized regression models including ridge regression, spline smoothing, and some functional regression models. We characterize when our proposed method should be expected to out-perform resampling-based nested cross-validation in various scaling regimes as well as in finite samples. Experiments for functional principle components regression demonstrate non-trivial cases in which our proposed method improves run times substantially.

stat.CO

Testing linear combinations of multiple variance components

We test the hypothesis that simulataneous linear contrasts of multiple variance components equal zero in a Gaussian variance components model via a parametric bootstrap. Applications include but are not limited to nested and crossed designs. The main technical contributions are a computationally efficient decomposition of the normalized residual log-likelihood that does not require the variance components to be non-negative or variance design matrices to be positive semi-definite, a modified Newton method for its minimization, and a method for efficient optimization and sampling under the null hypothesis that certain linear combinations of variance components equal zero. A special case of the proposed procedure is a test for multiple variance components simulataneously equalling zero, for which a likelihood ratio test was not previously available. However, the proposed procedure is significantly more general.

stat.ME

A Unified Framework for Multiple Exposure Distributed Lag Non-Linear Models for Air Pollution Epidemiology

This study quantifies the association between air pollution and mortality in Ontario, Canada. Exposure-response relationships in air pollution epidemiology are complex due to three features: time-lagged associations, non-linear associations, and multiple pollutants. To address the first two features, two distinct classes of distributed lag non-linear model (DLNM) have been proposed, but extending them to multiple exposures and selecting an appropriate model remain challenging. We propose a unified framework for multiple exposure DLNMs, integrating model specification, estimation, selection and stacking. The framework applies to four different model structures: two additive and two proposed single-index DLNMs, all applicable to general outcome types, including the mortality counts in the motivating application. We develop an estimation approach that applies to all four models. Choosing among the candidate DLNMs is challenging a priori, and we derive an AIC to select among them. As an alternative to selecting a single model, we also extend a model stacking approach to combine inferences across the four DLNMs and propose an implementation scalable to our dataset with 106,346 observations. In the motivating analysis, the four DLNMs yield different estimates, and the proposed stacking approach identifies significant associations between respiratory mortality and a mixture of PM2.5, O3 and NO2.

stat.ME

Fast approximate Bayesian inference of HIV indicators using PCA adaptive Gauss-Hermite quadrature

Naomi is a spatial evidence synthesis model used to produce district-level HIV epidemic indicators in sub-Saharan Africa. Multiple outcomes of policy interest, including HIV prevalence, HIV incidence, and antiretroviral therapy treatment coverage are jointly modelled using both household survey data and routinely reported health system data. The model is provided as a tool for countries to input their data to and generate estimates with during a yearly process supported by UNAIDS. Previously, inference has been conducted using empirical Bayes and a Gaussian approximation, implemented via the TMB R package. We propose a new inference method based on an extension of adaptive Gauss-Hermite quadrature to deal with more than 20 hyperparameters. Using data from Malawi, our method improves the accuracy of inferences for model parameters, while being substantially faster to run than Hamiltonian Monte Carlo with the No-U-Turn sampler. Our implementation leverages the existing TMB C++ template for the model's log-posterior, and is compatible with any model with such a template.

stat.ME

Candidate Dark Galaxy-2: Validation and Analysis of an Almost Dark Galaxy in the Perseus Cluster

Candidate Dark Galaxy-2 (CDG-2) is a potential dark galaxy consisting of four globular clusters (GCs) in the Perseus cluster, first identified in Li et al. (2025) through a sophisticated statistical method. The method searched for over-densities of GCs from a \textit{Hubble Space Telescope} (\textit{HST}) survey targeting Perseus. Using the same \textit{HST} images and the new imaging data from the \textit{Euclid} survey, we report the detection of extremely faint but significant diffuse emission around the four GCs of CDG-2. We thus have exceptionally strong evidence that CDG-2 is a galaxy. This is the first galaxy detected purely through its GC population. Under the conservative assumption that the four GCs make up the entire GC population, preliminary analysis shows that CDG-2 has a total luminosity of $L_{V, \mathrm{gal}}= 6.2\pm{3.0} \times 10^6 L_{\odot}$ and a minimum GC luminosity of $L_{V, \mathrm{GC}}= 1.03\pm{0.2}\times 10^6 L_{\odot}$. Our results indicate that CDG-2 is one of the faintest galaxies having associated GCs, while at least $\sim 16.6\%$ of its light is contained in its GC population. This ratio is likely to be much higher ($\sim 33\%$) if CDG-2 has a canonical GC luminosity function (GCLF). In addition, if the previously observed GC-to-halo mass relations apply to CDG-2, it would have a minimum dark matter halo mass fraction of $99.94\%$ to $99.98\%$. If it has a canonical GCLF, then the dark matter halo mass fraction is $\gtrsim 99.99\%$. Therefore, CDG-2 may be the most GC dominated galaxy and potentially one of the most dark matter dominated galaxies ever discovered.

astro-ph.GA

Estimating Associations Between Cumulative Exposure and Health via Generalized Distributed Lag Non-Linear Models using Penalized Splines

Quantifying associations between short-term exposure to ambient air pollution and health outcomes is an important public health priority. Many studies have investigated the association considering delayed effects within the past few days. Adaptive cumulative exposure distributed lag non-linear models (ACE-DLNMs) quantify associations between health outcomes and cumulative exposure that is specified in a data-adaptive way. While the ACE-DLNM framework is highly interpretable, it is limited to continuous outcomes and does not scale well to large datasets. Motivated by a large analysis of daily pollution and respiratory hospitalization counts in Canada between 2001 and 2018, we propose a generalized ACE-DLNM incorporating penalized splines, improving upon existing ACE-DLNM methods to accommodate general response types. We then develop a computationally efficient estimation strategy based on profile likelihood and Laplace approximate marginal likelihood with Newton-type methods. We demonstrate the performance and practical advantages of the proposed method through simulations. In application to the motivating analysis, the proposed method yields more stable inferences compared to generalized additive models with fixed exposures, while retaining interpretability.

stat.ME

Case-crossover designs and overdispersion with application in air pollution epidemiology

Over the last three decades, case-crossover designs have found many applications in health sciences, especially in air pollution epidemiology. They are typically used, in combination with partial likelihood techniques, to define a conditional logistic model for the responses, usually health outcomes, conditional on the exposures. Despite the fact that conditional logistic models have been shown equivalent, in typical air pollution epidemiology setups, to specific instances of the well-known Poisson time series model, it is often claimed that they cannot allow for overdispersion. This paper clarifies the relationship between case-crossover designs, the models that ensue from their use, and overdispersion. In particular, we propose to relax the assumption of independence between individuals traditionally made in case-crossover analyses, in order to explicitly introduce overdispersion in the conditional logistic model. As we show, the resulting overdispersed conditional logistic model coincides with the overdispersed, conditional Poisson model, in the sense that their likelihoods are simple re-expressions of one another. We further provide the technical details of a Bayesian implementation of the proposed case-crossover model, which we use to demonstrate, by means of a large simulation study, that standard case-crossover models can lead to dramatically underestimated coverage probabilities, while the proposed models do not. We also perform an illustrative analysis of the association between air pollution and morbidity in Toronto, Canada, which shows that the proposed models are more robust than standard ones to outliers such as those associated with public holidays.

stat.ME

Semi-parametric Benchmark Dose Analysis with Monotone Additive Models

Benchmark dose analysis aims to estimate the level of exposure to a toxin that results in a clinically-significant adverse outcome and quantifies uncertainty using the lower limit of a confidence interval for this level. We develop a novel framework for benchmark dose analysis based on monotone additive dose-response models. We first introduce a flexible approach for fitting monotone additive models via penalized B-splines and Laplace-approximate marginal likelihood. A reflective Newton method is then developed that employs de Boor's algorithm for computing splines and their derivatives for efficient estimation of the benchmark dose. Finally, we develop and assess three approaches for calculating benchmark dose lower limits: a naive one based on asymptotic normality of the estimator, one based on an approximate pivot, and one using a Bayesian parametric bootstrap. The latter approaches improve upon the naive method in terms of accuracy and are guaranteed to return a positive lower limit; the approach based on an approximate pivot is typically an order of magnitude faster than the bootstrap, although they are both practically feasible to compute. We apply the new methods to make inferences about the level of prenatal alcohol exposure associated with clinically significant cognitive defects in children using data from an NIH-funded longitudinal study. Software to reproduce the results in this paper is available at https://github.com/awstringer1/bmd-paper-code.

stat.ME

Exact Gradient Evaluation for Adaptive Quadrature Approximate Marginal Likelihood in Mixed Models for Grouped Data

A method is introduced for approximate marginal likelihood inference via adaptive Gaussian quadrature in mixed models with a single grouping factor. The core technical contribution is an algorithm for computing the exact gradient of the approximate log-marginal likelihood. This leads to efficient maximum likelihood via quasi-Newton optimization that is demonstrated to be faster than existing approaches based on finite-differenced gradients or derivative-free optimization. The method is specialized to Bernoulli mixed models with multivariate, correlated Gaussian random effects; here computations are performed using an inverse log-Cholesky parameterization of the Gaussian density that involves no matrix decomposition during model fitting, while Wald confidence intervals are provided for variance parameters on the original scale. Simulations give evidence of these intervals attaining nominal coverage if enough quadrature points are used, for data comprised of a large number of very small groups exhibiting large between-group heterogeneity. The Laplace approximation is well-known to give especially poor coverage and high bias for data comprised of a large number of small groups. Adaptive quadrature mitigates this, and the methods in this paper improve the computational feasibility of this more accurate method. All results may be reproduced using code available at \url{https://github.com/awstringer1/aghmm-paper-code}.

stat.ME

Poisson Cluster Process Models for Detecting Ultra-Diffuse Galaxies

We propose a novel set of Poisson Cluster Process (PCP) models to detect Ultra-Diffuse Galaxies (UDGs), a class of extremely faint, enigmatic galaxies of substantial interest in modern astrophysics. We model the unobserved UDG locations as parent points in a PCP, and infer their positions based on the observed spatial point patterns of their old star cluster systems. Many UDGs have somewhere from a few to hundreds of these old star clusters, which we treat as offspring points in our models. We also present a new framework to construct a marked PCP model using the marks of star clusters. The marked PCP model may enhance the detection of UDGs and offers broad applicability to problems in other disciplines. To assess the overall model performance, we design an innovative assessment tool for spatial prediction problems where only point-referenced ground truth is available, overcoming the limitation of standard ROC analyses where spatial Boolean reference maps are required. We construct a bespoke blocked Gibbs adaptive spatial birth-death-move MCMC algorithm to infer the locations of UDGs using real data from a \textit{Hubble Space Telescope} imaging survey. Based on our performance assessment tool, our novel models significantly outperform existing approaches using the Log-Gaussian Cox Process. We also obtained preliminary evidence that the marked PCP model improves UDG detection performance compared to the model without marks. Furthermore, we find evidence of a potential new ``dark galaxy'' that was not detected by previous methods.

astro-ph.GA

Model-based Smoothing with Integrated Wiener Processes and Overlapping Splines

In many applications that involve the inference of an unknown smooth function, the inference of its derivatives will often be just as important as that of the function itself. To make joint inferences of the function and its derivatives, a class of Gaussian processes called $p^{\text{th}}$ order Integrated Wiener's Process (IWP), is considered. Methods for constructing a finite element (FEM) approximation of an IWP exist but have focused only on the order $p = 2$ case which does not allow appropriate inference for derivatives, and their computational feasibility relies on additional approximation to the FEM itself. In this article, we propose an alternative FEM approximation, called overlapping splines (O-spline), which pursues computational feasibility directly through the choice of test functions, and mirrors the construction of an IWP as the Ospline results from the multiple integrations of these same test functions. The O-spline approximation applies for any order $p \in \mathbb{Z}^+$, is computationally efficient and provides consistent inference for all derivatives up to order $p-1$. It is shown both theoretically, and empirically through simulation, that the O-spline approximation converges to the true IWP as the number of knots increases. We further provide a unified and interpretable way to define priors for the smoothing parameter based on the notion of predictive standard deviation (PSD), which is invariant to the order $p$ and the placement of the knot. Finally, we demonstrate the practical use of the O-spline approximation through simulation studies and an analysis of COVID death rates where the inference is carried on both the function and its derivatives where the latter has an important interpretation in terms of the course of the pandemic.

stat.ME

Flexible Marginal Models for Dependent Data

Models for dependent data are distinguished by their targets of inference. Marginal models are useful when interest lies in quantifying associations averaged across a population of clusters. When the functional form of a covariate-outcome association is unknown, flexible regression methods are needed to allow for potentially non-linear relationships. We propose a novel marginal additive model (MAM) for modelling cluster-correlated data with non-linear population-averaged associations. The proposed MAM is a unified framework for estimation and uncertainty quantification of a marginal mean model, combined with inference for between-cluster variability and cluster-specific prediction. We propose a fitting algorithm that enables efficient computation of standard errors and corrects for estimation of penalty terms. We demonstrate the proposed methods in simulations and in application to (i) a longitudinal study of beaver foraging behaviour, and (ii) a spatial analysis of Loaloa infection in West Africa. R code for implementing the proposed methodology is available at https://github.com/awstringer1/mam.

stat.ME

Asymptotics of numerical integration for two-level mixed models

We study mixed models with a single grouping factor, where inference about unknown parameters requires optimizing a marginal likelihood defined by an intractable integral. Low-dimensional numerical integration techniques are regularly used to approximate these integrals, with inferences about parameters based on the resulting approximate marginal likelihood. For a generic class of mixed models that satisfy explicit regularity conditions, we derive the stochastic relative error rate incurred for both the likelihood and maximum likelihood estimator when adaptive numerical integration is used to approximate the marginal likelihood. We then specialize the analysis to well-specified generalized linear mixed models having exponential family response and multivariate Gaussian random effects, verifying that the regularity conditions hold, and hence that the convergence rates apply. We also prove that for models with likelihoods satisfying very weak concentration conditions that the maximum likelihood estimators from non-adaptive numerical integration approximations of the marginal likelihood are not consistent, further motivating adaptive numerical integration as the preferred tool for inference in mixed models. Code to reproduce the simulations in this paper is provided at https://github.com/awstringer1/aq-theory-paper-code.

stat.ME

Fast, Scalable Approximations to Posterior Distributions in Extended Latent Gaussian Models

We define a novel class of additive models, called Extended Latent Gaussian Models, that allow for a wide range of response distributions and flexible relationships between the additive predictor and mean response. The new class covers a broad range of interesting models including multi-resolution spatial processes, partial likelihood-based survival models, and multivariate measurement error models. Because computation of the exact posterior distribution is infeasible, we develop a fast, scalable approximate Bayesian inference methodology for this class based on nested Gaussian, Laplace, and adaptive quadrature approximations. We prove that the error in these approximate posteriors is op(1) under standard conditions, and provide numerical evidence suggesting that our method runs faster and scales to larger datasets than methods based on Integrated Nested Laplace Approximations and Markov Chain Monte Carlo, with comparable accuracy. We apply the new method to the mapping of malaria incidence rates in continuous space using aggregated data, mapping leukaemia survival hazards using a Cox Proportional-Hazards model with a continuously-varying spatial process, and estimating the mass of the Milky Way Galaxy using noisy multivariate measurements of the positions and velocities of star clusters in its orbit.

stat.ME

Stochastic Convergence Rates and Applications of Adaptive Quadrature in Bayesian Inference

We provide the first stochastic convergence rates for a family of adaptive quadrature rules used to normalize the posterior distribution in Bayesian models. Our results apply to the uniform relative error in the approximate posterior density, the coverage probabilities of approximate credible sets, and approximate moments and quantiles, therefore guaranteeing fast asymptotic convergence of approximate summary statistics used in practice. The family of quadrature rules includes adaptive Gauss-Hermite quadrature, and we apply this rule in two challenging low-dimensional examples. Further, we demonstrate how adaptive quadrature can be used as a crucial component of a modern approximate Bayesian inference procedure for high-dimensional additive models. The method is implemented and made publicly available in the aghq package for the R language, available on CRAN.

stat.ME

Implementing Approximate Bayesian Inference using Adaptive Quadrature: the aghq Package

The aghq package for implementing approximate Bayesian inference using adaptive quadrature is introduced. The method and software are described, and use of the package in making approximate Bayesian inferences in several challenging low- and high-dimensional models is illustrated. Examples include an infectious disease model; an astrostatistical model for estimating the mass of the Milky Way; two examples in non-Gaussian model-based geostatistics including one incorporating zero-inflation which is not easily fit using other methods; and a model for zero-inflated, overdispersed count data. The aghq package is especially compatible with the popular TMB interface for automatic differentiation and Laplace approximation, and existing users of that software can make approximate Bayesian inferences with aghq using very little additional code. The aghq package is available from CRAN and complete code for all examples in this paper can be found at https://github.com/awstringer1/aghq-software-paper-code.

stat.CO