arXiv ScienceSearch

arXiv subjects

Jonas Wallin

Publications and source records attributed to Jonas Wallin.

At least 19 recordsLinked to original sources

A bridge representation of Gaussian Whittle-Matérn fields on compact metric graphs

Gaussian Whittle-Matérn fields form a flexible class of Gaussian processes on compact metric graphs, where spatial dependence is governed by the geometry and connectivity of the network through a fractional-order stochastic partial differential equation. This paper develops a new bridge representation of these fields in the case of half-integer smoothness parameters, when the fields have Markov properties. This representation decomposes the field into a finite-dimensional graph component and independent Whittle-Matérn bridge processes on the individual edges. The resulting decomposition leads to efficient likelihood evaluation, kriging prediction, and simulation methods. We show that this improves numerical stability and can greatly reduce computation time compared to previous methods. A simulation study on a Chicago street-network graph illustrates the computational efficiency of the sampling method and an application to Madrid traffic intensity data demonstrates the practical gains for likelihood-based inference and prediction. The methods are implemented in the R package MetricGraph.

stat.CO

Gaussian Processes on Directed Metric Graphs

We introduce a statistical framework for Gaussian fields indexed at arbitrary edge locations on general compact directed metric graphs. The construction is based on a stochastic differential equation with a first-order operator and conditions at the vertices. We characterise well-posedness and identify the covariance reproducing kernel Hilbert space. We also connect the proposed framework to earlier stream-network models, showing that these arise from the same system under particular boundary conditions, and introduce new boundary conditions that yield more physically realistic processes. The differential-equation representation enables computationally efficient inference and prediction. This makes the method applicable to large data sets without approximation. Applications to temperature modelling on river networks and traffic speeds on road networks illustrate the framework, including the computational efficiency and improved performance under physically informed vertex conditions.

stat.ME

Bivariate geostatistical latent variable models for the analysis of antibody density data

The increasing availability of serosurveys that measure antibody responses to multiple antigens requires the development of methods that can exploit the full information content of such data, both biological and spatial. However, the non-Gaussian and potentially multimodal distributional behaviour of antibody responses makes the development of such methods inherently complex, especially in a multivariate setting. Here, we extend the latent variable framework of Giorgi and Wallin (2026), in which continuous antibody concentrations are modelled through an individual-level latent seroreactivity process that represents the level of immune activation to a given antigen. We focus primarily on the bivariate setting and set out a series of guiding principles that justify the resulting joint modelling structure. The proposed model captures distinct sources of correlation between antibody responses, arising both from shared exposure to the same environment and from biological processes occurring within the same host. Spatial dependence is introduced through a novel bivariate Matérn random field, which we use to construct a parsimonious class of cross-covariance functions between antigen-specific spatial processes. We illustrate the application of the framework to analyse data on bivariate antibody measurements from a malaria serosurvey in the Kenyan highlands. Results from the application and a simulation study show that ignoring this correlation substantially degrades inference on joint properties of the antibody distributions and on individual-level seroreactivity, but matters less when interest lies exclusively in each antibody's marginal distribution. Finally, we discuss how the framework could be extended to settings with more than two antigens, and highlight the modelling challenges that arise as the number of antigens grows.

stat.ME

Mesh Invariant Infinite Dimensional Adaptive MCMC for Latent Gaussian Processes

We introduce mesh-invariant adaptive Markov chain Monte Carlo methods for Gaussian-process posteriors arising in infinite-dimensional Bayesian inference. In function-space MCMC, posterior distributions are defined by a change of measure with respect to a Gaussian prior, making absolute continuity essential for valid proposal construction. Standard adaptive schemes that modify means or scales in the full discretized space can destroy this property, leading to proposal measures that become singular in the infinite-dimensional limit. To avoid this, we adapt only an active finite-dimensional subspace of the Gaussian-process representation while preserving the prior dynamics on inactive coordinates. This yields two adaptive proposals, pCNLV and pCNMV, which extend preconditioned Crank--Nicolson and Crank--Nicolson Langevin methods by learning posterior scale, and in pCNMV also posterior mean structure, on the data-informed subspace without introducing discretization-dependent Gaussian density ratios. The resulting samplers retain the mesh robustness of function-space methods while improving efficiency through local adaptation. Experiments on a Darcy-flow inverse problem and Bayesian logistic regression demonstrate consistent efficiency gains, including an approximately fourfold improvement in effective sampling efficiency for Darcy flow.

stat.CO

Hierarchical Bayesian Estimation of Covariance Matrices

We develop a hierarchical Bayesian framework for covariance matrix estimation built on a key observation: while equivariance under the full general linear group GL(p) is well known, it is an extremely restrictive property -- estimators equivariant to GL(p) are limited to scalar multiples of the sample covariance matrix and carry considerably larger risks than shrinkage estimators. By contrast, commonly used shrinkage estimators, including the Haff empirical Bayes estimator, and the Ledoit--Wolf estimators, are all equivariant under the smaller orthogonal group O(p). Exploiting this structure, we establish that the Haar measure Bayes rule in an oracle eigenvalue model is the minimum risk estimator within the class of O(p)-equivariant estimators, and derive oracle Bayes rules for the covariance and precision matrices under the squared Frobenius, Stein, and squared Stein loss functions. These oracle rules serve as theoretical benchmarks that dominate all commonly used estimators. To approximate them when the true eigenvalues are unknown, we introduce a hierarchical Bayes model that places a finite P'olya tree prior on the eigenvalue distribution and uses Gibbs sampling to generate posterior draws, yielding both shrinkage estimates for the eigenvalues and approximations to the oracle Bayes rules. Simulations suggest that the finite P'olya tree prior is able to recover the general form of the distribution of the eigenvalues, and confirm that the resulting estimators closely approach oracle performance, substantially outperforming classical competitors for both covariance and precision matrix estimation.

stat.ME

Efficient Solvers for SLOPE in R, Python, Julia, and C++

We present a suite of packages in R, Python, Julia, and C++ that efficiently solve the Sorted L-One Penalized Estimation (SLOPE) problem. The packages feature a highly efficient hybrid coordinate descent algorithm that fits generalized linear models (GLMs) and supports a variety of loss functions, including Gaussian, binomial, Poisson, and multinomial logistic regression. Our implementation is designed to be fast, memory-efficient, and flexible. The packages support a variety of data structures (dense, sparse, and out-of-memory matrices) and are designed to efficiently fit the full SLOPE path as well as handle cross-validation of SLOPE models, including the relaxed SLOPE. We present examples of how to use the packages and benchmarks that demonstrate the performance of the packages on both real and simulated data and show that our packages outperform existing implementations of SLOPE in terms of speed.

stat.CO

Identifying Network Hubs with the Partial Correlation Graphical LASSO

Graphical LASSO (GLASSO) is a widely used method for estimating sparse precision matrices and learning undirected graphical models in high-dimensional settings. Because GLASSO penalizes entries of the precision matrix directly, however, it is not scale-invariant. Partial Correlation Graphical LASSO (PCGLASSO), introduced by Carter et al. (2024), addresses this limitation by penalizing partial correlations, which directly characterize conditional dependence. In this paper, we study both statistical and computational properties of the PCGLASSO estimator. Our main contribution is the introduction of a scale-invariant irrepresentability condition for PCGLASSO and the proof that this condition is sufficient for consistent model selection. We further show that this condition is weaker than the corresponding irrepresentability condition for GLASSO, helping to explain the improved empirical behavior of PCGLASSO in settings such as hub-structured graphs. In addition, we develop two efficient algorithms for computing the estimator and analyze the nonconvex optimization problem underlying PCGLASSO, deriving conditions for global uniqueness and showing consistency of all minimizers.

math.ST

Adaptive Riemannian Manifold Hamiltonian Monte Carlo with Hierarchical Metric

Hamiltonian Monte Carlo (HMC) and its dynamic extensions, such as the No-U-Turn Sampler (NUTS), are powerful Markov chain Monte Carlo methods for sampling from complex, high-dimensional probability distributions. Riemannian manifold Hamiltonian Monte Carlo (RMHMC) extends HMC by allowing the mass matrix to depend on position, which can substantially improve mixing but also makes implementation considerably more challenging. In this paper, we study an adaptive hierarchical version of RMHMC that is well suited to many hierarchical sampling problems. A key feature of hierarchical RMHMC is that, unlike general RMHMC, it admits a closed-form explicit leapfrog integrator, enabling efficient implementation and direct use within dynamic HMC methods such as NUTS. We introduce an adaptive scheme that automatically tunes the parameters of the hierarchical mass matrix during simulation. Importantly, the target density need not exhibit any hierarchical or block structure; the hierarchy is instead imposed on the mass matrix as a modeling device to capture the local geometry of the target distribution. Numerical experiments demonstrate appealing empirical performance in high-dimensional Bayesian inference problems.

stat.CO

Asymptotic Theory for Graphical SLOPE: Precision Estimation and Pattern Convergence

This paper studies Graphical SLOPE for precision matrix estimation, with emphasis on its ability to recover both sparsity and clusters of edges with equal or similar strength. In a fixed-dimensional regime, we establish that the root-$n$ scaled estimation error converges to the unique minimizer of a strictly convex optimization problem defined through the directional derivative of the SLOPE penalty. We also establish convergence of the induced SLOPE pattern, thereby obtaining an asymptotic characterization of the clustering structure selected by the estimator. A comparison with GLASSO shows that the grouping property of SLOPE can substantially improve estimation accuracy when the precision matrix exhibits structured edge patterns. To assess the effect of departures from Gaussianity, we then analyze Gaussian-loss precision matrix estimation under elliptical distributions. In this setting, we derive the limiting distribution and quantify the inflation in variability induced by heavy tails relative to the Gaussian benchmark. We also study TSLOPE, based on the multivariate $t$-loss, and derive its limiting distribution. The results show that TSLOPE offers clear advantages over GSLOPE under heavy-tailed data-generating mechanisms. Simulation evidence suggests that these qualitative conclusions persist in high-dimensional settings, and an empirical application shows that SLOPE-based estimators, especially TSLOPE, can uncover economically meaningful clustered dependence structures.

math.ST

Controllable protein design with particle-based Feynman-Kac steering

Proteins underpin most biological function, and the ability to design them with tailored structures and properties is central to advances in biotechnology. Diffusion-based generative models have emerged as powerful tools for protein design, but steering them toward proteins with specified properties remains challenging. The Feynman-Kac (FK) framework provides a principled way to guide diffusion models using user-defined rewards. In this paper, we enable FK-based steering of RFdiffusion through the development of guiding potentials that leverage ProteinMPNN and structural relaxation to guide the diffusion process towards desired properties. We show that steering can be used to consistently improve predicted interface energetics and increase binder designability by $89.5\%$. Together, these results establish that diffusion-based protein design can be effectively steered toward arbitrary, non-differentiable objectives, providing a model-independent framework for controllable protein generation.

cs.LG

A Unified and Computationally Efficient Non-Gaussian Statistical Modeling Framework

Datasets that exhibit non-Gaussian characteristics are common in many fields, while the current modeling framework and available software for non-Gaussian models is limited. We introduce Linear Latent Non-Gaussian Models (LLnGMs), a unified and computationally efficient statistical modeling framework that extends a class of latent Gaussian models to allow for latent non-Gaussian processes. The framework unifies several popular models, from simple temporal models to complex spatial-temporal and multivariate models, facilitating natural non-Gaussian extensions. Computationally efficient Bayesian inference, with theoretical guarantees, is developed based on stochastic gradient descent estimation. The R package \texttt{ngme2}, which implements the framework, is presented and demonstrated through a wide range of applications including novel non-Gaussian spatial and spatio-temporal models.

stat.ME

A flexible class of latent variable models for the analysis of antibody response data

Existing approaches to modelling antibody concentration data are mostly based on finite mixture models that rely on the assumption that individuals can be divided into two distinct groups: seronegative and seropositive. Here, we challenge this dichotomous modelling assumption and propose a latent variable modelling framework in which the immune status of each individual is represented along a continuum of latent seroreactivity, ranging from minimal to strong immune activation. This formulation provides greater flexibility in capturing age-related changes in antibody distributions while preserving the full information content of quantitative measurements. We show that the proposed class of models can accommodate a large variety of model formulations, both mechanistic and regression-based, and also includes finite mixture models as a special case. We also propose a computationally efficient $L_2$-based estimator as an alternative to maximum likelihood estimation, which substantially reduces computational cost, and we establish its consistency. Through a case study on malaria serology, we demonstrate how the flexibility of the novel framework enables joint analyses across all ages while accounting for changes in transmission patterns. We conclude by outlining extensions of the proposed modelling framework and its relevance to other omics applications.

stat.ME

Geometric ergodicity of Gibbs samplers for linear latent models with GIG variance mixtures

We study geometric ergodicity of the Gibbs sampler for linear latent non-Gaussian models (LLnGMs), a class of hierarchical models in which conditional Gaussian structure is preserved through generalized inverse Gaussian (GIG) variance-mixture augmentation. Two complementary routes to geometric ergodicity are developed for the marginal chain on the mixing variables. First, we show that the associated Markov operator is trace-class, and hence admits a spectral gap, over a large portion of the GIG parameter space. Second, for the remaining boundary and heavy-tail regimes, we establish geometric ergodicity via drift and minorization, subject to an explicit null-smallness condition that quantifies how the drift interacts with the null space of the observation operator. Together, these results cover the full GIG parameter space, including the normal-inverse Gaussian, generalized asymmetric Laplace, and Student-$t$ special cases. The geometric ergodicity of this chain underpins the consistency of Gibbs-based stochastic-gradient estimators for maximum likelihood estimation, and we provide conditions that make the required integrability checks transparent. Numerical experiments illustrate the theoretical findings, contrasting mixing efficiency across parameter regimes and probing the role of the null-smallness constant.

math.ST

Scalable Ultra-High-Dimensional Quantile Regression with Genomic Applications

Modern datasets arising from social media, genomics, and biomedical informatics are often heterogeneous and (ultra) high-dimensional, creating substantial challenges for conventional modeling techniques. Quantile regression (QR) not only offers a flexible way to capture heterogeneous effects across the conditional distribution of an outcome, but also naturally produces prediction intervals that help quantify uncertainty in future predictions. However, classical QR methods can face serious memory and computational constraints in large-scale settings. These limitations motivate the use of parallel computing to maintain tractability. While extensive work has examined sample-splitting strategies in settings where the number of observations $n$ greatly exceeds the number of features $p$, the equally important (ultra) high-dimensional regime ($p >> n$) has been comparatively underexplored. To address this gap, we introduce a feature-splitting proximal point algorithm, FS-QRPPA, for penalized QR in high-dimensional regime. Leveraging recent developments in variational analysis, we establish a Q-linear convergence rate for FS-QRPPA and demonstrate its superior scalability in large-scale genomic applications from the UK Biobank relative to existing methods. Moreover, FS-QRPPA yields more accurate coefficient estimates and better coverage for prediction intervals than current approaches. We provide a parallel implementation in the R package fsQRPPA, making penalized QR tractable on large-scale datasets.

stat.ME

Asymptotic Distribution of Low-Dimensional Patterns Induced by Non-Differentiable Regularizers under General Loss Functions

This article investigates the asymptotic distribution of penalized estimators with non-differentiable penalties designed to recover low-dimensional pattern structures. Patterns play a central role in estimation, as they reveal the underlying structure of the parameter -- which coefficients are zero, which are equal, and how they are clustered. The main technical challenge stems from the discontinuous nature of these patterns (such as the sign function in the case of the Lasso penalty), a difficulty not previously addressed in the literature and only recently analyzed for the standard linear model. To overcome this, we extend classical results from empirical process theory for M-estimation by incorporating the distributional behavior of model patterns. We introduce a new mathematical framework for studying pattern convergence of regularized M-estimators. While classical approaches to distributional convergence rely on uniform conditions, our analysis employs a new local condition, stochastic Lipschitz differentiability (SLD), which controls fluctuations of the Taylor remainder. We demonstrate how this framework applies to a broad class of loss functions, covering generalized linear models (e.g., logistic and Poisson regression) and robust regression settings with non-smooth losses such as the Huber and quantile loss.

math.ST

The Choice of Normalization Influences Shrinkage in Regularized Regression

Regularized models are often sensitive to the scales of the features in the data and it has therefore become standard practice to normalize (center and scale) the features before fitting the model. But there are many different ways to normalize the features and the choice may have dramatic effects on the resulting model. In spite of this, there has so far been no research on this topic. In this paper, we begin to bridge this knowledge gap by studying normalization in the context of lasso, ridge, and elastic net regression. We focus on binary features and show that their class balances (proportions of ones) directly influences the regression coefficients and that this effect depends on the combination of normalization and regularization methods used. We demonstrate that this effect can be mitigated by scaling binary features with their variance in the case of the lasso and standard deviation in the case of ridge regression, but that this comes at the cost of increased variance of the coefficient estimates. For the elastic net, we show that scaling the penalty weights, rather than the features, can achieve the same effect. Finally, we also tackle mixes of binary and normal features as well as interactions and provide some initial results on how to normalize features in these cases.

stat.ML

Incorporating Correlated Nugget Effects in Multivariate Spatial Models: An Application to Argo Ocean Data

Accurate analysis of global oceanographic data, such as temperature and salinity profiles from the Argo program, requires geostatistical models capable of capturing complex spatial dependencies. This study introduces Gaussian and non-Gaussian hierarchical multivariate Matérn-SPDE models with correlated nugget effects to account for small-scale variability and measurement error correlations. Using simulations and Argo data, we demonstrate that incorporating correlated nugget effects significantly improves the accuracy of parameter estimation and spatial prediction in both Gaussian and non-Gaussian multivariate spatial processes. When applied to global ocean temperature and salinity data, our model yields lower correlation estimates between fields compared to models that assume independent noise. This suggests that traditional models may overestimate the underlying field correlation. By separating these effects, our approach captures fine-scale oceanic patterns more effectively. These findings show the importance of relaxing the assumption of independent measurement errors in multivariate hierarchical models.

stat.ME

Nonparametric Shrinkage Estimation in High Dimensional Generalized Linear Models via Polya Trees

Regularization in fitting regression models has been a highly active topic of research in the past few decades, but most of the existing methods are designed for particular situations, e.g. for the case of a sparse coefficient vector. We consider the problem of designing $\textit{universally}$ optimal regularized estimators in a given generalized linear model with fixed effects. First, we propose as a contender the Bayes estimator against an $\textit{ideal}$ prior that assigns equal mass to every permutation of the fixed coefficient vector, thus depending on the true coefficients only through their empirical CDF. We prove some optimality properties of this oracle estimator in both the frequentist and Bayesian frameworks. To compete with the oracle estimator, we posit a hierarchical Bayes model where the individual coefficients are modeled as i.i.d. draws from a common distribution $π$, which is in turn assigned a Polya tree prior to reflect indefiniteness. We demonstrate in examples that the posterior mean of $π$ under the postulated model adapts nonparametrically to the empirical CDF of the true coefficients. Correspondingly, the posterior means of the coefficients themselves are used to mimic the ideal estimator. Numerical experiments show that our method has better estimation and prediction accuracy compared to various parametric and nonparametric alternatives, from relatively standard $L_p$-regularized estimators to modern penalized-likelihood and Bayesian estimators for high dimensional regression.

stat.ME