arXiv ScienceSearch

arXiv · 2110.06122

Nonnegative spatial factorization

Abstract

Gaussian processes are widely used for the analysis of spatial data due to their nonparametric flexibility and ability to quantify uncertainty, and recently developed scalable approximations have facilitated application to massive datasets. For multivariate outcomes, linear models of coregionalization combine dimension reduction with spatial correlation. However, their real-valued latent factors and loadings are difficult to interpret because, unlike nonnegative models, they do not recover a parts-based representation. We present nonnegative spatial factorization (NSF), a spatially-aware probabilistic dimension reduction model that naturally encourages sparsity. We compare NSF to real-valued spatial factorizations such as MEFISTO and nonspatial dimension reduction methods using simulations and high-dimensional spatial transcriptomics data. NSF identifies generalizable spatial patterns of gene expression. Since not all patterns of gene expression are spatial, we also propose a hybrid extension of NSF that combines spatial and nonspatial components, enabling quantification of spatial importance for both observations and features. A TensorFlow implementation of NSF is available from https://github.com/willtownes/nsf-paper .

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

F. William Townes, Barbara E. Engelhardt. 2021-10-12. Nonnegative spatial factorization. https://arxiv.org/abs/2110.06122

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Partial identification and unmeasured confounding with multiple treatments and multiple outcomes

Estimating the health effects of multiple air pollutants is a crucial problem in public health, but one that is difficult due to unmeasured confounding bias. Motivated by this issue, we develop a framework for partial identification of causal effects in the presence of unmeasured confounding in settings with multiple treatments and multiple outcomes. Under a factor confounding assumption, we show that joint partial identification regions for multiple estimands can be more informative than considering partial identification for individual estimands one at a time. We show how assumptions related to the strength of confounding or magnitude of plausible effect sizes for one estimand can reduce the partial identification regions for other estimands. As a special case of this result, we explore how negative control assumptions reduce partial identification regions and discuss conditions under which point identification can be obtained. We develop novel computational approaches to finding partial identification regions under a variety of these assumptions. We then estimate the causal effect of PM$_{2.5}$ components on a variety of public health outcomes in the United States Medicare cohort, where we find that the detrimental effect of certain air pollutants are robust to the potential presence of unmeasured confounding bias.

stat.ME

Dynamic programming principle in cost-efficient sequential design: optimal update scheduling under cost constraints

We study sequential cost-efficient design in a situation where each update of covariates involves a fixed time cost typically considerable compared to a single measurement time. The problem arises from parameter estimation in switching measurements on superconducting Josephson junctions which are components needed in quantum computers and other superconducting electronics. In switching measurements, a sequence of current pulses is applied to the junction and a binary voltage response is observed. The measurement requires a very low temperature that can be kept stable only for a relatively short time, and therefore it is essential to use an efficient design. We use the dynamic programming principle from the mathematical theory of optimal control to solve the optimal update times. Specifically, we give a formulation for $D$-optimal experimental design based on the dynamic programming principle with a binary response model. Our simulations demonstrate the cost-efficiency compared to the previously used methods.

stat.ME

Consistent Bayesian Spatial Domain Partitioning Using Predictive Spanning Tree Methods

Bayesian model-based spatial clustering methods are widely used for their flexibility in estimating latent clusters with an unknown number of clusters while accounting for spatial proximity. Many existing methods are designed for clustering finite spatial units, limiting their ability to make predictions, or may impose restrictive geometric constraints on the shapes of subregions. Furthermore, the posterior clustering consistency theory of spatial clustering models remains largely unexplored in the literature. In this study, we propose a Spatial Domain Random Partition Model (Spat-RPM) and demonstrate its application for spatially clustered regression, which extends spanning tree-based Bayesian spatial clustering by partitioning the spatial domain into disjoint blocks and using spanning tree cuts to induce contiguous domain partitions. Under an infill-domain asymptotic framework, we introduce a new distance metric to study the posterior concentration of domain partitions. We show that Spat-RPM achieves a consistent estimation of domain partitions, including the number of clusters (which may go to infinity), and derive posterior concentration rates for partition, parameter, and prediction. We also establish conditions on the hyperparameters to achieve consistency, offering important practical guidance for hyperparameter selection. Finally, we examine the asymptotic properties of our model through simulation studies and apply it to Atlantic Ocean data.

stat.ME