arXiv ScienceSearch

arXiv · 2205.13734

An efficient tensor regression for high-dimensional data

Abstract

Most currently used tensor regression models for high-dimensional data are based on Tucker decomposition, which has good properties but loses its efficiency in compressing tensors very quickly as the order of tensors increases, say greater than four or five. However, for the simplest tensor autoregression in handling time series data, its coefficient tensor already has the order of six. This paper revises a newly proposed tensor train (TT) decomposition and then applies it to tensor regression such that a nice statistical interpretation can be obtained. The new tensor regression can well match the data with hierarchical structures, and it even can lead to a better interpretation for the data with factorial structures, which are supposed to be better fitted by models with Tucker decomposition. More importantly, the new tensor regression can be easily applied to the case with higher order tensors since TT decomposition can compress the coefficient tensors much more efficiently. The methodology is also extended to tensor autoregression for time series data, and nonasymptotic properties are derived for the ordinary least squares estimations of both tensor regression and autoregression. A new algorithm is introduced to search for estimators, and its theoretical justification is also discussed. Theoretical and computational properties of the proposed methodology are verified by simulation studies, and the advantages over existing methods are illustrated by two real examples.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yuefeng Si, Yingying Zhang, Yuxi Cai, Chunling Liu, Guodong Li. 2024-03-19. An efficient tensor regression for high-dimensional data. https://arxiv.org/abs/2205.13734

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A robust, scalable K-statistic for quantifying immune cell clustering in spatial proteomics data

Spatial summary statistics based on point process theory are widely used to quantify the spatial organization of cell populations in single-cell spatial proteomics data. Among these, Ripley's K is a popular metric for assessing whether cells are spatially clustered or are randomly dispersed. However, the key assumption of spatial homogeneity is frequently violated in spatial proteomics data, leading to overestimates of cell clustering and colocalization. To address this, we propose a novel method, termed KAMP (K adjustment by Analytical Moments of the Permutation distribution), for quantifying the spatial organization of cells in spatial proteomics samples. KAMP leverages background cells in each sample along with a new closed-form representation of the first and second moments of the permutation null distribution of Ripley's K. Our method is robust to inhomogeneity, computationally efficient even in large datasets, and provides approximate p-values to test spatial clustering and colocalization. Methodological developments are motivated by a spatial proteomics study of women with ovarian cancer; in the subset with sufficient B cells and macrophages, KAMP provides exploratory, scale-specific evidence linking B cell-macrophage colocalization with overall patient survival. Notably, we also find evidence that using K without correcting for sample inhomogeneity may bias hazard ratio estimates in downstream analyses.

stat.ME

A Heavily Right Strategy for Statistical Inference with Dependent Studies in Arbitrary Dimensions

We leverage recent advances in heavy-tail approximations for global hypothesis testing with dependent studies to construct approximate confidence regions without modeling or estimating their dependence structures. A non-rejection region is a confidence region but it may not be convex. Convexity is appealing because it ensures any one-dimensional linear projection of the region is a confidence interval, easy to compute and interpret. We show why convexity fails for nearly all heavy-tail combination tests proposed in recent years, including the influential Cauchy combination test. These insights motivate a heavily right strategy: truncating the left half of the Cauchy distribution to obtain the Half-Cauchy combination test. The harmonic mean test also corresponds to a heavily right distribution with a Cauchy-like tail, namely a Pareto distribution with unit power. We prove that both approaches guarantee convexity when individual studies are summarized by Hotelling $T^2$ or $χ^{2}$ statistics (regardless of the validity of this summary) and provide efficient, exact algorithms for implementation. Applying these methods, we develop a divide-and-combine strategy for mean estimation in any dimension and construct simultaneous confidence intervals in a network meta-analysis for treatment effect comparisons across multiple clinical trials. We also present many open problems and conclude with epistemic reflections.

stat.ME

Recent advances in the Bradley--Terry Model: theory, algorithms, and applications

This article surveys recent advances in the Bradley-Terry (BT) model and its extensions. We focus on statistical and computational aspects, with particular emphasis on the regime in which both the number of objects and the volume of comparisons tend to infinity-a setting relevant to large-scale applications. The main topics include asymptotic theory for statistical estimation and inference, together with the corresponding computational algorithms. We also review applications of these models in sports analytics, social choice, psychometrics, and machine learning. Finally, we discuss several key challenges and outline directions for future research.

stat.ME