arXiv ScienceSearch

arXiv · 2512.15802

Consensus dimension reduction via multi-view learning

Abstract

A plethora of dimension reduction methods have been developed to visualize high-dimensional data in low dimensions. However, different dimension reduction methods often output different and possibly conflicting visualizations of the same data. This problem is further exacerbated by the choice of hyperparameters, which may substantially impact the resulting visualization. To obtain a more robust and trustworthy dimension reduction output, we advocate for a consensus approach, which summarizes multiple visualizations into a single consensus dimension reduction visualization. Here, we leverage ideas from multi-view learning in order to identify the patterns that are most stable or shared across the many different dimension reduction visualizations, or views, and subsequently visualize this shared structure in a single low-dimensional plot. We demonstrate that this consensus visualization effectively identifies and preserves the shared low-dimensional data structure through both simulated and real-world case studies. We further highlight our method's robustness to the choice of dimension reduction method and hyperparameters -- a highly-desirable property when working towards trustworthy and reproducible data science.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Bingxue An, Tiffany M. Tang. 2025-12-16. Consensus dimension reduction via multi-view learning. https://arxiv.org/abs/2512.15802

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Partial identification and unmeasured confounding with multiple treatments and multiple outcomes

Estimating the health effects of multiple air pollutants is a crucial problem in public health, but one that is difficult due to unmeasured confounding bias. Motivated by this issue, we develop a framework for partial identification of causal effects in the presence of unmeasured confounding in settings with multiple treatments and multiple outcomes. Under a factor confounding assumption, we show that joint partial identification regions for multiple estimands can be more informative than considering partial identification for individual estimands one at a time. We show how assumptions related to the strength of confounding or magnitude of plausible effect sizes for one estimand can reduce the partial identification regions for other estimands. As a special case of this result, we explore how negative control assumptions reduce partial identification regions and discuss conditions under which point identification can be obtained. We develop novel computational approaches to finding partial identification regions under a variety of these assumptions. We then estimate the causal effect of PM$_{2.5}$ components on a variety of public health outcomes in the United States Medicare cohort, where we find that the detrimental effect of certain air pollutants are robust to the potential presence of unmeasured confounding bias.

stat.ME

Dynamic programming principle in cost-efficient sequential design: optimal update scheduling under cost constraints

We study sequential cost-efficient design in a situation where each update of covariates involves a fixed time cost typically considerable compared to a single measurement time. The problem arises from parameter estimation in switching measurements on superconducting Josephson junctions which are components needed in quantum computers and other superconducting electronics. In switching measurements, a sequence of current pulses is applied to the junction and a binary voltage response is observed. The measurement requires a very low temperature that can be kept stable only for a relatively short time, and therefore it is essential to use an efficient design. We use the dynamic programming principle from the mathematical theory of optimal control to solve the optimal update times. Specifically, we give a formulation for $D$-optimal experimental design based on the dynamic programming principle with a binary response model. Our simulations demonstrate the cost-efficiency compared to the previously used methods.

stat.ME

Consistent Bayesian Spatial Domain Partitioning Using Predictive Spanning Tree Methods

Bayesian model-based spatial clustering methods are widely used for their flexibility in estimating latent clusters with an unknown number of clusters while accounting for spatial proximity. Many existing methods are designed for clustering finite spatial units, limiting their ability to make predictions, or may impose restrictive geometric constraints on the shapes of subregions. Furthermore, the posterior clustering consistency theory of spatial clustering models remains largely unexplored in the literature. In this study, we propose a Spatial Domain Random Partition Model (Spat-RPM) and demonstrate its application for spatially clustered regression, which extends spanning tree-based Bayesian spatial clustering by partitioning the spatial domain into disjoint blocks and using spanning tree cuts to induce contiguous domain partitions. Under an infill-domain asymptotic framework, we introduce a new distance metric to study the posterior concentration of domain partitions. We show that Spat-RPM achieves a consistent estimation of domain partitions, including the number of clusters (which may go to infinity), and derive posterior concentration rates for partition, parameter, and prediction. We also establish conditions on the hyperparameters to achieve consistency, offering important practical guidance for hyperparameter selection. Finally, we examine the asymptotic properties of our model through simulation studies and apply it to Atlantic Ocean data.

stat.ME