arXiv Science⌕ Search

arXiv · 2609.38141

Global Synchronization for Multi-Source Data Integration under Blockwise Missing Patterns

Abstract

Multi-source data integration problems over datasets from different sources covering different but possibly overlapping sets of entities have become increasingly important in many real-world areas, including genomics, single-cell analysis, and healthcare research. In such problems, one often first learns a low-dimensional representation of the entities within each source and then integrates these representations across sources. As the representations from different sources are only identifiable up to some transformation, how to align them across sources using the sources' overlapping entities becomes a key challenge. Existing methods align the sources in a sequential or tree-structured manner, and are therefore sensitive to the chosen order and exploit only part of the available overlapping information. Motivated by this limitation, we propose Global Synchronized Multiple Matrix Integration (GSMMI), which formulates this alignment problem as a global synchronization problem and jointly aligns all sources using all pairwise overlaps at once, thereby making full use of all overlapping information across the sources. We develop an efficient iterative algorithm for GSMMI that is fast and scalable to the large-scale data arising in these applications. We show both theoretically and empirically that GSMMI improves alignment accuracy, with clear improvements even under modest overlap structure. Moreover, we develop GSMMI to be broadly applicable across data types, covering symmetric positive semidefinite, symmetric indefinite, and asymmetric or rectangular matrices, and even settings where sources overlap only in their rows or only in their columns, making it suitable for a wide variety of application scenarios.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Runbing Zheng, Dmitriy Kunisky. 2026-09-29. Global Synchronization for Multi-Source Data Integration under Blockwise Missing Patterns. https://arxiv.org/abs/2609.38141

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Bayesian Neural-Net-Assisted Multi-Treatment Mixture Cure Survival Model with Application in Pediatric Oncology

Estimating covariate-conditional treatment effects in multi-arm oncology studies is complicated when treatment arms have common distributional features and a non-negligible fraction of patients achieve long-term remission. We propose a joint mixture cure model with covariate-dependent mixtures of log-normal kernels with treatment-specific inclusions. Both linear and neural-network-assisted non-linear covariate links are proposed. Specifically, the susceptible survival distributions use a common finite dictionary of log-normal components, and then a binary inclusion matrix determines which components are active in each treatment arm. All parameters, including the hidden bases of the neural network, are learned jointly, while the output coefficients remain treatment- or component-specific. Posterior inference is performed using gradient-based MCMC, and treatment effects are summarized by covariate-conditional differences in restricted mean survival time (RMST). Variable importance is assessed using thresholded marginal best linear projections with data partitioning. Across two simulation settings, the proposed method demonstrates good finite-sample performance, with lower RMST-based estimation error than flexsurvcure. Compared with pairwise grf fits, the proposed method yields lower RMST-contrast MSE in most comparisons while ensuring mutually coherent multi-treatment contrasts. Finally, the application to the AALL0434 trial reveals covariate-dependent patterns in RMST posterior across methotrexate-based regimens and provides new insights into how these differences vary with patient covariates, highlighting the method's practical utility for studying heterogeneous treatment effects in pediatric oncology trials.

stat.ME↗

A Three-Stage PCA Procedure for Sequentially Arriving High-Dimensional Data

We develop a three-stage adaptive procedure for principal component analysis (PCA) when high-dimensional observations are collected sequentially and additional sampling incurs a cost. The procedure balances PCA compression loss against sampling cost while selecting the retained dimension through a prescribed explained-variance criterion. Starting from a pilot sample, an intermediate stage updates the PCA quantities before determining the final sample size, thereby avoiding reliance on unknown population eigenvalues. Under suitable regularity conditions, we establish both first- and second-order efficiency relative to the population oracle. Comparison with the corresponding two-stage rule shows that the additional recalibration yields sharper second-order control and reduces the influence of the pilot stage on the final sampling decision. The theory allows the ambient dimension to exceed the sample size under appropriate covariance and spectral conditions. Simulation studies demonstrate the strong finite-sample performance of the procedure across increasing dimensions and several dense covariance structures. As a real-data application, we conduct a retrospective study of gene-expression data from 32 cancer-type cohorts in The Cancer Genome Atlas, illustrating both cost-effective early stopping and settings in which additional observations are recommended.

stat.ME↗

Batting Average as the Product of Two Rates: Skill, Luck, and the Disappearance of the .400 Hitter

Batting average factors exactly as BA = c times f, where c = (AB - SO)/AB is the rate of avoiding a strikeout and f = H/(AB - SO) is the rate at which non-strikeout at-bats become hits. Using the 256 major-league hitters with at least 300 at-bats in 2025, we show that the two factors behave very differently. Strikeout avoidance is highly repeatable, with a median year-to-year correlation of 0.86 over 19 consecutive-season pairs from 2004 to 2025. The finishing rate is not: its median correlation is 0.44, and batting average itself (0.44) is no more repeatable than its noisier factor. A bivariate logistic-normal random-effects model fit to the 2025 season estimates the correlation between the two talents at -0.68 (95\% profile interval -0.85 to -0.50), far stronger than the raw correlation of -0.44, and implies single-season reliabilities of 0.91 for c and 0.37 for f. The model also yields a closed-form bivariate shrinkage estimator in which a hitter's strikeout rate informs the estimate of his finishing rate. Applied to decades of American and National League data, the decomposition revisits Gould's explanation for the disappearance of the .400 hitter. Since the dead-ball era, the talent variance of strikeout avoidance has grown roughly 4.6-fold while that of finishing has halved, and the correlation between them has moved from near zero to about -0.53. That emerging trade-off, rather than a general narrowing of talent, accounts for the reduced spread of modern batting averages.

stat.ME↗