arXiv Science⌕ Search

arXiv · 2609.38681

A Three-Stage PCA Procedure for Sequentially Arriving High-Dimensional Data

Abstract

We develop a three-stage adaptive procedure for principal component analysis (PCA) when high-dimensional observations are collected sequentially and additional sampling incurs a cost. The procedure balances PCA compression loss against sampling cost while selecting the retained dimension through a prescribed explained-variance criterion. Starting from a pilot sample, an intermediate stage updates the PCA quantities before determining the final sample size, thereby avoiding reliance on unknown population eigenvalues. Under suitable regularity conditions, we establish both first- and second-order efficiency relative to the population oracle. Comparison with the corresponding two-stage rule shows that the additional recalibration yields sharper second-order control and reduces the influence of the pilot stage on the final sampling decision. The theory allows the ambient dimension to exceed the sample size under appropriate covariance and spectral conditions. Simulation studies demonstrate the strong finite-sample performance of the procedure across increasing dimensions and several dense covariance structures. As a real-data application, we conduct a retrospective study of gene-expression data from 32 cancer-type cohorts in The Cancer Genome Atlas, illustrating both cost-effective early stopping and settings in which additional observations are recommended.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Partha Sarkar, Sairam Rayapolu, Bhargab Chattopadhyay. 2026-09-30. A Three-Stage PCA Procedure for Sequentially Arriving High-Dimensional Data. https://arxiv.org/abs/2609.38681

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A General Design-Based Framework and Estimator for Randomized Experiments

We describe a widely applicable design-based framework for drawing causal inference in randomized experiments. Causal effects are defined as linear functionals evaluated at unit-level potential outcome functions. Assumptions about the potential outcome functions are encoded as function spaces. This makes the framework expressive, allowing experimenters to formulate and investigate a wide range of causal questions that previously could not be investigated with design-based methods. The framework is particularly suited for complex, non-discrete interventions and causal interference. We describe a class of estimators for estimands defined using the framework and investigate their properties. We provide necessary and sufficient conditions for unbiasedness and consistency. We also describe a class of conservative variance estimators, which facilitate the construction of confidence intervals. In order to demonstrate the value of our approach in practice, we provide several illustrative examples of causal investigations which can be handled within our framework, but that could not be addressed using conventional design-based methods.

stat.ME↗

Bayesian Neural-Net-Assisted Multi-Treatment Mixture Cure Survival Model with Application in Pediatric Oncology

Estimating covariate-conditional treatment effects in multi-arm oncology studies is complicated when treatment arms have common distributional features and a non-negligible fraction of patients achieve long-term remission. We propose a joint mixture cure model with covariate-dependent mixtures of log-normal kernels with treatment-specific inclusions. Both linear and neural-network-assisted non-linear covariate links are proposed. Specifically, the susceptible survival distributions use a common finite dictionary of log-normal components, and then a binary inclusion matrix determines which components are active in each treatment arm. All parameters, including the hidden bases of the neural network, are learned jointly, while the output coefficients remain treatment- or component-specific. Posterior inference is performed using gradient-based MCMC, and treatment effects are summarized by covariate-conditional differences in restricted mean survival time (RMST). Variable importance is assessed using thresholded marginal best linear projections with data partitioning. Across two simulation settings, the proposed method demonstrates good finite-sample performance, with lower RMST-based estimation error than flexsurvcure. Compared with pairwise grf fits, the proposed method yields lower RMST-contrast MSE in most comparisons while ensuring mutually coherent multi-treatment contrasts. Finally, the application to the AALL0434 trial reveals covariate-dependent patterns in RMST posterior across methotrexate-based regimens and provides new insights into how these differences vary with patient covariates, highlighting the method's practical utility for studying heterogeneous treatment effects in pediatric oncology trials.

stat.ME↗

Slice Monte Carlo Integration

Numerical integration involving expensive target functions is a common bottleneck in Bayesian inference and simulation. When a cheap surrogate is available, standard approaches such as reweighting or importance sampling often suffer from high variance and inefficient use of function evaluations. We introduce Slice Monte Carlo integration (S$\ell$MC), a method that leverages a Nested Sampling-like procedure on the surrogate to partition the space into informative strata, or slices, while generating samples in the parameter space drawn from the prior within each slice. This enables stratified Monte Carlo integration of the expensive target function over the surrogate-induced partition, yielding an efficient estimate of the target integral. The surrogate level sets therefore make the induced ordering of the parameter space the relevant information for S$\ell$MC, rather than the pointwise target-to-surrogate ratios governing importance-sampling efficiency. Another key advantage of S$\ell$MC is the decoupling of slice volume estimation from the evaluation of the target, allowing the refinement of the slice-volume estimates without additional target evaluations. We investigate the properties of S$\ell$MC, demonstrate how to efficiently generate posterior samples, and assess its robustness in controlled Gaussian benchmark families with continuously tunable surrogate-target mismatch.

stat.ME↗