arXiv ScienceSearch

arXiv · 2411.15691

Semi-supervised inference using unlabeled summary statistics

Abstract

Semi-supervised inference assumes access to a labeled dataset together with a large unlabeled dataset in which the outcome variable is missing, and it is widely used to improve statistical efficiency and support generalizability across populations. In many modern applications, however, individual-level unlabeled data may not be directly accessible due to privacy restrictions, data-sharing limits, or storage constraints, while summary statistics such as sample means and covariances from the unlabeled population are often available. In this work, we study this constrained semi-supervised setting where, in addition to labeled data with individualized observations, auxiliary information from the unlabeled population is available only through summary statistics. We propose new semi-supervised inference methods for mean estimation under both covariate-independent and covariate-dependent labeling and show that unlabeled summaries can still improve efficiency and help correct selection bias. The proposed methods apply in high dimensions and are robust to model misspecification. Valid inference is obtained under sparsity conditions comparable to those required by semi-supervised methods that assume access to individual-level unlabeled samples. Our approach relies on a specialized cross-fitting procedure, where sample splitting is applied only to the labeled data, which removes the need for individualized unlabeled covariates. We further extend this framework to average treatment effect estimation, enabling generalizability and transportability of causal conclusions in this constrained semi-supervised setting.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Facheng Yu, Zhen Qi, Yuqian Zhang. 2026-06-02. Semi-supervised inference using unlabeled summary statistics. https://arxiv.org/abs/2411.15691

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Functional BART with Shape Priors: A Bayesian Tree Approach to Constrained Functional Regression

Motivated by the remarkable success of Bayesian additive regression trees (BART) in regression modelling, we propose a novel nonparametric Bayesian method, termed Functional BART (FBART), tailored specifically for function-on-scalar regression. FBART leverages spline-based representations for functional responses coupled with a flexible tree-based partitioning structure, effectively capturing complex and heterogeneous relationships between response curves and scalar predictors. To facilitate efficient posterior inference, we develop a customized Bayesian backfitting algorithm. Additionally, we extend FBART by introducing shape constraints (e.g., monotonicity or convexity) on the response curves, enabling enhanced estimation and prediction when prior shape information is available. The use of shape priors ensures that posterior samples respect the specified functional constraints. Under mild regularity conditions, we establish posterior convergence rates for both FBART and its shape-constrained variant, demonstrating rate adaptivity to unknown smoothness. Extensive simulation studies and analyses of two real datasets illustrate the superior estimation accuracy and predictive performance of our proposed methods compared to existing state-of-the-art alternatives.

stat.ME

An integer programming-based approach to construct exact two-sample binomial tests with maximum power

Comparing binomial proportions is a foundational task in clinical trials; however, practitioners often use a liberal likelihood-based test or the conservative Fisher's exact test. To increase power without sacrificing validity, we use a linear integer program to construct finite-sample exact binomial tests with maximum power. Our proposed Average Power Knapsack (APK) test finds a decision boundary that maximizes average power, guaranteeing it cannot uniformly be improved in terms of power, while enforcing type I error rate control across the null hypothesis parameter space using Lipschitz continuity. Our numerical evaluation reveals consistent pointwise power gains of up to 35% for the APK test over Fisher's exact test, as well as systematic power improvements over state-of-the-art exact Berger and Boos implementations of the mid-p-value and Z-pooled tests. When compared against non-exact procedures exhibiting moderate type I error rate inflation, the exact APK test maintains comparable power. For sample sizes where APK construction becomes computationally intensive, Fisher's mid-p-value test emerges as the best non-exact alternative. To support practical application, we developed an interactive R Shiny tool that enables seamless implementation of our optimized exact tests for binomial proportions.

stat.ME

On the limitations of causal inference with current-treatment Cox models

Cox models with time-varying treatments often include only the current treatment level. Translating such a model into causally meaningful intervention-specific survival probabilities relies on the Markov property: that the hazard is independent of treatment history conditional on current treatment. For the Markov property to not be population-specific, it needs to hold also conditional on any unmeasured prognostic heterogeneity (frailty). Using a discrete-time argument, earlier work concluded that the Markov property can hold both conditionally and marginally only in the absence of a treatment effect or an effect of the unmeasured heterogeneity. We broaden this argument by developing a continuous-time framework that encompasses both proposed extensions of the Kaplan-Meier curve and current-treatment marginal structural Cox models, and sharpen it by characterizing precisely the conditions under which the Markov property can hold both conditionally and marginally. Specifically, we show that it requires the absence of treatment-frailty interaction on the additive hazard scale. As frailty is inherently unmeasured, such a no-interaction assumption cannot be verified. Consequently, causal interpretation of a current-treatment marginal structural Cox model rests on a strong and unverifiable structural assumption. Illustrating this point, we construct a setting in which a Cox model is correctly specified conditionally on current treatment only, although the marginal Markov property fails. Transforming the fitted model to the survival scale does not recover the true intervention-specific survival probabilities. Thus, moving from the hazard scale to the survival scale is not in itself sufficient to obtain a causal interpretation in the time-varying treatment setting.

stat.ME