arXiv ScienceSearch

arXiv · 2502.15584

Improving variable selection properties with data integration and transfer learning

Abstract

We study variable selection (also called support recovery) in high-dimensional sparse linear regression when one has external information on which variables are likely to be associated with the response. Consistent recovery is only possible under somewhat restrictive conditions on sample size, dimension, signal strength, and sparsity. We investigate how these conditions can be relaxed by incorporating said external information. A key application that we consider is structural transfer learning, where variables selected in one or more source datasets are used to guide variable selection in a target dataset. We introduce a family of likelihood penalties that depend on the external information, motivated by connections to Bayesian variable selection. We show that these methods achieve variable selection consistency in regimes where any method ignoring external information fails, and that they achieve consistency at faster rates. We first quantify the potential gains under ideal, oracle-chosen, penalties. We then propose computationally efficient empirical Bayes procedures that learn suitable penalties from the data. We prove that these procedures have improved variable selection properties compared to methods that do not use external information. We illustrate our approach using simulations and a genomics application, where results from mouse experiments are used to inform variable selection for gene expression data in humans.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Paul Rognon-Vael, David Rossell, Piotr Zwiernik. 2026-02-13. Improving variable selection properties with data integration and transfer learning. https://arxiv.org/abs/2502.15584

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

On the minimax-rate optimality of approximate Bayesian computation in nonparametric problems

Approximate Bayesian computation (ABC) replaces likelihood evaluation by simulation and comparison of observed and synthetic data. We establish minimax-rate guarantees for nonparametric ABC under random-series priors with simulable finite-dimensional coordinates. The contraction theorem uses local prior mass, bounds on ABC acceptance probabilities, and control of prior mass outside a sieve. In fixed-design orthogonal-series regression with centered $g$-and-$k$ errors, an infinite Gaussian series prior with a compact scale hyperprior yields minimax-rate contraction and a minimax-rate clipped posterior mean. In compound Poisson decompounding, only random sums are observed and the target is the underlying jump density. With an unknown count intensity in a fixed compact subinterval of $(0,π/2)$, we prove stability of the zero-count-augmented trigonometric population summaries and use a square-root Gaussian series prior on the space of probability density functions. Over bounded periodic Sobolev classes of smoothness $α>d/2$, a polynomially enlarged synthetic sample yields ABC contraction at rate $n^{-α/(2α+d)}$ and posterior mean squared risk of order $n^{-2α/(2α+d)}$, matching a lower bound for the aggregate-observation model. Rejection-ABC Monte Carlo approximations inherit these rates under sufficient sampling budgets.

math.ST

On the continuity of the Tukey depth function for fuzzy data

The practical use of statistical depth for fuzzy data requires regularity, inferential stability, and computational feasibility. This paper studies these aspects for the first time for the fuzzy depth, in particular for the fuzzy Tukey depth. In particular, we investigate the continuity of this depth with respect to each of its arguments. As a function of the elements of the underlying space, we prove its upper semicontinuity under the main metrics used in fuzzy spaces. As a function of the distribution with respect to which the depth is computed, we establish the almost sure uniform consistency of its empirical version for continuous fuzzy random variables. These results ensure closed depth regions and support the use of sample depth values to approximate their population counterparts. We also derive consistency of empirical deepest-point estimators and study a finite-grid approximation of the depth, supported by theoretical and simulation results.

math.ST

Bernstein-smoothed estimation and bootstrap inference for the lower-tail Spearman's rho curve

This paper studies a Bernstein-smoothed plug-in estimator for the lower-tail Spearman's rho curve, a rank-based measure of local concordance defined through a normalized copula integral over the lower-left square $[0,p]^2$. The estimator applies the lower-tail Spearman functional to the empirical Bernstein copula and introduces a degree parameter that controls finite-sample regularization. The contribution is target-specific estimation and inference for the lower-tail Spearman's rho curve rather than a new general-purpose copula estimator. We establish uniform strong consistency on compact intervals away from zero whenever $m \to \infty$. Under a target-specific integrated Bernstein-bias condition and $\sqrt n/m \to 0$, we derive functional weak convergence with the same first-order Gaussian limit as the empirical copula-based estimator; this degree regime includes $m = \lfloor n^{2/3}\rfloor$. For inference, we justify a smoothed beta bootstrap based on sampling from the empirical beta copula and construct pointwise confidence intervals from the absolute centered bootstrap root. Monte Carlo experiments for the Farlie--Gumbel--Morgenstern, Gaussian, Clayton, and Frank copulas assess pointwise coverage over the threshold grid and show that Bernstein smoothing generally reduces integrated variance and often lowers mean integrated squared error under weak to moderate dependence. Additional common-sample experiments compare the proposed estimator with the empirical beta and empirical checkerboard Bernstein copula estimators in terms of pointwise error, integrated error, computation time, selected degrees, numerical stability, and pointwise coverage. A sensitivity analysis shows that $m = \lfloor n^{2/3}\rfloor$ is a simple theoretically admissible default that avoids the most severe oversmoothing. A descriptive application to the Loss--ALAE insurance claims data illustrates our method.

math.ST