arXiv Science⌕ Search

arXiv · 2610.02578

High-Dimensional Asymptotics and Dataset Selection for Private Transfer Learning

Abstract

To commit to buying external data or participate in collaborative learning, one must decide whether the additional data will improve prediction enough to justify the cost. This comes with several challenges: (i) the decision often relies only on aggregated statistics available publicly, rather than individual-level data; (ii) covariate and model shifts can induce negative transfer, so the additional data deteriorates rather than improves performance; (iii) if the data is sensitive, its privatization requires the injection of noise, which can also offset the benefit of a larger sample size. In this paper, we model the problem of dataset selection through high-dimensional regression with multiple heterogeneous sources and a weighted ridge estimator. Our approach uses only summary statistics and it gives privacy guarantees either on labels only or jointly on features and labels, in terms of $ρ$-zero-concentrated differential privacy. The main technical contribution is a deterministic equivalent of the test error, which captures the interactions between sample size, covariance structure, model shift, regularization and privacy noise. Our theory allows to optimize hyperparameters (weights and ridge regularizers) and, more broadly, to decide when private external datasets are useful without accessing the data itself but only relying on population-level quantities. This provides a theoretically tractable foundation for private transfer learning, which we support via experiments on both synthetic and real-world datasets.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Filip Kovačević, Edwige Cyffers, Stefano Sarao Mannelli, Marco Mondelli. 2026-10-01. High-Dimensional Asymptotics and Dataset Selection for Private Transfer Learning. https://arxiv.org/abs/2610.02578

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Convergence of Statistical Estimators via Mutual Information Bounds

Recent advances in statistical learning theory have revealed profound connections between mutual information (MI) bounds, PAC-Bayesian theory, and Bayesian nonparametrics. This work introduces a mutual information bound for statistical models, and derives from it convergence rates for fractional posteriors, for their variational approximations, and for the maximum likelihood estimator. The observations are assumed independent but not identically distributed, and the model is not assumed well-specified, so that the bounds are oracle inequalities and cover regression with a fixed or a conditioned design; the independent and identically distributed, well-specified case is recovered by dropping an index. We illustrate the method on two applications. In the Gaussian sequence model the rate is minimax in both the radius of the Sobolev ball and the sample size, which only appears through its product with the temperature. In logistic regression, where the model is not conjugate and the variational approximation is computed by a stochastic gradient method, the bound applies to the output of the algorithm rather than to an idealized minimizer, and attains the parametric order in the Renyi risk with no logarithmic factor, the statistical and the optimization error being separated.

stat.ML↗

Mitigating Over-squashing without Rewiring: A Sheaf Effective Resistance Perspective

Graph Neural Networks (GNNs) often struggle to capture long-range dependencies due to over-squashing -- a phenomenon in which the repeated compression of node embeddings into finite-size messages causes representations to collapse. Over-squashing is most often diagnosed as a property of the graph topology, with effective resistance serving as a principled measure of the bottleneck. We provide a complementary view on the matter: building on cellular sheaves, we introduce sheaf effective resistance, a generalization of effective resistance that depends on the sheaf attached to the graph, and we prove that for flat vector bundles, the over-squashing sensitivity in the Jacobian sense is upper bounded by a quantity related to the sheaf effective resistance between the nodes. The bottleneck thus need not lie in the graph itself: it can be relocated, and reduced, by adjusting the sheaf. We instantiate this idea in FlatNSD, a simple message-passing variant of Neural Sheaf Diffusion, and show that it implicitly learns to modulate total sheaf effective resistance, performing well on benchmarks designed to stress over-squashing without altering the original graph topology.

stat.ML↗

Risk-Calibrated Proposal Transport for Finite-Particle Diffusion Steering

Inference-time steering combines pretrained diffusion experts or rewards without retraining by changing the dynamics that transport noise to data. Feynman-Kac correction compensates for proposal mismatch through importance-weighted sequential Monte Carlo (SMC), whose finite-particle behavior depends on the proposal. Variance-controlling guidance (VCG) improves that proposal by fitting a linear drift correction to minimize empirical log-weight-rate variance. Although its population optimum cannot worsen residual variance, finite-particle VCG can nearly eliminate its fitting residual while increasing residual risk on new states by orders of magnitude. The resulting update can degrade unweighted generation or accelerate particle collapse. We show that the centered Feynman-Kac rate is the normalized transport residual and that expected out-of-fit benefit is exactly population headroom minus coefficient-estimation penalty. Under regularity assumptions, a Wasserstein analysis bounds the unweighted proposal's terminal error using this residual. These results motivate Risk-Calibrated Proposal Transport (RCPT), which uses deletion leave-one-out residuals to calibrate the retained fraction of the VCG update, adding no model calls and only small linear-algebra overhead. Experiments on 2D checker distributions, scaffold decoration, molecular property optimization, and class-conditional CIFAR-10 generation demonstrate recovery from harmful fitted updates. Across molecular and image domains, RCPT mitigates harmful fitted updates and improves a broad range of terminal metrics relative to uncalibrated VCG.

stat.ML↗