arXiv ScienceSearch

arXiv · 2604.22391

Conformalized Super Learner

Abstract

The Super Learner (SL) is a widely used ensemble method that combines point predictions from a library of learners based on their predictive performance. Interval predictions are of considerable practical interest because they allow uncertainty in predictions produced by an individual learner or an ensemble to be quantified. Several methods have been proposed for constructing interval predictions based on the SL, however, these approaches are typically justified using asymptotic arguments or rely on computationally intensive procedures such as the bootstrap. Conformal prediction (CP) is a machine learning framework for constructing prediction intervals with finite-sample and asymptotic coverage guarantees under mild conditions. We propose coupling CP with the SL through a natural construction that mirrors the original SL framework, using individual learner weights and combining learner-specific conformity scores via a weighted majority vote. We characterize the properties of the resulting SL-based prediction intervals for continuous outcomes. We cover settings under exchangeability, potential violations of exchangeability, and data-generating mechanisms exhibiting heteroscedasticity, sparsity, and other forms of distributional heterogeneity. A comprehensive simulation study shows that the conformalized SL achieves valid finite-sample coverage with competitive performance relative to the true data-generating mechanism. A central contribution of this work is an application to predicting creatinine levels using socio-demographic, biometric, and laboratory measurements. This example demonstrates the benefits of an ensemble with carefully selected learners designed to capture key aspects of complex regression functions, including non-linear effects, interactions, sparsity, heteroscedasticity, and robustness to outliers.

Explore related subjects

Keep this discovery

BibTeXRIS

Zhanli Wu, Fabrizio Leisen, Miguel-Angel Luque-Fernandez, F. Javier Rubio. 2026-09-08. Conformalized Super Learner. https://arxiv.org/abs/2604.22391

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

Optimal Slice-Adaptive Tuning of Hybrid Slice Sampling

Slice sampling is a Markov chain Monte Carlo algorithm that draws its next state uniformly from a "slice"---a super-level set of the target density function---at each iteration, thereby providing automatic local adaptivity to the scale of the target. In practice the exact slice is not known, so general-purpose implementations use an approximate slice that is grown from a starting interval of length $w>0$, with a computational cost that depends on $w$. This work presents an analysis of the average per-iteration number of target density evaluations, as a function of $w$, of hybrid slice sampling with various slice-finding schemes for targets with contiguous slices. The paper uses the results of the analysis to develop automated, slice-adaptive tuning schemes along with suboptimality bounds and asymptotic convergence guarantees. Simulations demonstrate that the tuning schemes reliably yield near-optimal slice-adaptive tuning with essentially no dependence on the initial setting of $w$.

stat.CO

Recovering Weak Signals with Normalizing Flows

In many scientific disciplines, weak signals of interest are obscured by dominant nuisance signals that are several orders of magnitude stronger. Recovering these weak signals requires subtracting the dominant ones; however, this calibration process inherently distorts or partially suppresses the underlying signal of interest. To address this problem, we propose the use of normalizing flow models to reconstruct calibration-affected weak signals. By leveraging the statistical invariance of the target signals and assuming minimal initial suppression, our framework effectively recovers the lost signal components. We provide a comprehensive theoretical overview of this normalizing flow-based recovery method and demonstrate its efficacy using simulated data.

stat.ML

Residual-augmented flow matching operators for probabilistic partial differential equations

Learning surrogate models for physical systems with latent uncertainty remains challenging in data-scarce regimes: deterministic neural operators fail to characterize uncertainty, while generative approaches require large ensembles of high-fidelity solution operator simulations and often sacrifice resolution generalizability. In this work, we propose a residual-augmented probabilistic operator learning framework that casts flow-matching-based generative modeling in infinite-dimensional function spaces while leveraging inexpensive low-fidelity solution operators as an inductive bias. Rather than learning the full high-fidelity stochastic solution operator directly, the proposed framework learns probabilistic residual operators that characterize the discrepancy between low- and high-fidelity solutions. By parameterizing the vector field in flow matching using neural operators conditioned on both the known system input and low-fidelity solution, the framework amortizes probabilistic inference across input conditions while enabling uncertainty-aware and resolution-generalizable predictions across spatial discretizations. Numerical experiments on stochastic advection, Burgers', and Darcy flow systems demonstrate that the residual-augmented formulation improves predictive accuracy under the same high-fidelity data budget, while the probabilistic operator learning formulation enables accurate characterization of uncertainty in low-data regimes compared to learning high-fidelity stochastic operators directly from data.

stat.CO