arXiv Science⌕ Search

arXiv · 2610.00769

Integration of external predictions for efficient estimation of the risk ratio

Abstract

We consider the augmentation of randomized experiments or trials with data from observational studies for the purpose of improving statistical precision. In particular, we focus on the Hybrid Augmented Inverse Probability Weighting estimator, designed to integrate predictions from several foundation models while preserving valid statistical inference. In this article, we extend the Augmented Inverse Probability Weighting framework to the estimation of the risk ratio, the quotient of the absolute risk of an exposed to an unexposed group. Our approach allows one to use information from black-box foundation models trained on external and possibly unstructured data, yielding an estimator of the risk ratio whose asymptotic variance is never larger than the one of the default estimator based on experimental data alone.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Victoria Mezger, Nils Krüger, Georg Hahn. 2026-09-30. Integration of external predictions for efficient estimation of the risk ratio. https://arxiv.org/abs/2610.00769

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

How many labelers do you have? A closer look at gold-standard labels

The construction of most supervised learning datasets revolves around collecting multiple labels for each instance, then aggregating the labels to form a type of "true" label. We question the wisdom of this pipeline by developing a (stylized) theoretical model of this process and analyzing its statistical consequences, showing how access to non-aggregated label information can make training well-calibrated models more feasible than it is with cleaned labels. The entire story, however, is subtle, and the contrasts between aggregated and fuller label information depend on the particulars of the problem, where estimators that use aggregated information exhibit robust but slower rates of convergence, while estimators that can effectively leverage all labels converge more quickly if they have fidelity to (or can learn) the true labeling process. The theory makes several predictions for real-world datasets, including when non-aggregate labels should improve learning performance, which we test to corroborate the validity of our predictions.

math.ST↗

A new class of colored Gaussian graphical models with explicit normalizing constants

We study Bayesian model selection in colored Gaussian graphical models (CGGMs), which combine sparsity of conditional independencies with symmetry constraints encoded by vertex- and edge-colored graphs. A key computational bottleneck in Bayesian inference for CGGMs is the evaluation of the Diaconis-Ylvisaker normalizing constants, given by gamma-type integrals over cones of precision matrices with prescribed zeros and equality constraints. We introduce a new subclass of RCON models for which these normalizing constants admit closed-form expressions. On the algebraic side, we identify conditions on the space of colored precision matrices that guarantee tractability of the associated integrals, leading to the notions of Block-Cholesky spaces (BC-spaces) and Diagonally Commutative Block-Cholesky spaces (DCBC-spaces). On the combinatorial side, we characterize the colored graphs inducing such spaces via a color perfect elimination ordering and a 2-path regularity condition, and define the resulting Color Elimination-Regular (CER) graphs and their symmetric variants. Our framework extends the theory of uncolored decomposable graphs to the colored setting, and the CER class contains all RCOP models associated with decomposable graphs. For models with a single vertex color, our framework reveals a close connection between DCBC-spaces and Bose-Mesner algebras. For models defined on BC-spaces, we derive finite product formulas for the normalizing constants in terms of gamma functions, matrix determinants, and model-specific structure constants. For RCON models whose matrix spaces are DCBC-spaces, we provide efficient methods for computing all structure constants and determinant factors required to evaluate these formulas. Our results substantially broaden the range of CGGMs amenable to Bayesian structure learning in high-dimensional applications.

math.ST↗

Randomized Spectral Inference for Hyperuniformity

We test hyperuniformity from one large realization of a stationary point process. Hyperuniformity is equivalent to the average $A_r$ of the structure factor over $B_r$ vanishing as $r\downarrow0$, so the target is a low-frequency average rather than the value of the structure factor at the origin, which need not exist. We estimate $A_r$ from squared Fourier coefficients at frequencies drawn uniformly from $B_r$; if these coefficients are asymptotically Gaussian at almost every fixed frequency, the squared coefficients converge to independent variables with mean $A_r$. Assuming $A_r=s+c r^α+\varepsilon_r$ with $|\varepsilon_r|\leq Lr^β$ and given exponents $0<α<β$, a two-radius extrapolation removes the $r^α$ term and estimates $s$ with deterministic error of order $r^β$. This gives confidence bounds and a one-sided test of $H_0:s=0$ with asymptotic level at most $γ$, consistent against every fixed $s>0$ as $R\to\infty$, then the number of sampled frequencies tends to infinity, and then $r\downarrow0$. Essentially free translation actions with completely positive entropy satisfy the Fourier assumption.

math.ST↗