arXiv ScienceSearch

arXiv · 2608.19849

Distributional Extrapolation for Interactions

Abstract

Predicting combinatorial effects from limited-range observations is a fundamental challenge in many scientific domains, including drug discovery and hyperparameter optimization. We study combinatorial extrapolation, where training data consists of axis-aligned samples with only one active covariate, while test-time inputs involve multiple simultaneously active covariates. We introduce DExtrI, a method for extrapolating interaction effects beyond the support of the training data. We provide theoretical guarantees characterizing when such extrapolation is possible. Empirical results on synthetic and real-world datasets demonstrate that DExtrI successfully generalizes to unseen combinations of covariates. Our approach enables applications such as predicting previously untested drug combinations and improving the efficiency of hyperparameter optimization.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Marin Šola, Xinwei Shen, Peter Bühlmann. 2026-08-20. Distributional Extrapolation for Interactions. https://arxiv.org/abs/2608.19849

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Modelling structural zeros in compositional data via a zero-censored multivariate normal model

We present a new model for analyzing compositional data with structural zeros. Inspired by \cite{butler2008} who suggested a model in the presence of zero values in the data we propose a model that treats the zero values in a different manner. Instead of projecting every zero value towards a vertex, we project them onto their corresponding edge and fit a zero-censored multivariate model.

stat.ME

Inventing a coin-flip classifier: Accounting for multiplicity in machine learning benchmark performance

State-of-the-art (SOTA) performance refers to the highest performance achieved by some model on a test sample, preferably under controlled conditions such as public data (reproducibility) or public challenges (independent sample). Thousands of classifiers are applied, and the highest performance becomes the new reference point for a particular problem. In effect, this set-up is an estimate of the expected best performance among all classifiers applied to a random sample; a sample maximum estimate. In this paper, we argue that SOTA should instead be estimated by the expected performance of the best classifier, which can be done without knowing which classifier it is. Our contribution is the formal distinction between the two, and an investigation into the practical consequences of using the former to estimate the latter. This is done by presenting sample maximum estimator distributions for non-identical and dependent classifiers. We illustrate the impact on real world examples from public challenges.

stat.ME

A computationally efficient multivariate volatility model for many assets

This paper develops a flexible and computationally efficient multivariate volatility model that accommodates dynamic conditional correlations and volatility spillover effects among financial assets. The new model has desirable properties such as identifiability and computational tractability for many assets.A sufficient condition for strict stationarity is derived for the process.Two quasi-maximum likelihood estimation methods are proposed for the new model without and with low-rank constraints on the coefficient matrices, respectively, and the asymptotic properties of both estimators are established. Moreover, a selection-consistent Bayesian information criterion is developed for order selection, and testing for volatility spillover effects is discussed. The finite-sample performance of the proposed methodology is evaluated in simulations for small and moderate dimensions.Its usefulness and inference tools are illustrated by two empirical examples for 5 stock markets and 17 industry portfolios.

stat.ME