arXiv Science⌕ Search

arXiv · 2609.39274

Adaptive inference for functionals of M-estimands

Abstract

Reinforcement learning and contextual bandit algorithms have become increasingly common in sequential decision-making applications. When these methods are deployed in high-stakes domains, there is growing interest not only in learning effective policies, but also in conducting statistical inference for quantities learned under adaptive data collection. However, classical procedures applied naively in these settings can fail: even when estimators are unbiased, their variance becomes path-dependent and as a result may not be asymptotically normal. A growing literature has emerged to ameliorate this problem, but solutions tend to be problem specific and often rely on correct specification of a working model. In this work, we develop a unified framework for constructing asymptotically valid confidence intervals to cover smooth functionals of nonparametric M-estimands under adaptive sampling. Under Neyman orthogonality, we provide two novel methods for performing inference: (1) a self-normalized statistic based on the realized quadratic variation of the influence function and (2) a statistic using a plug-in estimate of the conditional variance based on reweighted influence function increments. Our results allow for flexible nonparametric estimation of nuisance parameters and remain valid under model misspecification. Our theory is supported by a simulation study for a dynamic pricing application which demonstrates that this method can produce asymptotically valid confidence intervals where standard methods fail.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

James Leiner, Aurélien Bibaut, Nathan Kallus, Aaditya Ramdas, Koulik Khamaru. 2026-09-30. Adaptive inference for functionals of M-estimands. https://arxiv.org/abs/2609.39274

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Construction of optimal tests for symmetry on the torus and their quantitative error bounds

In this paper, we investigate the general problem of assessing symmetry in data points on the hyper-dimensional torus, a question that originally emerged in applications from bioinformatics and directional statistics. We develop optimal tests for symmetry for both scenarios where the center of symmetry is known and where it is unknown. Our new tests are not only valid under a given parametric hypothesis but also under a very broad class of symmetric distributions. The asymptotic behavior of the proposed tests is studied both under the null hypothesis and local alternatives. A key contribution of our paper is that we accompany our asymptotic results with error guarantees by deriving quantitative bounds on the distributional distance between the exact (unknown) distribution of the test statistic and its asymptotic counterpart by leveraging Stein's method. The finite-sample performance of the tests is evaluated through simulation studies, and their practical utility in bioinformatics is demonstrated via an application to protein folding data.

math.ST↗

Axioms for testing with data-dependent levels, e-values and p-values

The emerging literature on hypothesis testing with data-dependent and post-hoc significance levels relies on a particular extension of the Type-I error to data-dependent levels. Existing arguments for this extension are heuristic, and primarily motivated by a resulting connection to the e-value. Our first contribution is to show that it is uniquely characterized by three axioms: law-invariance, calibration to classical testing, and a mixing axiom. Inspired by a combination of Birnbaum's conditionality principle and Savage's sure-thing principle, the mixing axiom assumes that a test produced by randomly selecting between (in)valid tests must be (in)valid. Our second contribution is to show that three analogous axioms characterize the e-value as a continuous generalization of a test in a decision-theoretic framework. We recover the p-value by dropping part of the mixing axiom, showing that e-values correspond to those p-values for which a random choice between two invalid p-values cannot lead to a valid p-value. Finally, we show that the relationship between e-values and post-hoc testing goes through under much weaker axioms.

math.ST↗

Stochastic Inversion of Multivariate Uniform-Distribution-Preserving Transformations

A multivariate transformation of the unit cube with component transformations that are piecewise continuously differentiable and uniform distribution preserving (udp) is considered. A stochastic inverse transformation is defined using randomization to overcome the non-injective nature of the udp transformations. The inverse transformation preserves the uniform margins of a random vector distributed according to a copula and yields different copulas for different randomizations. A copula density transformation result for the multivariate stochastic inverse is proved and illustrated in the bivariate case.

math.ST↗