arXiv Science⌕ Search

arXiv · 2610.11279

Anchored multiple testing: a transparent use of e-closure to improve FDR procedures

Abstract

The recent e-closure method can recover every procedure that controls FDR (and other expectation losses). But the recovered e-collection is ``self-referential'' and gives no insight on how to improve the procedure (if improvable), and some recent improvements have been somewhat opaque. We introduce an elementary new technique called anchoring that exploits looseness in existing FDR proofs to enlarge a baseline multiple-testing procedure's self-referential local e-value. The resulting e-closure thus transparently retains every baseline discovery (and usually adding more) and controlling the false discovery rate under the same conditions as the baseline. To show that this principle is broadly applicable, we use it to improve a large suite of multiple testing procedures: (i) Anchored-BH dominates the Benjamini-Hochberg (BH) procedure under PRDS while being incomparable to Goeman's recent closed-BH, (ii) Anchored-BY dominates the Benjamini-Yekutieli (BY) procedure under arbitrary dependence while being incomparable to closed-BY, (iii) For two-sided Gaussian p-values (under appropriate covariance conditions), Anchored-2BH dominates running BH twice at half the level on two one-sided p-values, (iv) Anchored-dBH dominates dependence-adjusted BH, (v) Anchored e-BH dominates e-BH and is incomparable to closed e-BH, (vi) Anchored SeqStep+ improves the original (including selective and adpative variants) while preserving ordered rejection structures. All of these are accomplished in sorting or quadratic time. The appendix shows how to dominate Shifted-BH (for two-sided arbitrarily correlated Gaussians) and NDBH (under negative dependent p-values).

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Aaditya Ramdas. 2026-10-08. Anchored multiple testing: a transparent use of e-closure to improve FDR procedures. https://arxiv.org/abs/2610.11279

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Inference in generalized linear models with robustness to misspecified variances

Generalized linear models usually assume a common dispersion parameter, an assumption that is seldom true in practice. Consequently, standard parametric methods may suffer appreciable loss of type I error control. As an alternative, we present a semi-parametric group-invariance method based on sign flipping of score contributions. Our method requires only the correct specification of the mean model, but is robust against any misspecification of the variance. We present tests for single as well as multiple regression coefficients. The test is asymptotically valid but shows excellent performance in small samples. We illustrate the method using RNA sequencing count data, for which it is difficult to model the overdispersion correctly. The method is available in the R library flipscores.

stat.ME↗

Sample-Efficient "Clustering and Conquer" Procedures for Parallel Large-Scale Ranking and Selection

This work aims to improve the sample efficiency of parallel large-scale ranking and selection (R&S) problems by leveraging correlation information. We modify the commonly used "divide and conquer" framework in parallel computing by adding a correlation-based clustering step, transforming it into "clustering and conquer". Theoretically, we develop a novel gradient-based analysis framework and show that this seemingly simple modification substantially improves the performance of large-scale R&S procedures. Our approach enjoys two key advantages: (1) it does not require highly accurate correlation estimation or precise clustering, and (2) it can be seamlessly integrated with various existing fixed-precision and fixed-budget R&S procedures while achieving optimal sample complexity. We also introduce a new parallel clustering algorithm tailored to large-scale settings. Finally, in large-scale AI applications such as neural architecture search, our methods demonstrate superior performance.

stat.ME↗

A Dynamic Factor Model for Multivariate Counting Process Data

We propose a dynamic multiplicative factor model for process data arising from complex problem-solving items, an emerging type of data in large-scale educational assessment. The proposed model can be viewed as an extension of the classical frailty models developed in survival analysis for multivariate recurrent event times, but with two important distinctions: (i) the factor (frailty) is of primary interest; (ii) covariates are internal and embedded in the factor. It allows us to explore low-dimensional structure with meaningful interpretation. We show that the proposed model is generically identifiable and that the maximum likelihood estimators are consistent and asymptotically normal. Furthermore, to obtain a parsimonious model and to improve the interpretation of parameters, variable selection and estimation for both fixed and random effects are developed through suitable penalisation. The computation is carried out using a stochastic EM algorithm with elliptical slice sampling in the stochastic E-step and coordinate descent in the M-step. Simulation studies demonstrate that the proposed approach effectively recovers the true structure. The proposed method is applied to the analysis of the log file of an item from the Programme for the International Assessment of Adult Competencies (PIAAC), and meaningful relationships are identified.

stat.ME↗