arXiv ScienceSearch

arXiv subjects

Jack Freestone

Publications and source records attributed to Jack Freestone.

4 recordsLinked to original sources

Response-guided knockoffs for directional FDR control in linear models

We consider the problem of feature selection in linear models with finite-sample control of the false discovery rate (FDR). While existing knockoff-based methods control the directional FDR, which penalises incorrect sign estimates, they do not target discoveries in a pre-specified direction, and their knockoff constructions are entirely response-agnostic. We introduce the response-guided knockoff filter, which leverages a noise-perturbed version of the response to guide knockoff construction toward features likely to have the target sign, while provably controlling the directional FDR. The method operates under a weaker sample-size requirement $n > p + 2$, compared to $n \geq 2p$ required by existing fixed-X generators. Simulations and HIV drug resistance experiments demonstrate power gains over existing methods.

stat.ME

Risk-Limiting Audits for Parliamentary Majorities

Existing methods for risk-limiting audits typically focus on certifying individual contests. In parliamentary elections, however, the politically relevant outcome is often whether a party has won enough seats to form government, not whether every reported seat outcome is correct. Extending on the work of Mohanty et al. (2019), we formulate the certification of a parliamentary majority as a partial conjunction testing problem: it is enough to verify that the reported winning party truly won at least a majority of its reported seats. Building on the SHANGRLA auditing framework, we construct a sequential audit statistic for the majority outcome by combining seat-level statistics. We then propose adaptive sampling strategies that allocate auditing effort across seats, including variants that learn to avoid spending excessive effort on seats that appear unlikely to have been truly won. Using simulations based on synthetic and real data, from the 2014 Indian Lok Sabha election, we show that auditing the parliamentary majority can substantially reduce the number of ballots inspected (by almost a thousand-fold) compared to certifying every reported winning seat.

stat.AP

A semi-supervised framework for diverse multiple hypothesis testing scenarios

Standard multiple testing procedures are designed to report a list of discoveries, or suspected false null hypotheses, given the hypotheses' p-values or test scores. Recently there has been a growing interest in enhancing such procedures by combining additional information with the primary p-value or score. In line with this idea, we develop RESET (REScoring via Estimating and Training), which splits the data into a training part and an estimating part so that any semi-supervised learning approach can factor in the available side information while maintaining finite-sample error-rate control. Our practical implementation, RESET Ensemble, selects from an ensemble of classification algorithms so that it is compatible with a range of multiple testing scenarios without the need for the user to select the appropriate one. We apply RESET to both p-value and competition based multiple testing problems and show that RESET is (1) power-wise competitive, (2) fast compared to most tools and (3) able to uniquely achieve finite sample false discovery rate or false discovery exceedance control, depending on the user's preference.

stat.ME

Bounding the FDP in competition-based control of the FDR

Competition-based approach to controlling the false discovery rate (FDR) recently rose to prominence when, generalizing it to sequential hypothesis testing, Barber and Cand\`es used it as part of their knockoff-filter. Control of the FDR implies that the, arguably more important, false discovery proportion is only controlled in an average sense. We present TDC-SB and TDC-UB that provide upper prediction bounds on the FDP in the list of discoveries generated when controlling the FDR using competition. Using simulated and real data we show that, overall, our new procedures offer significantly tighter upper bounds than ones obtained using the recently published approach of Katsevich and Ramdas, even when the latter is further improved using the interpolation concept of Goeman et al.

stat.ME