arXiv ScienceSearch

arXiv subjects

Stefan Wager

Publications and source records attributed to Stefan Wager.

At least 19 recordsLinked to original sources

What Would it Cost to End Extreme Poverty?

We study poverty minimization via direct transfers, framing this as a statistical learning problem while retaining the information constraints faced by real-world programs. Using nationally representative household consumption surveys from 34 countries that together account for 76% of the world's poor, we estimate that reducing the poverty rate to 1% (from a baseline of 13%) would cost $211 B nominal per year. This is 4.0 times the corresponding reduction in the aggregate poverty gap, but only 19% of the cost of universal basic income. Extrapolated globally, the results imply a cost of 0.28% of global GDP to (approximately) end extreme poverty.

econ.GN

PLRD: Partially Linear Regression Discontinuity Inference

Regression discontinuity designs have become one of the most popular research designs in empirical economics. We argue, however, that the widely used approaches to building confidence intervals in regression discontinuity designs often exhibit suboptimal behavior in practice. We propose a new estimator, the partially linear regression discontinuity (PLRD) estimator that, in set of a simulation studies carefully calibrated to twelve high-profile applications of regression discontinuity designs, has substantially lower estimation error than available comparison methods. Throughout our experiments, the confidence intervals built using PLRD are both valid and short. We also provide large-sample guarantees for PLRD. Our simulation study serves as a general template for how new econometric methods can be credibly evaluated relative to the existing alternatives by constructing simulation designs that generate synthetic data indistinguishable from the original data using the Wasserstein generative adversarial network methodology.

econ.EM

Experimenting under Stochastic Congestion

We study randomized experiments in a service system when stochastic congestion can arise from temporarily limited supply or excess demand. Such congestion gives rise to cross-unit interference between the waiting customers, and analytic strategies that do not account for this interference may be biased. In current practice, one of the most widely used ways to address stochastic congestion is to use switchback experiments that alternatively turn a target intervention on and off for the whole system. We find, however, that under a queueing model for stochastic congestion, the standard way of analyzing switchbacks is inefficient, and that estimators that leverage the queueing model can be materially more accurate. Additionally, we show how the queueing model enables estimation of total policy gradients from unit-level randomized experiments, thus giving practitioners an alternative experimental approach they can use without needing to pre-commit to a fixed switchback length before data collection.

eess.SY

Learning from a Biased Sample

The empirical risk minimization approach to data-driven decision making requires access to training data drawn under the same conditions as those that will be faced when the decision rule is deployed. However, in a number of settings, we may be concerned that our training sample is biased in the sense that some groups (characterized by either observable or unobservable attributes) may be under- or over-represented relative to the general population; and in this setting empirical risk minimization over the training set may fail to yield rules that perform well at deployment. We propose a model of sampling bias called conditional $Γ$-biased sampling, where observed covariates can affect the probability of sample selection arbitrarily much but the amount of unexplained variation in the probability of sample selection is bounded by a constant factor. Applying the distributionally robust optimization framework, we propose a method for learning a decision rule that minimizes the worst-case risk incurred under a family of test distributions that can generate the training distribution under $Γ$-biased sampling. We apply a result of Rockafellar and Uryasev to show that this problem is equivalent to an augmented convex risk minimization problem. We give statistical guarantees for learning a model that is robust to sampling bias via the method of sieves, and propose a deep learning algorithm whose loss function captures our robust learning target. We empirically validate our proposed method in a case study on prediction of mental health scores from health survey data and a case study on ICU length of stay prediction.

stat.ME

Treatment effect estimation under convergent network interference

Under network interference, a unit's observed outcome depends on the treatment assignment of its neighboring units in an exposure graph. Existing design-based asymptotic theory typically considers local interference by restricting neighborhood sizes in the exposure graph. Such methods do not apply to dense exposure graphs, so prior work has often adopted a superpopulation approach instead, imposing regularity through random-graph models. In this paper, we introduce a notion of convergence for a sequence of finite populations under anonymous interference. Building on the graph limit framework of Lovász and Szegedy, we show that large-scale geometry of the exposure graph can provide a source of regularity beyond sparsity assumptions or random-graph modeling. Under Bernoulli assignment, our convergence notion yields asymptotic normality of standard estimators for the average direct effect, even on dense, non-random exposure graphs. As a special case, graphon-based random-graph models studied in prior work generate finite populations that converge in our sense. Under these models, graph randomness generates exposure graphs with stable large-scale geometry, while first-order uncertainty in average direct effect estimation is driven by treatment assignment.

math.ST

Data Fusion for High-Resolution Estimation

High-resolution estimates of population health indicators are critical for precision public health. We propose a method for high-resolution estimation that fuses distinct data sources: an unbiased, low-resolution data source (e.g. aggregated administrative data) and a potentially biased, high-resolution data source (e.g. individual-level online survey responses). We assume that the potentially biased, high-resolution data source is generated from the population under a model of sampling bias where observables can have arbitrary impact on the probability of response but the difference in the log probabilities of response between units with the same observables is linear in the difference between sufficient statistics of their observables and outcomes. Our data fusion method learns a distribution that is closest (in the sense of KL divergence) to the online survey distribution and consistent with the aggregated administrative data and our model of sampling bias. This approach significantly reduces bias in high-resolution estimates compared to baselines that rely on a single data source alone on a testbed that includes repeated measurements of three indicators measured by both the (online) Household Pulse Survey and ground-truth data sources at two geographic resolutions over the same time period.

stat.ME

What is the Long-Term Value of Reliability?

We describe Chronos LTV, a system to measure the long-term impact of delays and other service defects on key business metrics. We use Markov decision processes to model customer interactions over time, and formalize our target estimand as the marginal policy effect with respect to moving the average delay rate. Given this setup, we show that we can identify long-term effects under a sequential unconfoundedness assumption where delays are as good as random given observed order characteristics; and can estimate these effects using a simple covariate-balancing algorithm.

stat.ME

Optimal Targeting in Dynamic Systems

Modern treatment targeting methods often rely on estimating a conditional average treatment effect (CATE) using machine learning tools. While effective in identifying who benefits from treatment on the individual level, these approaches typically overlook system-level dynamics that may arise when treatments induce strain on shared capacity. We study the problem of targeting in Markovian systems, where treatment decisions must be made one at a time as units arrive, and early decisions can impact later outcomes through delayed or limited access to resources. We show that optimal policies in such settings compare CATE-like quantities to state-specific thresholds, where each threshold reflects the expected cumulative impact on the system of treating an additional individual in the given state. We propose an algorithm that augments standard CATE estimation with state-level value iteration to estimate these thresholds from observational data. Theoretical results establish consistency and convergence guarantees, and empirical studies demonstrate that our method improves long-run outcomes considerably relative to individual-level CATE targeting rules and generic offline reinforcement learning algorithms.

stat.ME

Treatment Allocation under Uncertain Costs

We consider the problem of learning how to optimally allocate treatments whose cost is uncertain and can vary with pre-treatment covariates. This setting may arise in medicine if we need to prioritize access to a scarce resource that different patients would use for different amounts of time, or in marketing if we want to target discounts whose cost to the company depends on how much the discounts are used. Here, we show that the optimal treatment allocation rule under budget constraints is a thresholding rule based on priority scores (those with a higher score are treated first), and we propose a number of practical methods for learning these priority scores using data from a randomized trial. Our formal results leverage a statistical connection between our problem and that of learning heterogeneous treatment effects under endogeneity using an instrumental variable. We find our method to perform well in a number of empirical evaluations.

stat.ME

Estimating Dynamic Marginal Policy Effects under Sequential Unconfoundedness

We develop methods for estimating how infinitesimal policy changes affect long-term outcomes in dynamic systems. We show that dynamic marginal policy effects (MPEs) can be identified via tractable reduced-form expressions, and can be estimated under a general sequential unconfoundedness assumption. We also propose a doubly robust estimator for dynamic MPEs. Our approach does not require observing full dynamic state information (as is typically assumed for off-policy evaluation in Markov decision processes), and does not incur an exponential curse of horizon (as is typical in non-Markovian off-policy evaluation). We demonstrate practicality and robustness of our approach in a number of simulations, including one motivated by a dynamic pricing application where people use past prices to form a reference level for current prices.

stat.ME

Non-parametric Causal Inference in Dynamic Thresholding Designs

We consider causal inference in dynamic settings where treatment is assigned by thresholding a state variable that can change over time. There is a large literature on regression-discontinuity methods building on the fact that, in the static setting, treatment assignment via threshold crossing induces a quasi-experimental design that enables pragmatic causal inference. But dynamic settings involve challenges not present in the static setting, e.g., past treatments may affect current state and thus future treatments, and so existing regression-discontinuity methods do not apply. Here, we show that dynamic thresholding designs identify a marginal policy effect that nests the classical regression-discontinuity parameter in the static setting; and propose a tailored local linear regression estimator that is consistent for this marginal policy effect. We demonstrate our approach using an experiment that emulates real-world optimization of thresholds for continuous glucose monitoring using data generated from an FDA-approved simulator.

stat.ME

Unlocking Deep Demand Flexibility via Dynamic Signals

The rapid proliferation of distributed energy resources (DERs) and the electrification of residential loads offer significant potential for grid flexibility but pose stability challenges under static pricing regimes. Specifically, high levels of automation under static Time-of-Use (TOU) tariffs often induce ``device synchronization,'' where simultaneous responses from home energy management systems (HEMS) create artificial demand peaks that threaten grid stability. This paper proposes a privacy-preserving, one-way dynamic signaling framework to unlock deep demand flexibility from HEMS. We utilize a feedback-based learning algorithm that updates day-ahead price profiles based on aggregate substation demand and environmental contexts, effectively closing the loop between utility objectives and aggregated edge behaviors. The framework is rigorously validated using high-fidelity simulations on an 84-bus distribution network populated with hundreds of HEMS controlling diverse devices, including HVAC, PV, batteries, and flexible loads. Results demonstrate that the proposed mechanism achieves substantial reductions in both peak demand and total load variation. Extensive analyses across diverse climates and scalable deployments confirm the framework's robustness, indicating that dynamic pricing acts as a force multiplier for DERs, with peak shaving potential increasing significantly under high renewable penetration scenarios.

math.OC

Neyman Jackknife: Design-Based Variance Estimation for Causal Inference under Interference

We propose a framework, the Neyman Jackknife, for conservative variance estimation in finite-population causal inference under interference. Our approach provides a general, flexible blueprint that enables conservative variance estimation whenever we are able to recompute our target estimator with some treatment assignments omitted. In classical settings, our approach recovers estimators closely related to the Neyman estimator under SUTVA and the Newey-West HAC variance estimator for time series. Numerical experiments suggest that our general-purpose framework yields variance estimators that can match or even surpass the performance of baselines that were purpose-built for specific applications.

stat.ME

Nonparametric Regression Discontinuity Designs with Survival Outcomes

Quasi-experimental evaluations are central for generating real-world causal evidence and complementing insights from randomized trials. The regression discontinuity design (RDD) is a quasi-experimental design that can be used to estimate the causal effect of treatments that are assigned based on a running variable crossing a threshold. Such threshold-based rules are ubiquitous in healthcare, where predictive and prognostic biomarkers frequently guide treatment decisions. However, standard RD estimators rely on complete outcome data, an assumption often violated in time-to-event analyses where censoring arises from loss to follow-up. To address this issue, we propose a nonparametric approach that leverages doubly robust censoring corrections and can be paired with existing RD estimators. Our approach can handle multiple survival endpoints, long follow-up times, and covariate-dependent variation in survival and censoring. We discuss the relevance of our approach across multiple areas of applications and demonstrate its usefulness through simulations and the prostate component of the Prostate, Lung, Colorectal and Ovarian (PLCO) Cancer Screening Trial where our new approach offers several advantages, including higher efficiency and robustness to misspecification. We have also developed an open-source software package, $\texttt{rdsurvival}$, for the $\texttt{R}$ language.

stat.ML

Sequentially-Rerandomized Switchback Experiments

Large-scale online platforms and marketplace systems often evaluate new policies through experiments that randomize treatment across operational units (e.g., geographies, regions, or clusters) over many time periods. In these settings, standard A/B testing can be inefficient or unreliable due to a limited number of units, substantial cross-unit heterogeneity, non-stationarity, and potential carryover across periods. We propose Sequentially-Rerandomized Switchback Experiments (SRSB), a new experimental design that helps mitigate these challenges. SRSB re-randomizes treatment at each time period such as to enforce balance on pre-specified prognostic variables constructed from past observations. In the absence of carryover, SRSB improves precision by leveraging temporal dependence through balancing lagged outcomes and covariates; we develop finite-sample randomization inference under a sharp null as well as asymptotic inference as the number of periods grows. We then extend SRSB to settings with first-order carryover and introduce a blocked SRSB variant that rerandomizes within strata defined by the previous treatment to form stable and comparable "stay" groups. Extensive simulations demonstrate the practical gains and robustness of SRSB relative to standard switchback designs.

stat.ME

Switchback Experiments under Geometric Mixing

The switchback is an experimental design that measures treatment effects by repeatedly turning an intervention on and off for a whole system. Switchback experiments are a robust way to overcome cross-unit spillover effects; however, they are vulnerable to bias from temporal carryovers. In this paper, we consider properties of switchback experiments in Markovian systems that mix at a geometric rate. We find that, in this setting, standard switchback designs suffer considerably from carryover bias: Their estimation error decays as $T^{-1/3}$ in terms of the experiment horizon $T$, whereas in the absence of carryovers a faster rate of $T^{-1/2}$ would have been possible. We also show, however, that judicious use of burn-in periods can considerably improve the situation, and enables errors that decay as $\log(T)^{1/2}T^{-1/2}$. Our formal results are mirrored in an empirical evaluation.

stat.ME

Off-Policy Evaluation in Markov Decision Processes under Weak Distributional Overlap

Doubly robust methods hold considerable promise for off-policy evaluation in Markov decision processes (MDPs) under sequential ignorability: They have been shown to converge as $1/\sqrt{T}$ with the horizon $T$, to be statistically efficient in large samples, and to allow for modular implementation where preliminary estimation tasks can be executed using standard reinforcement learning techniques. Existing results, however, make heavy use of a strong distributional overlap assumption whereby the stationary distributions of the target policy and the data-collection policy are within a bounded factor of each other -- and this assumption is typically only credible when the state space of the MDP is bounded. In this paper, we re-visit the task of off-policy evaluation in MDPs under a weaker notion of distributional overlap, and introduce a class of truncated doubly robust (TDR) estimators which we find to perform well in this setting. When the distribution ratio of the target and data-collection policies is square-integrable (but not necessarily bounded), our approach recovers the large-sample behavior previously established under strong distributional overlap. When this ratio is not square-integrable, TDR is still consistent but with a slower-than-$1/\sqrt{T}$-rate; furthermore, this rate of convergence is minimax over a class of MDPs defined only using mixing conditions. We validate our approach numerically and find that, in our experiments, appropriate truncation plays a major role in enabling accurate off-policy evaluation when strong distributional overlap does not hold.

stat.ML

Admissibility of Completely Randomized Trials: A Large-Deviation Approach

When an experimenter has the option of running an adaptive trial, is it admissible to ignore this option and run a non-adaptive trial instead? We provide a negative answer to this question in the best-arm identification problem, where the experimenter aims to allocate measurement efforts judiciously to confidently deploy the most effective treatment arm. We find that, whenever there are at least three treatment arms, there exist simple adaptive designs that universally and strictly dominate non-adaptive completely randomized trials. This dominance is characterized by a notion called efficiency exponent, which quantifies a design's statistical efficiency when the experimental sample is large. Our analysis focuses on the class of batched arm elimination designs, which progressively eliminate underperforming arms at pre-specified batch intervals. We characterize simple sufficient conditions under which these designs universally and strictly dominate completely randomized trials. These results resolve the second open problem posed in Qin [2022].

stat.ML