arXiv ScienceSearch

arXiv subjects

Sokbae Lee

Publications and source records attributed to Sokbae Lee.

At least 19 recordsLinked to original sources

Binary Classification with the Maximum Score Model and Linear Programming

This paper presents a computationally efficient method for binary classification using Manski's (1975, 1985) maximum score model when covariates are discretely distributed and parameters are partially but not point identified. We establish minimax-regret-optimal classification rules that take account of partial identification of the model's parameters. We bound misclassification probabilities and expected excess regret induced by sampling uncertainty. We also describe an extension of our method to continuous covariates. Our approach avoids the computational difficulty of maximum score estimation by reformulating the problem as two linear programs. Compared to parametric and nonparametric methods, our method balances extrapolation ability with minimal distributional assumptions. Monte Carlo simulations and empirical applications demonstrate its effectiveness and practical relevance.

econ.EM

Learning the Effect of Persuasion via Difference-In-Differences

We develop a difference-in-differences framework to measure the persuasive impact of informational treatments on behavior in staggered treatment settings. We introduce two causal parameters, the forward and backward average persuasion rates on the treated, which refine the average treatment effect on the treated. The forward rate excludes cases of "preaching to the converted," while the backward rate omits "talking to a brick wall" cases. The backward rate coincides with the probability of necessity from the literature on probabilities of causation. We identify both persuasion rates under a no-backlash condition and a parallel-trends assumption imposed on a known transformation of response probabilities, taking the identity link as the baseline and nonlinear links as sensitivity checks. We develop estimation and inference using GMM and a limited-information method. We demonstrate the usefulness of our framework with an application to a Chinese curriculum reform introduced across provinces at different times.

econ.EM

SAUSS: Stochastic Approximation with Unbiased Simulated Scores for Limited Dependent Variable Models

Multinomial choice models allow flexible substitution patterns but become computationally demanding with many alternatives or observations. With a fixed per-observation simulation budget, simulated maximum likelihood introduces simulation bias, while each optimization step requires a full-sample likelihood evaluation. We propose Stochastic Approximation with Unbiased Simulated Scores (SAUSS), an averaged stochastic approximation based on conditionally unbiased mini-batch score estimates. Each iteration uses a fixed mini-batch regardless of sample size. For multinomial probit, accept-reject sampling provides exact conditional draws and unbiased score estimates for any fixed number of accepted draws. Under local conditions, asymptotic theory for the averaged estimator and the partial-sum process of the SAUSS iterates incorporates mini-batch and simulation variability and supports random-scaling and plug-in inference. In simulations and an application, SAUSS gives comparable results in less than 1% of the computation time of simulated maximum likelihood. SAUSS extends to limited dependent variable models with conditional-expectation score representations and exact conditional sampling.

stat.ME

Cognitive Ability and Tournament Entry: Evidence from Three Korean Populations

We compare tournament entry among three Korean groups raised in different institutional environments: South Korea, North Korea, and China. Experiments with a random bonus show that North Korean (NK) refugees enter tournaments less often than South Korean (SK) participants and Korean-Chinese immigrants. The conditional NK-SK difference becomes statistically insignificant when Raven score is included. A choice model with probability weighting suggests that lower cognitive ability is associated with lower expected performance, more pessimistic beliefs, and, within the model, greater aversion to competition.

econ.GN

Statistical inference in two-stage observation models including algorithmic randomness

Randomized algorithms, such as random sampling, random projections, and stochastic optimization, are increasingly used to reduce the computational cost of modern statistical analysis. These algorithms introduce algorithmic randomness in addition to the sampling randomness in the data, and this extra source of variation complicates statistical inference. We develop a framework for valid inference in such two-stage observation models, where data are first generated from an underlying population process and are then analyzed through a randomized algorithm. Our method, called sub-randomization, runs the randomized algorithm multiple times at different computational scales and uses the auxiliary runs to approximate the conditional algorithmic error distribution. In important converging-scale settings, the procedure avoids estimating the limiting covariance matrix or other nuisance parameters in the limiting law. We illustrate the method in two settings where standard approaches can fail to achieve nominal coverage: inference from repeated observations with highly correlated noise, and confidence sets for the minimizers of stochastic optimization problems computed using momentum methods, with particular emphasis on the stochastic heavy ball algorithm.

stat.ME

Individual Shrinkage for Random Effects

This paper develops an approach to random effects estimation and individual-level forecasting in micropanels that targets individual accuracy rather than aggregate performance. The conventional shrinkage methods used in the literature, such as the James-Stein estimator and Empirical Bayes, target aggregate performance and can lead to inaccurate decisions at the individual level. We propose a class of shrinkage estimators with individual weights (IW) that leverage an individual's own history, instead of the cross-sectional dimension. This approach can help overcome the "tyranny of the majority" inherent in existing methods, while relying on weaker assumptions. A key contribution is addressing the challenge of obtaining feasible weights from short time-series data under parameter heterogeneity. We discuss the theoretical optimality of IW and recommend using feasible weights determined through a Minimax Regret analysis in practice.

econ.EM

Bounding Treatment Effects by Pooling Limited Information across Observations

We provide novel bounds on average treatment effects (on the treated) that are valid under an unconfoundedness assumption. Our bounds are designed to be robust in challenging situations, for example, when the conditioning variables take on a large number of different values in the observed sample, or when the overlap condition is violated. This robustness is achieved by only using limited "pooling" of information across observations. Namely, the bounds are constructed as sample averages over functions of the observed outcomes such that the contribution of each outcome only depends on the treatment status of a limited number of observations. No information pooling across observations leads to so-called "Manski bounds", while unlimited information pooling leads to standard inverse propensity score weighting. We explore the intermediate range between these two extremes and provide corresponding inference methods. We show in Monte Carlo experiments and through two empirical application that our bounds are indeed robust and informative in practice.

econ.EM

Treatment Effects with Targeting Instruments

Multivalued treatments are commonplace in applications. We explore the use of discrete-valued instruments to control for selection bias in this setting. Our discussion revolves around the concept of targeting: which instruments target which treatments. It allows us to establish conditions under which counterfactual averages and treatment effects are point- or partially-identified for composite complier groups. We explore the additional identifying power of a positive selection assumption. We illustrate its usefulness by revisiting the findings of Kline and Walters (2016) on the Head Start Impact Study. We derive informative bounds that suggest less beneficial effects of Head Start expansions than their parametric estimates.

econ.EM

Leave No One Undermined: Policy Targeting with Regret Aversion

While the importance of personalized policymaking is widely recognized, fully personalized implementation remains rare in practice, often due to legal, fairness or cost concerns. We study the problem of policy targeting for a regret-averse planner when training data gives a rich set of observables while the assignment rules can only depend on its subset. Our regret-averse criterion reflects a planner's concern about regret inequality across the population. This, in general, leads to a fractional optimal rule due to treatment effect heterogeneity beyond the average treatment effects conditional on the subset of observables. We propose a debiased empirical risk minimization approach to learn the optimal rule from data and establish favorable, new upper and lower bounds for the excess risk, indicating a convergence rate of 1/n and asymptotic efficiency in certain cases. We apply our approach to the National JTPA Study and the International Stroke Trial.

econ.EM

Bounding the Effect of Persuasion with Monotonicity Assumptions: Reassessing the Impact of TV Debates

Televised debates between presidential candidates are often regarded as the exemplar of persuasive communication. Yet, recent evidence from Le Pennec and Pons (2023) indicates that they may not sway voters as strongly as popular belief suggests. We revisit their findings through the lens of the persuasion rate and introduce a robust framework that does not require exogenous treatment, parallel trends, or credible instruments. Instead, we leverage plausible monotonicity assumptions to partially identify the persuasion rate and related parameters. Our results reaffirm that the sharp upper bounds on the persuasive effects of TV debates remain modest.

econ.EM

Empirical Bayes Estimation in Heterogeneous Coefficient Panel Models

We develop an empirical Bayes (EB) G-modeling framework for short-panel linear models with nonparametric prior for the random intercepts, slopes, dynamics, and non-spherical error variances. We establish identification and consistency of the nonparametric maximum likelihood estimator (NPMLE) under general conditions, and provide low-level sufficient conditions for several models of empirical interest. Conditions for regret consistency of the EB estimators are also established. The NPMLE is computed using a Wasserstein-Fisher-Rao gradient flow algorithm adapted to panel regressions. Using data from the Panel Study of Income Dynamics, we find that the slope coefficient for potential experience is substantially heterogeneous and negatively correlated with the random intercept, and that error variances and autoregressive coefficients vary significantly across individuals. The EB estimates reduce mean squared prediction errors relative to individual maximum likelihood estimates.

econ.EM

Policy Learning with Confidence

This paper introduces a rule for policy selection in the presence of estimation uncertainty, explicitly accounting for estimation risk. The rule belongs to the class of risk-aware rules on the efficient decision frontier, characterized as policies offering maximal estimated welfare for a given level of estimation risk. Among this class, the proposed rule is chosen to provide a reporting guarantee, ensuring that the welfare delivered exceeds a threshold with a pre-specified confidence level. We apply this approach to the allocation of a limited budget among social programs using estimates of their marginal value of public funds and associated standard errors.

econ.EM

SLIM: Stochastic Learning and Inference in Overidentified Models

We propose SLIM (Stochastic Learning and Inference in overidentified Models), a scalable stochastic approximation framework for nonlinear GMM. SLIM forms iterative updates from independent mini-batches of moments and their derivatives, producing unbiased directions that ensure almost-sure convergence. It requires neither a consistent initial estimator nor global convexity and accommodates both fixed-sample and random-sampling asymptotics. We further develop an optional second-order refinement achieving full-sample GMM efficiency and inference procedures based on random scaling and plug-in methods, including plug-in, debiased plug-in, and online versions of the Sargan--Hansen $J$-test tailored to stochastic learning. In Monte Carlo experiments based on a nonlinear demand system with 576 moment conditions, 380 parameters, and $n = 10^5$, SLIM solves the model in under 1.4 hours, whereas full-sample GMM in Stata on a powerful laptop converges only after 18 hours. The debiased plug-in $J$-test delivers satisfactory finite-sample inference, and SLIM scales smoothly to $n = 10^6$.

econ.EM

Persuasion Effects in Regression Discontinuity Designs

We develop a framework for identifying and estimating persuasion effects in regression discontinuity (RD) designs. The RD persuasion rate measures the probability that individuals at the threshold would take the action if exposed to a persuasive message, given that they would not take the action without exposure. We present identification results for both sharp and fuzzy RD designs, derive sharp bounds under various data scenarios, and extend the analysis to local compliers. Estimation and inference rely on local polynomial regression, enabling straightforward implementation with standard RD tools. Applications to public health and media illustrate its empirical relevance.

econ.EM

Individual Welfare Analysis: Random Quasilinear Utility, Independence, and Confidence Bounds

We introduce a novel framework for individual-level welfare analysis. It builds on a parametric model for continuous demand with a quasilinear utility function, allowing for heterogeneous coefficients and unobserved individual-good-level preference shocks. We obtain bounds on the individual-level consumer welfare loss at any confidence level due to a hypothetical price increase, solving a scalable optimization problem constrained by a novel confidence set under an independence restriction. This confidence set is computationally simple and robust to weak instruments, nonlinearity, and partial identification. The validity of the confidence set is guaranteed by our new results on the joint limiting distribution of the independence test by Chatterjee (2021). These results together with the confidence set may have applications beyond welfare analysis. Monte Carlo simulations and two empirical applications on gasoline and food demand demonstrate the effectiveness of our method.

econ.EM

Inference for parameters identified by conditional moment restrictions using a generalized Bierens maximum statistic

Many economic panel and dynamic models, such as rational behavior and Euler equations, imply that the parameters of interest are identified by conditional moment restrictions. We introduce a novel inference method without any prior information about which conditioning instruments are weak or irrelevant. Building on Bierens (1990), we propose penalized maximum statistics and combine bootstrap inference with model selection. Our method optimizes asymptotic power by solving a data-dependent max-min problem for tuning parameter selection. Extensive Monte Carlo experiments, based on an empirical example, demonstrate the extent to which our inference procedure is superior to those available in the literature.

econ.EM

The ET Interview: Professor Joel L. Horowitz

Joel L. Horowitz has made profound contributions to many areas in econometrics and statistics. These include bootstrap methods, semiparametric and nonparametric estimation, specification testing, nonparametric instrumental variables estimation, high-dimensional models, functional data analysis, and shape restrictions, among others. Originally trained as a physicist, Joel made a pivotal transition to econometrics, greatly benefiting our profession. Throughout his career, he has collaborated extensively with a diverse range of coauthors, including students, departmental colleagues, and scholars from around the globe. Joel was born in 1941 in Pasadena, California. He attended Stanford for his undergraduate studies and obtained his Ph.D. in physics from Cornell in 1967. He has been Charles E. and Emma H. Morrison Professor of Economics at Northwestern University since 2001. Prior to that, he was a faculty member at the University of Iowa (1982-2001). He has served as a co-editor of Econometric Theory (1992-2000) and Econometrica (2000-2004). He is a Fellow of the Econometric Society and of the American Statistical Association, and an elected member of the International Statistical Institute. The majority of this interview took place in London during June 2022.

econ.EM

Group Shapley Value and Counterfactual Simulations in a Structural Model

We propose a variant of the Shapley value, the group Shapley value, to interpret counterfactual simulations in structural economic models by quantifying the importance of different components. Our framework compares two sets of parameters, partitioned into multiple groups, and applying group Shapley value decomposition yields unique additive contributions to the changes between these sets. The relative contributions sum to one, enabling us to generate an importance table that is as easily interpretable as a regression table. The group Shapley value can be characterized as the solution to a constrained weighted least squares problem. Using this property, we develop robust decomposition methods to address scenarios where inputs for the group Shapley value are missing. We first apply our methodology to a simple Roy model and then illustrate its usefulness by revisiting two published papers.

econ.EM