arXiv ScienceSearch

arXiv subjects

Ryo Okui

Publications and source records attributed to Ryo Okui.

13 recordsLinked to original sources

Inference methods for unit-specific coefficients in panel data models with latent group structure

This paper introduces statistical inference procedures for unit-specific coefficients in panel data models, where the coefficients exhibit a latent group structure. The proposed methods achieve efficiency gains by clustering units into a small number of groups, while explicitly accounting for the statistical uncertainty of group assignments. The core idea is to integrate standard inference procedures, such as the $t$-test and Wald tests, with confidence sets for group membership. Two methods are proposed: the first takes the minimum of the test statistics over the confidence set for group membership, and the second corrects for bias caused by possible group misassignment. The former can produce shorter but possibly disconnected sets, while the latter guarantees connected, interpretable intervals at some cost in length. We also develop standard errors that are adjusted for possible group misassignment and valid even with short time periods, which may be of independent interest. Monte Carlo simulations demonstrate that our approach yields narrower confidence sets for units with relatively large error variances than unit-by-unit time-series methods. In contrast, ignoring statistical uncertainty in the group membership estimation leads to distortions in size and coverage. We illustrate the method with an empirical example that estimates the effect of the minimum wage in each U.S. state.

econ.EM

Robust Inference Methods for Latent Group Panel Models under Possible Group Non-Separation

We develop robust inference methods for general linear hypotheses in linear panel data models with latent group structure in the coefficients. We employ a selective conditional inference approach based on the conditional distribution of coefficient estimates given the group structure estimated from the data. The resulting inference procedures remain valid even when group separation fails (i.e., when the distributional properties of the group-specific coefficients are not established) and, because they account for uncertainty in estimating the group structure, they also improve on conventional asymptotic procedures in finite samples when separation does hold. Our tests are exactly valid under Gaussian errors with known variances and asymptotically valid under general error distributions. Unlike much of the post-clustering inference literature, which focuses on testing group homogeneity, our framework accommodates arbitrary linear restrictions on the group-specific coefficients. Inverting the conditional tests yields selective confidence sets with valid coverage conditional on the estimated group structure. We illustrate the methods through Monte Carlo simulations and an application to growth convergence clubs. Simulations demonstrate accurate size control and good power in finite samples, including in the presence of serial correlation and cross-sectional dependence. The applications show sharp differences between the traditional inference methods and robust methods proposed in this paper, illustrating the importance of taking the estimated group structure into account.

econ.EM

Inference on effect size after multiple hypothesis testing

Significant treatment effects are often emphasized when interpreting and summarizing empirical findings in studies that estimate multiple, possibly many, treatment effects. Under this kind of selective reporting, conventional treatment effect estimates may be biased and their corresponding confidence intervals may undercover the true effect sizes. We propose new estimators and confidence intervals that provide valid inferences on the effect sizes of the significant effects after multiple hypothesis testing. Our methods are based on the principle of selective conditional inference and complement a wide range of tests, including step-up tests and bootstrap-based step-down tests. Our approach is scalable, allowing us to study an application with over 370 estimated effects. We justify our procedure for asymptotically normal treatment effect estimators. We provide two empirical examples that demonstrate bias correction and confidence interval adjustments for significant effects. The magnitude and direction of the bias correction depend on the correlation structure of the estimated effects and whether the interpretation of the significant effects depends on the (in)significance of other effects.

econ.EM

Location Characteristics of Conditional Selective Confidence Intervals via Polyhedral Methods

We examine the location properties of a conditional selective confidence interval constructed via the polyhedral method. The interval is derived from the distribution of a test statistic conditional on the event of statistical significance. For a one-sided test, its behavior depends on whether the parameter is highly or only marginally significant. In the highly significant case, the interval closely resembles the conventional confidence interval that ignores selection. By contrast, when the parameter is only marginally significant, the interval may shift far to the left of zero, potentially excluding all a priori plausible parameter values. This "location problem" does not arise if significance is determined by a two-sided test or by a one-sided test with randomized response (e.g., data carving).

math.ST

A Uniform Confidence Band for the Marginal Treatment Effect Function

This paper presents a method for constructing uniform confidence bands for the marginal treatment effect (MTE) function. The shape of the MTE function provides insight into how the unobserved propensity to receive treatment relates to the treatment effect. Our approach visualizes the statistical uncertainty of an estimated function, facilitating inferences about the function's shape. The proposed method is computationally inexpensive and requires only minimal information: sample size, standard errors, kernel function, and bandwidth. We derive a Gaussian approximation for a local quadratic estimator and consider the approximation of the distribution of its supremum in polynomial order. Monte Carlo simulations demonstrate that our bands provide the desired coverage and are less conservative than those based on the Gumbel approximation. An empirical application based on the rural electrification program is included.

econ.EM

Recovering latent linkage structures and spillover effects with structural breaks in panel data models

This paper introduces a framework to analyze time-varying spillover effects in panel data. We consider panel models where a unit's outcome depends not only on its own characteristics (private effects) but also on the characteristics of other units (spillover effects). The linkage of units is allowed to be latent and may shift at an unknown breakpoint. We propose a novel procedure to estimate the breakpoint, linkage structure, spillover and private effects. We address the high-dimensionality of spillover effect parameters using penalized estimation, and estimate the breakpoint with refinement. We establish the super-consistency of the breakpoint estimator, ensuring that inferences about other parameters can proceed as if the breakpoint were known. The private effect parameters are estimated using a double machine learning method. The proposed method is applied to estimate the cross-country R&D spillovers, and we find that the R&D spillovers become sparser after the financial crisis.

econ.EM

Latent group structure in linear panel data models with endogenous regressors

This paper concerns the estimation of linear panel data models with endogenous regressors and a latent group structure in the coefficients. We consider instrumental variables estimation of the group-specific coefficient vector. We show that direct application of the Kmeans algorithm to the generalized method of moments objective function does not yield unique estimates. We newly develop and theoretically justify two-stage estimation methods that apply the Kmeans algorithm to a regression of the dependent variable on predicted values of the endogenous regressors. The results of Monte Carlo simulations demonstrate that two-stage estimation with the first stage modeled using a latent group structure achieves good classification accuracy, even if the true first-stage regression is fully heterogeneous. We apply our estimation methods to revisiting the relationship between income and democracy.

econ.EM

Convergence rate of estimators of clustered panel models with misclassification

We study kmeans clustering estimation of panel data models with a latent group structure and $N$ units and $T$ time periods under long panel asymptotics. We show that the group-specific coefficients can be estimated at the parametric root $NT$ rate even if error variances diverge as $T \to \infty$ and some units are asymptotically misclassified. This limit case approximates empirically relevant settings and is not covered by existing asymptotic results.

econ.EM

Panel Data Analysis with Heterogeneous Dynamics

This paper proposes a model-free approach to analyze panel data with heterogeneous dynamic structures across observational units. We first compute the sample mean, autocovariances, and autocorrelations for each unit, and then estimate the parameters of interest based on their empirical distributions. We then investigate the asymptotic properties of our estimators using double asymptotics and propose split-panel jackknife bias correction and inference based on the cross-sectional bootstrap. We illustrate the usefulness of our procedures by studying the deviation dynamics of the law of one price. Monte Carlo simulations confirm that the proposed bias correction is effective and yields valid inference in small samples.

econ.EM

Kernel Estimation for Panel Data with Heterogeneous Dynamics

This paper proposes nonparametric kernel-smoothing estimation for panel data to examine the degree of heterogeneity across cross-sectional units. We first estimate the sample mean, autocovariances, and autocorrelations for each unit and then apply kernel smoothing to compute their density functions. The dependence of the kernel estimator on bandwidth makes asymptotic bias of very high order affect the required condition on the relative magnitudes of the cross-sectional sample size (N) and the time-series length (T). In particular, it makes the condition on N and T stronger and more complicated than those typically observed in the long-panel literature without kernel smoothing. We also consider a split-panel jackknife method to correct bias and construction of confidence intervals. An empirical application and Monte Carlo simulations illustrate our procedure in finite samples.

econ.EM

Heterogeneous structural breaks in panel data models

This paper develops a new model and estimation procedure for panel data that allows us to identify heterogeneous structural breaks. We model individual heterogeneity using a grouped pattern. For each group, we allow common structural breaks in the coefficients. However, the number, timing, and size of these breaks can differ across groups. We develop a hybrid estimation procedure of the grouped fixed effects approach and adaptive group fused Lasso. We show that our method can consistently identify the latent group structure, detect structural breaks, and estimate the regression parameters. Monte Carlo results demonstrate the good performance of the proposed method in finite samples. An empirical application to the relationship between income and democracy illustrates the importance of considering heterogeneous structural breaks.

econ.EM

Confidence set for group membership

Our confidence set quantifies the statistical uncertainty from data-driven group assignments in grouped panel models. It covers the true group memberships jointly for all units with pre-specified probability and is constructed by inverting many simultaneous unit-specific one-sided tests for group membership. We justify our approach under $N, T \to \infty$ asymptotics using tools from high-dimensional statistics, some of which we extend in this paper. We provide Monte Carlo evidence that the confidence set has adequate coverage in finite samples.An empirical application illustrates the use of our confidence set.

econ.EM

Doubly Robust Uniform Confidence Band for the Conditional Average Treatment Effect Function

In this paper, we propose a doubly robust method to present the heterogeneity of the average treatment effect with respect to observed covariates of interest. We consider a situation where a large number of covariates are needed for identifying the average treatment effect but the covariates of interest for analyzing heterogeneity are of much lower dimension. Our proposed estimator is doubly robust and avoids the curse of dimensionality. We propose a uniform confidence band that is easy to compute, and we illustrate its usefulness via Monte Carlo experiments and an application to the effects of smoking on birth weights.

stat.ME