arXiv ScienceSearch

arXiv subjects

Tom Boot

Publications and source records attributed to Tom Boot.

7 recordsLinked to original sources

Transformer-based CoVaR: Systemic Risk in Textual Information

Conditional Value-at-Risk (CoVaR) quantifies systemic financial risk by measuring the loss quantile of one asset, conditional on another asset experiencing distress. We develop a Transformer-based methodology that integrates financial news articles directly with market data to improve CoVaR estimates. Unlike approaches that use predefined sentiment scores, our method incorporates raw text embeddings generated by a large language model (LLM). We prove explicit error bounds for our Transformer CoVaR estimator, showing that accurate CoVaR learning is possible even with small datasets. Using U.S. market returns and Reuters news items from 2006--2013, our out-of-sample results show that textual information impacts the CoVaR forecasts. With better predictive performance, we identify a pronounced negative dip during market stress periods across several equity assets when comparing the Transformer-based CoVaR to both the CoVaR without text and the CoVaR using traditional sentiment measures. Our results show that textual data can be used to effectively model systemic risk without requiring prohibitively large data sets.

econ.EM

Diffusion index forecasts under weaker loadings: PCA, ridge regression, and random projections

We study the accuracy of forecasts in the diffusion index forecast model with possibly weak loadings. The default option to construct forecasts is to estimate the factors through principal component analysis (PCA) on the available predictor matrix, and use the estimated factors to forecast the outcome variable. Alternatively, we can directly relate the outcome variable to the predictors through either ridge regression or random projections. We establish that forecasts based on PCA, ridge regression and random projections are consistent for the conditional mean under the same assumptions on the strength of the loadings. However, under weaker loadings the convergence rate is lower for ridge and random projections if the time dimension is small relative to the cross-section dimension. We assess the relevance of these findings in an empirical setting by comparing relative forecast accuracy for monthly macroeconomic and financial variables using different window sizes. The findings support the theoretical results, and at the same time show that regularization-based procedures may be more robust in settings not covered by the developed theory.

econ.EM

Inference on LATEs with covariates

In theory, two-stage least squares (TSLS) identifies a weighted average of covariate-specific local average treatment effects (LATEs) from a saturated specification, without making parametric assumptions on how available covariates enter the model. In practice, TSLS is severely biased as saturation leads to a large number of control dummies and an equally large number of, arguably weak, instruments. This paper derives asymptotically valid tests and confidence intervals for the weighted average of LATEs that is targeted, yet missed by saturated TSLS. The proposed inference procedure is robust to unobserved treatment effect heterogeneity, covariates with rich support, and weak identification. We find LATEs statistically significantly different from zero in applications in criminology, finance, health, and education.

econ.EM

Identification- and many moment-robust inference via invariant moment conditions

Identification-robust hypothesis tests are commonly based on the continuous updating GMM objective function. When the number of moment conditions grows proportionally with the sample size, the large-dimensional weighting matrix prohibits the use of conventional asymptotic approximations and the behavior of these tests remains unknown. We show that the structure of the weighting matrix opens up an alternative route to asymptotic results when, under the null hypothesis, the distribution of the moment conditions satisfies a symmetry condition known as reflection invariance. We provide several examples in which the invariance follows from standard assumptions. Our results show that existing tests will be asymptotically conservative, and we propose an adjustment to attain nominal size in large samples. We illustrate our findings through simulations for various linear and nonlinear models, and an empirical application on the effect of the concentration of financial activities in banks on systemic risk.

econ.EM

Uniform Inference in Linear Error-in-Variables Models: Divide-and-Conquer

It is customary to estimate error-in-variables models using higher-order moments of observables. This moments-based estimator is consistent only when the coefficient of the latent regressor is assumed to be non-zero. We develop a new estimator based on the divide-and-conquer principle that is consistent for any value of the coefficient of the latent regressor. In an application on the relation between investment, (mismeasured) Tobin's $q$ and cash flow, we find time periods in which the effect of Tobin's $q$ is not statistically different from zero. The implausibly large higher-order moment estimates in these periods disappear when using the proposed estimator.

econ.EM

Unbiased estimation of the OLS covariance matrix when the errors are clustered

When data are clustered, common practice has become to do OLS and use an estimator of the covariance matrix of the OLS estimator that comes close to unbiasedness. In this paper we derive an estimator that is unbiased when the random-effects model holds. We do the same for two more general structures. We study the usefulness of these estimators against others by simulation, the size of the $t$-test being the criterion. Our findings suggest that the choice of estimator hardly matters when the regressor has the same distribution over the clusters. But when the regressor is a cluster-specific treatment variable, the choice does matter and the unbiased estimator we propose for the random-effects model shows excellent performance, even when the clusters are highly unbalanced.

econ.EM

Scalable simultaneous inference in high-dimensional linear regression models

The computational complexity of simultaneous inference methods in high-dimensional linear regression models quickly increases with the number variables. This paper proposes a computationally efficient method based on the Moore-Penrose pseudoinverse. Under a symmetry assumption on the available regressors, the estimators are normally distributed and accompanied by a closed-form expression for the standard errors that is free of tuning parameters. We study the numerical performance in Monte Carlo experiments that mimic the size of modern applications for which existing methods are computationally infeasible. We find close to nominal coverage, even in settings where the imposed symmetry assumption does not hold. Regularization of the pseudoinverse via a ridge adjustment is shown to yield possible efficiency gains.

math.ST