arXiv ScienceSearch

arXiv subjects

Matteo Barigozzi

Publications and source records attributed to Matteo Barigozzi.

At least 19 recordsLinked to original sources

Mean Square Errors of factors extracted using principal components, linear projections, and Kalman filter

Factor extraction from systems of variables with a large cross-sectional dimension, $N$, is often based on either Principal Components (PC)-based procedures, or Kalman filter (KF)-based procedures. Measuring the uncertainty of the extracted factors is important when, for example, they have a direct interpretation and/or they are used to summarized the information in a large number of potential predictors. In this paper, we compare the finite $N$ mean square errors (MSEs) of PC and KF factors extracted under different structures of the idiosyncratic cross-correlations. We show that the MSEs of PC-based factors, implicitly based on treating the true underlying factors as deterministic, are larger than the corresponding MSEs of KF factors, obtained by treating the true factors as either serially independent or autocorrelated random variables. We also study and compare the MSEs of PC and KF factors estimated when the idiosyncratic components are wrongly considered as if they were cross-sectionally homoscedastic and/or uncorrelated. The relevance of the results for the construction of confidence intervals for the factors are illustrated with simulated data.

econ.EM

Measuring the Euro Area Output Gap

We measure the Euro Area (EA) output gap and potential output using a non-stationary dynamic factor model estimated on a large dataset of macroeconomic and financial variables. Our results indicate that, between 2012 and 2024, the EA economy was consistently tighter than suggested by institutional estimates, implying that its weak growth reflects a potential output problem rather than a business-cycle one. Moreover, we find that the decline in trend inflation-rather than economic slack-kept core inflation below 2% before the pandemic, while demand forces explain at least 30% of the post-pandemic rise in core inflation.

econ.EM

Predicting Energy Demand with Tensor Factor Models

Hourly consumption from multiple providers displays pronounced intra-day, intra-week, and annual seasonalities, as well as strong cross-sectional correlations. We introduce a novel approach for forecasting high-dimensional U.S. electricity demand data by accounting for multiple seasonal patterns via tensor factor models. To this end, we restructure the hourly electricity demand data into a sequence of weekly tensors. Each weekly tensor is a three-mode array whose dimensions correspond to the hours of the day, the days of the week, and the number of providers. This multi-dimensional representation enables a factor decomposition that distinguishes among the various seasonal patterns along each mode: factor loadings over the hour dimension highlight intra-day cycles, factor loadings over the day dimension capture differences across weekdays and weekends, and factor loadings over the provider dimension reveal commonalities and shared dynamics among the different entities. We rigorously compare the predictive performance of our tensor factor model against several benchmarks, including traditional vector factor models and cutting-edge functional time series methods. The results consistently demonstrate that the tensor-based approach delivers superior forecasting accuracy at different horizons and provides interpretable factors that align with domain knowledge. Beyond its empirical advantages, our framework offers a systematic way to gain insight into the underlying processes that shape electricity demand patterns. In doing so, it paves the way for more nuanced, data-driven decision-making and can be adapted to address similar challenges in other high-dimensional time series applications.

stat.AP

Estimation of large approximate dynamic matrix factor models based on the EM algorithm and Kalman filtering

This paper considers an approximate dynamic matrix factor model that accounts for the time series nature of the data by explicitly modelling the time evolution of the factors. We study estimation of the model parameters based on the Expectation Maximization (EM) algorithm, implemented jointly with the Kalman smoother which gives estimates of the factors. We establish the consistency of the estimated loadings and factor matrices as the sample size $T$ and the matrix dimensions $p_1$ and $p_2$ diverge to infinity. We then extend this approach to: (a) the case of arbitrary patterns of missing data and (b) the presence of common stochastic trends. The finite sample properties of the estimators are assessed through a large simulation study and two applications on: (i) a financial dataset of volatility proxies and (ii) a macroeconomic dataset covering the main euro area countries.

stat.ME

Large datasets for the Euro Area and its member countries and the dynamic effects of the common monetary policy

We introduce EA-MD-QD, a new publicly available dataset comprising 1136 macroeconomic time series for the euro area (EA) and its ten largest member countries observed at monthly or quarterly frequency. Since January 2024, EA-MD-QD has been updated monthly and continuously revised, providing a valuable resource for policy analysis in the EA. Using EA-MD-QD, we study country-specific impulse responses to an EA-wide monetary policy shock. Results reveal moderate yet significant cross-country heterogeneity, with differences between so-called core countries, such as France and Germany, and peripheral countries, such as Italy and Spain, in their price and interest rate responses, together with meaningful differences in real activity, while stock price responses are relatively homogeneous. Evidence points to homeownership and saving behavior as potential drivers of the observed cross-country differences.

econ.EM

Moving sum procedure for multiple change point detection in large factor models

This paper proposes a moving sum methodology for detecting multiple change points in high-dimensional time series under a factor model, where changes are attributed to those in loadings as well as emergence or disappearance of factors. We establish the asymptotic null distribution of the proposed test for family-wise error control, and show the consistency of the procedure for multiple change point estimation. Simulation studies and an application to a large dataset of volatilities demonstrate the competitive performance of the proposed method.

stat.ME

The Dynamic, the Static, and the Weak: Factor models and the analysis of high-dimensional time series

Several fundamental and closely interconnected issues related to factor models are reviewed and discussed: dynamic versus static loadings, rate-strong versus rate-weak factors, the concept of weakly common component recently introduced by Gersing et al. (2023), the irrelevance of cross-sectional ordering and the assumption of cross-sectional exchangeability, the impact of undetected strong factors, and the problem of combining common and idiosyncratic forecasts. Conclusions all point to the advantages of the General Dynamic Factor Model approach of Forni et al. (2000) over the widely used Static Approximate Factor Model introduced by Chamberlain and Rothschild (1983).

econ.EM

Tail-robust factor modelling of vector and tensor time series in high dimensions

We study the problem of factor modelling vector- and tensor-valued time series in the presence of heavy tails in the data, which produce extreme observations with non-negligible probability. We propose to combine a two-step procedure for tensor decomposition with data truncation, which is easy to implement and does not require an iterative search for a numerical solution. Departing away from the light-tail assumptions often adopted in the time series factor modelling literature, we derive the consistency and asymptotic normality of the proposed estimators while assuming the existence of the $(2 + 2\epsilon)$-th moment only for some $\epsilon \in (0, 1)$. Our rates explicitly depend on $\eps$ characterising the effect of heavy tails, and on the chosen level of truncation. We also propose a consistent criterion for determining the number of factors. Simulation studies and applications to two macroeconomic datasets demonstrate the good performance of the proposed estimators.

stat.ME

General Spatio-Temporal Factor Models for High-Dimensional Random Fields on a Lattice

Motivated by the need for analysing large spatio-temporal panel data, we introduce a novel dimensionality reduction methodology for $n$-dimensional random fields observed across a number $S$ spatial locations and $T$ time periods. We call it General Spatio-Temporal Factor Model (GSTFM). First, we provide the probabilistic and mathematical underpinning needed for the representation of a random field as the sum of two components: the common component (driven by a small number $q$ of latent factors) and the idiosyncratic component (mildly cross-correlated). We show that the two components are identified as $n\to\infty$. Second, we propose an estimator of the common component and derive its statistical guarantees (consistency and rate of convergence) as $\min(n, S, T )\to\infty$. Third, we propose an information criterion to determine the number of factors. Estimation makes use of Fourier analysis in the frequency domain and thus we fully exploit the information on the spatio-temporal covariance structure of the whole panel. Synthetic data examples illustrate the applicability of GSTFM and its advantages over the extant generalized dynamic factor model that ignores the spatial correlations.

stat.ME

Dynamic Factor Models: a Genealogy

Dynamic factor models have been developed out of the need of analyzing and forecasting time series in increasingly high dimensions. While mathematical statisticians faced with inference problems in high-dimensional observation spaces were focusing on the so-called spiked-model-asymptotics, econometricians adopted an entirely and considerably more effective asymptotic approach, rooted in the factor models originally considered in psychometrics. The so-called dynamic factor model methods, in two decades, has grown into a wide and successful body of techniques that are widely used in central banks, financial institutions, economic and statistical institutes. The objective of this chapter is not an extensive survey of the topic but a sketch of its historical growth, with emphasis on the various assumptions and interpretations, and a family tree of its main variants.

econ.EM

Asymptotic equivalence of Principal Components and Quasi Maximum Likelihood estimators in Large Approximate Factor Models

This paper investigates the properties of Quasi Maximum Likelihood estimation of an approximate factor model for an $n$-dimensional vector of stationary time series. We prove that the factor loadings estimated by Quasi Maximum Likelihood are asymptotically equivalent, as $n\to\infty$, to those estimated via Principal Components. Both estimators are, in turn, also asymptotically equivalent, as $n\to\infty$, to the unfeasible Ordinary Least Squares estimator we would have if the factors were observed. We also show that the usual sandwich form of the asymptotic covariance matrix of the Quasi Maximum Likelihood estimator is asymptotically equivalent to the simpler asymptotic covariance matrix of the unfeasible Ordinary Least Squares. All these results hold in the general case in which the idiosyncratic components are cross-sectionally heteroskedastic, as well as serially and cross-sectionally weakly correlated. The intuition behind these results is that as $n\to\infty$ the factors can be considered as observed, thus showing that factor models enjoy a blessing of dimensionality.

econ.EM

The Canonical Decomposition of Factor Models: Weak Factors are Everywhere

We derive a novel canonical decomposition of factor models encompassing both the static factor model - where factors are loaded only contemporaneously - and the Generalised Dynamic Factor Model - where factors are loaded with lags. This decomposition features a new term: the weak common component, defined as the difference between the dynamic and static common components. It is driven by (possibly infinitely many) non-pervasive weak factors which belong to the dynamically common space. Through theoretical and empirical examples - both on U.S. macroeconomic indicators and global financial volatilities - we show that, in general, the weak common component shall not be neglected. Furthermore, we show that, by accounting for the presence of weak common components, we are likely to obtain Impulse Response Functions with more plausible shapes than those obtained from purely static approaches. In addition, we provide consistent estimators for all terms of the canonical decomposition and for the weak factors.

econ.EM

Hierarchical DCC-HEAVY Model for High-Dimensional Covariance Matrices

We introduce a HD DCC-HEAVY class of hierarchical-type factor models for high-dimensional covariance matrices, employing the realized measures built from higher-frequency data. The modelling approach features straightforward estimation and forecasting schemes, independent of the cross-sectional dimension of the assets under consideration, and accounts for sophisticated asymmetric dynamics in the covariances. Empirical analyses suggest that the HD DCC-HEAVY models have a better in-sample fit and deliver statistically and economically significant out-of-sample gains relative to the existing hierarchical factor model and standard benchmarks. The results are robust under different frequencies and market conditions.

econ.EM

Robust Tensor Factor Analysis

We consider (robust) inference in the context of a factor model for tensor-valued sequences. We study the consistency of the estimated common factors and loadings space when using estimators based on minimising quadratic loss functions. Building on the observation that such loss functions are adequate only if sufficiently many moments exist, we extend our results to the case of heavy-tailed distributions by considering estimators based on minimising the Huber loss function, which uses an $L_{1}$ -norm weight on outliers. We show that such class of estimators is robust to the presence of heavy tails, even when only the second moment of the data exists. We also propose a modified version of the eigenvalue-ratio principle to estimate the dimensions of the core tensor and show the consistency of the resultant estimators without any condition on the relative rates of divergence of the sample size and dimensions. Extensive numerical studies are conducted to show the advantages of the proposed methods over the state-of-the-art ones especially under the heavy-tailed cases. An import/export dataset of a variety of commodities across multiple countries is analyzed to show the practical usefulness of the proposed robust estimation procedure. An R package ``RTFA" implementing the proposed methods is available on R CRAN.

stat.ME

Quasi Maximum Likelihood Estimation of High-Dimensional Factor Models: A Critical Review

We review Quasi Maximum Likelihood estimation of factor models for high-dimensional panels of time series. We consider two cases: (1) estimation when no dynamic model for the factors is specified (Bai and Li, 2012, 2016); (2) estimation based on the Kalman smoother and the Expectation Maximization algorithm thus allowing to model explicitly the factor dynamics (Doz et al., 2012, Barigozzi and Luciani, 2019). Our interest is in approximate factor models, i.e., when we allow for the idiosyncratic components to be mildly cross-sectionally, as well as serially, correlated. Although such setting apparently makes estimation harder, we show, in fact, that factor models do not suffer of the {\it curse of dimensionality} problem, but instead they enjoy a {\it blessing of dimensionality} property. In particular, given an approximate factor structure, if the cross-sectional dimension of the data, $N$, grows to infinity, we show that: (i) identification of the model is still possible, (ii) the mis-specification error due to the use of an exact factor model log-likelihood vanishes. Moreover, if we let also the sample size, $T$, grow to infinity, we can also consistently estimate all parameters of the model and make inference. The same is true for estimation of the latent factors which can be carried out by weighted least-squares, linear projection, or Kalman filtering/smoothing. We also compare the approaches presented with: Principal Component analysis and the classical, fixed $N$, exact Maximum Likelihood approach. We conclude with a discussion on efficiency of the considered estimators.

econ.EM

Multidimensional dynamic factor models

This paper generalises dynamic factor models for multidimensional dependent data. In doing so, it develops an interpretable technique to study complex information sources ranging from repeated surveys with a varying number of respondents to panels of satellite images. We specialise our results to model microeconomic data on US households jointly with macroeconomic aggregates. This results in a powerful tool able to generate localised predictions, counterfactuals and impulse response functions for individual households, accounting for traditional time-series complexities depicted in the state-space literature. The model is also compatible with the growing focus of policymakers for real-time economic analysis as it is able to process observations online, while handling missing values and asynchronous data releases.

econ.EM

fnets: An R Package for Network Estimation and Forecasting via Factor-Adjusted VAR Modelling

The package fnets for the R language implements the suite of methodologies proposed by Barigozzi et al. (2022) for the network estimation and forecasting of high-dimensional time series under a factor-adjusted vector autoregressive model, which permits strong spatial and temporal correlations in the data. Additionally, we provide tools for visualising the networks underlying the time series data after adjusting for the presence of factors. The package also offers data-driven methods for selecting tuning parameters including the number of factors, vector autoregressive order and thresholds for estimating the edge sets of the networks of interest in time series analysis. We demonstrate various features of fnets on simulated datasets as well as real data on electricity prices.

stat.CO

Principal Component Analysis for High-Dimensional Approximate Factor Models in Time Series: Assumptions, Asymptotic Theory, and Identification

We consider estimation of large approximate factor models in high-dimensional panels of stationary time series using Principal Component Analysis (PCA). We review the key results establishing the necessary and sufficient conditions for consistency and asymptotic normality of the estimators, which hold when both the cross-sectional dimension $n$ and the sample size $T$ tend to infinity. Special emphasis is placed on identification. First, we show that the common and idiosyncratic components are identified only in the limit $n\to\infty$. Second, we discuss the restrictions required to uniquely determine factors and loadings and examine their consequences for statistical inference.

econ.EM