arXiv Science⌕ Search

arXiv · 2610.03181

Cellwise, Blockwise, and Casewise Robust Multiblock PCA for Sustainable and Inclusive Wellbeing in the EU

Abstract

Gross Domestic Product (GDP) is widely used to guide economic and social decision-making, but it provides only a partial view of wellbeing. For this reason, the European Union (EU) has launched initiatives to monitor sustainable and inclusive wellbeing beyond GDP, developing indicator frameworks that cover dimensions such as health, education, environment, and social inclusion. These indicators are naturally grouped into thematic areas, and policymakers are interested in understanding how these areas contribute to global wellbeing and which indicators explain differences across countries. However, such data are high-dimensional, contain missing values, and may include anomalies affecting entire observations, specific thematic areas, or individual indicators. We introduce blockwise outliers and propose bloccPCA, a robust multiblock PCA method that simultaneously handles casewise, blockwise, and cellwise outliers, as well as missing values. The method provides robust global components to summarize the overall structure of wellbeing, while preserving thematic-area contributions through robust blockcomponents. It also yields diagnostic tools to identify whether anomalies arise at the case, block, or cell level. Monte Carlo simulations and an application to the EU wellbeing dataset show that bloccPCA provides valuable insights into sustainable and inclusive wellbeing.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Arthur Daisomont, Fabio Centofanti, Mia Hubert. 2026-10-02. Cellwise, Blockwise, and Casewise Robust Multiblock PCA for Sustainable and Inclusive Wellbeing in the EU. https://arxiv.org/abs/2610.03181

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Self-Tuned Rejection Sampling within Gibbs and a Case Study in Small Area Estimation

When formulating a Gibbs sampler, some conditionals may be unfamiliar distributions without well-known variate generation routines. Rejection sampling may be used to draw from such distributions exactly; however, it can be challenging to obtain practical proposal distributions. A practical proposal is one where accepted draws are not extremely rare occurrences and which is not too computationally intensive to use repeatedly within the Gibbs sampler. Consequently, approximate methods such as Metropolis-Hastings steps tend to be used in this setting. This work revisits the vertical weighted strips (VWS) method of proposal construction from arXiv:2401.09696 for univariate conditionals within Gibbs. VWS constructs a finite mixture based on the form of the target density and provides an upper bound on the rejection probability. The rejection probability can be reduced by refining terms in the finite mixture. Naïvely constructing a new proposal for each target encountered in a Gibbs sampler can be computationally impractical. Instead, we consider proposal distributions which persist over the Gibbs sampler and tune themselves gradually to avoid very high rejection probabilities while discarding mixture terms with low contribution. We explore a motivating application in small area estimation, applied to the estimation of county-level population counts of school-aged children in poverty. Here, a Gibbs sampler for a Bayesian model of interest includes a family of unfamiliar densities to be drawn for each observation in the data. Self-tuned VWS is applied to obtain exact draws within Gibbs while keeping the computational workload of proposal maintenance under control.

stat.ME↗

The Whittle likelihood for mixed models with application to groundwater level time series

Understanding the processes that influence groundwater levels is crucial for forecasting and responding to hazards such as groundwater droughts. Mixed models, which combine a fixed mean, expressed using independent predictors, with autocorrelated random errors, are used for inference, forecasting and filling in missing values in groundwater level time series. Estimating parameters of mixed models using maximum likelihood has high computational complexity. For large datasets, this leads to restrictive simplifying assumptions such as fixing certain free parameters in practical implementations. In this paper, we propose a method to jointly estimate all parameters of mixed models using the Whittle likelihood, a frequency-domain quasi-likelihood. Our method is robust to missing and non-Gaussian data and can handle much larger data sizes. We demonstrate the utility of our method both in a simulation study and with real-world data, comparing against maximum likelihood and an alternative two-stage approach that estimates fixed and random effect parameters separately.

stat.ME↗

Empirical Bayes Method for Large-Scale Multiple Testing with Heteroscedastic Errors

In this paper, we address the normal mean inference problem, which involves testing multiple means of normal random variables with heteroscedastic variances. Most existing empirical Bayes methods for this setting are developed under restrictive assumptions, such as the scaled inverse-chi-squared prior for variances and unimodality for the non-null mean distribution. However, when either of these assumptions is violated, these methods often fail to control the false discovery rate (FDR) at the target level or suffer from a substantial loss of power. To overcome these limitations, we propose a new empirical Bayes method, gg-Mix, which assumes only independence between the normal means and variances, without imposing any structural restrictions on their distributions. We thoroughly evaluate the FDR control and power of gg-Mix through extensive numerical studies and demonstrate its superior performance compared to existing methods. Finally, we apply gg-Mix to three real data examples to further illustrate the practical advantages of our approach.

stat.ME↗