arXiv Science⌕ Search

arXiv · 2610.03006

Psychometric Tests: Quantifying the Consequences of Low Reliability and Improving Reliability Estimation

Abstract

We introduce a set of consequence-based measures that quantify the misclassification arising from imperfect reliability in multi-item psychometric scales. Particular attention is given to the probability of individuals in extreme latent-trait percentiles being correctly identified from their observed scores. These cost functions provide a principled and interpretable way to characterise how inadequate reliability distorts classification and reduces the informational value of test scores. We also examine methodological issues in estimating reliability under common-factor models, with emphasis on McDonald's omega and the construction of accurate confidence intervals. To address the computational burden of model fitting, especially in small samples, we derive a modified version of Cronbach's alpha that closely approximates omega when the common-factor model holds. This estimator has comparable sampling variability to omega while requiring no parameter estimation, offering a practical and computationally efficient alternative. y

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rose Baker. 2026-10-02. Psychometric Tests: Quantifying the Consequences of Low Reliability and Improving Reliability Estimation. https://arxiv.org/abs/2610.03006

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Self-Tuned Rejection Sampling within Gibbs and a Case Study in Small Area Estimation

When formulating a Gibbs sampler, some conditionals may be unfamiliar distributions without well-known variate generation routines. Rejection sampling may be used to draw from such distributions exactly; however, it can be challenging to obtain practical proposal distributions. A practical proposal is one where accepted draws are not extremely rare occurrences and which is not too computationally intensive to use repeatedly within the Gibbs sampler. Consequently, approximate methods such as Metropolis-Hastings steps tend to be used in this setting. This work revisits the vertical weighted strips (VWS) method of proposal construction from arXiv:2401.09696 for univariate conditionals within Gibbs. VWS constructs a finite mixture based on the form of the target density and provides an upper bound on the rejection probability. The rejection probability can be reduced by refining terms in the finite mixture. Naïvely constructing a new proposal for each target encountered in a Gibbs sampler can be computationally impractical. Instead, we consider proposal distributions which persist over the Gibbs sampler and tune themselves gradually to avoid very high rejection probabilities while discarding mixture terms with low contribution. We explore a motivating application in small area estimation, applied to the estimation of county-level population counts of school-aged children in poverty. Here, a Gibbs sampler for a Bayesian model of interest includes a family of unfamiliar densities to be drawn for each observation in the data. Self-tuned VWS is applied to obtain exact draws within Gibbs while keeping the computational workload of proposal maintenance under control.

stat.ME↗

The Whittle likelihood for mixed models with application to groundwater level time series

Understanding the processes that influence groundwater levels is crucial for forecasting and responding to hazards such as groundwater droughts. Mixed models, which combine a fixed mean, expressed using independent predictors, with autocorrelated random errors, are used for inference, forecasting and filling in missing values in groundwater level time series. Estimating parameters of mixed models using maximum likelihood has high computational complexity. For large datasets, this leads to restrictive simplifying assumptions such as fixing certain free parameters in practical implementations. In this paper, we propose a method to jointly estimate all parameters of mixed models using the Whittle likelihood, a frequency-domain quasi-likelihood. Our method is robust to missing and non-Gaussian data and can handle much larger data sizes. We demonstrate the utility of our method both in a simulation study and with real-world data, comparing against maximum likelihood and an alternative two-stage approach that estimates fixed and random effect parameters separately.

stat.ME↗

Empirical Bayes Method for Large-Scale Multiple Testing with Heteroscedastic Errors

In this paper, we address the normal mean inference problem, which involves testing multiple means of normal random variables with heteroscedastic variances. Most existing empirical Bayes methods for this setting are developed under restrictive assumptions, such as the scaled inverse-chi-squared prior for variances and unimodality for the non-null mean distribution. However, when either of these assumptions is violated, these methods often fail to control the false discovery rate (FDR) at the target level or suffer from a substantial loss of power. To overcome these limitations, we propose a new empirical Bayes method, gg-Mix, which assumes only independence between the normal means and variances, without imposing any structural restrictions on their distributions. We thoroughly evaluate the FDR control and power of gg-Mix through extensive numerical studies and demonstrate its superior performance compared to existing methods. Finally, we apply gg-Mix to three real data examples to further illustrate the practical advantages of our approach.

stat.ME↗