arXiv ScienceSearch

arXiv · 2608.10305

COMPACT: Spectral Adjustment Scores from a Complete and Irreducible Causal Criterion

Abstract

Observational datasets frequently contain many baseline variables, yet investigators estimating causal effects may not know which variables to include in the adjustment set. Confounding information may also be distributed weakly across many variables. Propensity scores can simplify adjustment by reducing high-dimensional covariates to a scalar with binary treatment. Although the propensity score is the coarsest balancing score, this distributional optimality does not imply maximal specificity over causal graphs. We instead examine all causal graphs among a candidate score, treatment, and outcome while allowing latent variables. Under faithfulness, we identify the largest set of unconditional and conditional dependence relations whose truth is invariant to whether treatment causes the outcome, leaving treatment-effect estimation to the downstream analysis. This criterion defines the maximally specific graph class expressible through these relations. We then develop the proposed algorithm, which operationalizes the criterion through a generalized eigenvalue problem whose score space targets the span of a balancing coordinate and an outcome-guided coordinate. We show that sufficiently informative proxies can recover this span without direct observation of the adjustment variables, characterize the resulting estimation and causal errors, and establish bootstrap validity for the complete procedure. Simulations and a real-data application demonstrate superior performance over several alternatives.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Eric V. Strobl. 2026-08-10. COMPACT: Spectral Adjustment Scores from a Complete and Irreducible Causal Criterion. https://arxiv.org/abs/2608.10305

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Robust Estimation and Inference with Categorical Data

Categorical data pose a distinctive robustness challenge: contingency-table cells need not have a meaningful magnitude, ordering, or metric, so departures must instead be assessed through discrepancies between observed cell frequencies and model-implied probabilities. We develop $C$-estimation, a unifying framework for robust estimation in structured categorical models. $C$-estimation limits the influence of large frequency discrepancies and can be applied to unconditional, composite, and regression models of categorical data. Building on minimum-disparity estimation, the framework also accommodates clipped nonsmooth loss functions. A common asymptotic theory establishes Fisher consistency, consistency for the population target and asymptotic normality under contamination, and sandwich covariance estimation. At a correctly specified model, regular $C$-estimators retain the first-order efficiency of maximum likelihood. To quantify global robustness, we derive computable lower and upper envelopes for maximum-bias curves and, for a Huber-like loss, connect its clipping constants to a global robustness bound. Simulations support the theory and illustrate the estimators' robustness. An application to questionnaire data illustrates robust estimation of a latent factor model and identifies misfitting response strings that may reflect careless responding. A software implementation is provided.

stat.ME

Multi-Attribute Preferences: A Transfer Learning Approach

We introduce a transfer-learning method based on the Bradley--Terry model for multi-attribute pairwise-comparison data. The aim is to estimate the log-worth parameters of one primary attribute while using information from related secondary attributes. The method first pools the primary data with data from informative secondary attributes. It then corrects this pooled estimate using the primary likelihood, with a ridge penalty controlling the size of the correction. When the informative set is unknown, we use held-out primary data to select secondary attributes. For a known informative set and under a pooled Bradley--Terry compatibility condition, we derive high-probability $\ell_\infty$ and $\ell_2$ error bounds. Under additional conditions, these bounds can have a smaller asymptotic order than the corresponding primary-only Bradley--Terry bounds. We also establish asymptotic normality for a one-step estimator. A simulation study evaluates the method under more general settings, and an application to consumer preferences for eba, a cassava-derived food product, illustrates its use and interpretation. An R package implementing the method is available at https://CRAN.R-project.org/package=BTTL.

stat.ME

Spatially Dependent Indian Buffet Processes

We develop a new stochastic process called spatially dependent Indian buffet processes (sIBP) for binary feature matrices of unbounded columns with spatial correlations between subjects, and propose general spatial factor models for various multivariate response variables. We introduce spatial dependency through the stick-breaking representation of the original Indian buffet process (IBP; Griffiths and Ghahramani, 2005, 2011) and latent Gaussian process for the logit-transformed breaking proportions to capture underlying spatial correlation. We show that sIBP retains the sparsity and finite-feature behavior of the original IBP, while its joint feature allocation probabilities are affected by spatial correlation. Using binomial expansion and Polya-gamma data augmentation, we provide an efficient Gibbs sampler for posterior computation. The usefulness of our sIBP is demonstrated through simulation studies and two applications for large-dimensional multinomial data of areal dialects and geographical distribution of multiple tree species.

stat.ME