arXiv · 2610.11584
Embedding-Bias in Conditional Independence Testing
Abstract
To test conditional independence of $X$ and $Y$ given a text or an image $Z$, one conditions on an embedding $ψ(Z)$ in place of $Z$. The embedded test is valid if $Z$ is independent of $X$ or of $Y$ given $ψ(Z)$, which cannot be confirmed from data, and when this fails, the rejection probability under the null hypothesis can tend to one. We study this failure, and show that focusing on a specific form of dependence relaxes what the embedding must retain. For a residual correlation test inspired by the Generalised Covariance Measure, validity only requires that the parts of $\mathbb{E}[X \mid Z]$ and $\mathbb{E}[Y \mid Z]$ missed by $\mathbb{E}[X \mid ψ(Z)]$ and $\mathbb{E}[Y \mid ψ(Z)]$ are uncorrelated. Otherwise, we treat the discarded information as an omitted variable. Under the null hypothesis, the bias equals the absolute correlation of the missed parts times the geometric mean of two partial $R^2$ values. This identity yields a robust test valid under a declared tolerance for the geometric mean, which, like a sensitivity parameter, is not identified from the data. On synthetic data and text embeddings, the robust test holds its level approximately. On text generated by a language model, under an exact null hypothesis, every embedding, even the generator's own states, biases the embedded test.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Nikolaj Thams, Anton Rask Lundborg. 2026-10-08. Embedding-Bias in Conditional Independence Testing. https://arxiv.org/abs/2610.11584
Cite the original work for its findings. Save a collection to share your selection of sources.