arXiv · 2609.14934
Cross-Block Conditioning in Deep Boltzmann Machines for Statistical Data Fusion
Abstract
Statistical data fusion combines two panels that share a block of covariates but observe disjoint outcome blocks, and in its traditional form no row observes both outcomes at once. That rules out the discriminative criterion one would rather train a Deep Boltzmann Machine with, since multi-prediction training needs ground truth for whatever it holds out. We propose observed-block multi-prediction, which restricts the multi-prediction objective to targets drawn from what each row actually observes. It is well defined for any missingness pattern and reduces to the original criterion when rows are complete. Having a discriminative criterion that survives the setting lets us ask whether the joint model is needed at all, by separating what it contributes into a representation part and an inference part. On two datasets of different kinds, a consumer purchase panel and public-domain census microdata, over grids in sample size and covariate width spanning 40 cells and 200 runs per method, almost none of the fine-tuned DBM's advantage comes from generative pre-training, which is confined to the smallest sample size on one dataset and absent on the other. It comes from conditioning on one outcome block when predicting the other. This term amounts to +0.19 and +0.36 percentage points, is positive in all 40 cells, never decays as the panels grow (it is flat on one dataset and grows on the other), and requires neither a second hidden layer nor more inference. Against baselines tuned on validation and given the same conditioning, the fine-tuned DBM is the best method in 37 of the 40 cells. The imputers that can also condition on the other outcome block mostly lose accuracy when they do, whereas the DBM gains in every cell; since fusion data cannot validate that choice, this is the property that matters.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Junichiro Niimi. 2026-09-16. Cross-Block Conditioning in Deep Boltzmann Machines for Statistical Data Fusion. https://arxiv.org/abs/2609.14934
Cite the original work for its findings. Save a collection to share your selection of sources.