arXiv ScienceSearch

arXiv · 2510.17553

Relaxing the Assumption of Strongly Non-Informative Linkage Error in Secondary Regression Analysis of Linked Files

Abstract

Data analysis of files that are a result of linking records from multiple sources are often affected by linkage errors. Records may be linked incorrectly, or their links may be missed. In consequence, it is essential that such errors are taken into account to ensure valid post-linkage inference. Here, we propose an extension to a general framework for regression with linked covariates and responses based on a two-component mixture model, which was developed in prior work. This framework addresses the challenging case of secondary analysis in which only the linked data is available and information about the record linkage process is limited. The extension considered herein relaxes the assumption of strongly non-informative linkage in the framework according to which linkage does not depend on the covariates used in the analysis, which may be limiting in practice. The effectiveness of the proposed extension is investigated by simulations and a case study.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Priyanjali Bukke, Martin Slawski. 2025-10-20. Relaxing the Assumption of Strongly Non-Informative Linkage Error in Secondary Regression Analysis of Linked Files. https://arxiv.org/abs/2510.17553

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Robust Estimation and Inference with Categorical Data

Categorical data pose a distinctive robustness challenge: contingency-table cells need not have a meaningful magnitude, ordering, or metric, so departures must instead be assessed through discrepancies between observed cell frequencies and model-implied probabilities. We develop $C$-estimation, a unifying framework for robust estimation in structured categorical models. $C$-estimation limits the influence of large frequency discrepancies and can be applied to unconditional, composite, and regression models of categorical data. Building on minimum-disparity estimation, the framework also accommodates clipped nonsmooth loss functions. A common asymptotic theory establishes Fisher consistency, consistency for the population target and asymptotic normality under contamination, and sandwich covariance estimation. At a correctly specified model, regular $C$-estimators retain the first-order efficiency of maximum likelihood. To quantify global robustness, we derive computable lower and upper envelopes for maximum-bias curves and, for a Huber-like loss, connect its clipping constants to a global robustness bound. Simulations support the theory and illustrate the estimators' robustness. An application to questionnaire data illustrates robust estimation of a latent factor model and identifies misfitting response strings that may reflect careless responding. A software implementation is provided.

stat.ME

Multi-Attribute Preferences: A Transfer Learning Approach

We introduce a transfer-learning method based on the Bradley--Terry model for multi-attribute pairwise-comparison data. The aim is to estimate the log-worth parameters of one primary attribute while using information from related secondary attributes. The method first pools the primary data with data from informative secondary attributes. It then corrects this pooled estimate using the primary likelihood, with a ridge penalty controlling the size of the correction. When the informative set is unknown, we use held-out primary data to select secondary attributes. For a known informative set and under a pooled Bradley--Terry compatibility condition, we derive high-probability $\ell_\infty$ and $\ell_2$ error bounds. Under additional conditions, these bounds can have a smaller asymptotic order than the corresponding primary-only Bradley--Terry bounds. We also establish asymptotic normality for a one-step estimator. A simulation study evaluates the method under more general settings, and an application to consumer preferences for eba, a cassava-derived food product, illustrates its use and interpretation. An R package implementing the method is available at https://CRAN.R-project.org/package=BTTL.

stat.ME

Spatially Dependent Indian Buffet Processes

We develop a new stochastic process called spatially dependent Indian buffet processes (sIBP) for binary feature matrices of unbounded columns with spatial correlations between subjects, and propose general spatial factor models for various multivariate response variables. We introduce spatial dependency through the stick-breaking representation of the original Indian buffet process (IBP; Griffiths and Ghahramani, 2005, 2011) and latent Gaussian process for the logit-transformed breaking proportions to capture underlying spatial correlation. We show that sIBP retains the sparsity and finite-feature behavior of the original IBP, while its joint feature allocation probabilities are affected by spatial correlation. Using binomial expansion and Polya-gamma data augmentation, we provide an efficient Gibbs sampler for posterior computation. The usefulness of our sIBP is demonstrated through simulation studies and two applications for large-dimensional multinomial data of areal dialects and geographical distribution of multiple tree species.

stat.ME