arXiv ScienceSearch

arXiv · 2609.04243

Multi-dimensional Bias in Modeling Multi-dimensional Preferences: Evaluating the Ability of Synthetic Agents to Replace Human Participants in Conjoint Experiments

Abstract

Despite growing interest in using LLMs to add robustness or reduce data-collection costs in survey experiments, their efficacy in conjoint design---an increasingly popular method in political science---remains underexplored. This paper addresses that gap by investigating whether synthetic agents can reproduce the multi-dimensional human preference patterns that conjoint is designed to capture. It replicates published conjoint studies and compares the results generated by synthetic agents with original human data along three dimensions: representational correspondence, inferential correspondence, and procedural stability. Our analysis evaluates the alignment of choice distributions as well as the statistical and substantive similarity of estimates, and the results are uneven across these dimensions and studies replicated. This implies that the validity of synthetic participants should be considered claim-dependent and hierarchical. Reproducing a figure or obtaining strong sign agreement is evidence of similar aggregate outputs, but not enough to support replacing human respondents. Our results suggest that the discipline as a whole must first map this innovation's boundaries across various levels before considering synthetic agents a robust substitute for human samples.

Explore related subjects

Keep this discovery

BibTeXRIS

Ho Ting Hung, Nachiket Midha, Victor Y. Wu, Yiwen Zhang. 2026-08-13. Multi-dimensional Bias in Modeling Multi-dimensional Preferences: Evaluating the Ability of Synthetic Agents to Replace Human Participants in Conjoint Experiments. https://arxiv.org/abs/2609.04243

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

Bias-Corrected Subspace Intersection: Minimax-Optimal Shared Subspace Estimation in Multi-View Data

Estimating a low-dimensional subspace shared across noisy data matrices is a fundamental problem in multi-view matrix estimation. We study this problem under the two-view JIVE model, where each data matrix contains shared and view-specific low-rank components. We demonstrate that standard plug-in subspace intersection, including AJIVE, suffers from a second-order bias caused by direction-dependent leakage of the empirical singular vectors. We propose bias-corrected subspace intersection (BCSI), which removes this bias before estimating the shared subspace. We establish finite-sample risk bounds for BCSI that accommodate unequal view dimensions, signal strengths, and view-specific ranks and require no condition-number assumptions on the signal matrices. When the shared and view-specific ranks are comparable, these bounds match our minimax lower bounds up to universal constants. The resulting minimax rate contains a new second-order term, arising from quadratic leakage perturbations relative to the shrinking spectral gap when the view-specific subspaces are nearly aligned. This term is absent from previous JIVE minimax lower bounds. Numerical experiments demonstrate the advantage of BCSI over AJIVE when the leakage bias is pronounced. Along the way, we establish a nonasymptotic concentration result for the bias-corrected leakage Gram matrix of a rectangular spiked matrix, which may be of independent interest.

stat.ME

A simple derivation of the Kalman filter

In this lecture note, we present a concise and self-contained derivation of the discrete-time Kalman filter equations that requires only a basic understanding of least squares estimation. The treatment is designed to minimize mathematical overhead while preserving both rigor and generality.

math.OC

Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs

Large Language Models (LLMs) exhibit systematic biases across demographic groups. Auditing is proposed as an accountability tool for black-box LLM applications, but suffers from resource-intensive query access. We conceptualise auditing as uncertainty estimation over a target fairness metric and introduce BAFA, the Bounded Active Fairness Auditor for query-efficient auditing of black-box LLMs. BAFA maintains a version space of surrogate models consistent with queried scores and computes uncertainty intervals for fairness metrics (e.g., $Δ$ AUC) via constrained empirical risk minimisation. Active query selection narrows these intervals to reduce estimation error. We evaluate BAFA on two standard fairness dataset case studies: \textsc{CivilComments} and \textsc{Bias-in-Bios}, comparing against stratified sampling, power sampling, and ablations. BAFA achieves target error thresholds with up to 40$\times$ fewer queries than stratified sampling (e.g., 144 vs 5,956 queries at $\varepsilon=0.02$ for \textsc{CivilComments}) for tight thresholds, demonstrates substantially better performance over time, and shows lower variance across runs. These results suggest that active sampling can reduce resources needed for independent fairness auditing with LLMs, supporting continuous model evaluations.

cs.LG