Confidence Composition for Multiagent Language Model Systems
Multiagent language model systems, such as collaborative reasoning and debate, produce multiple correlated candidate answers and confidence signals. However, these signals are usually calibrated only at the individual agent level, and provide no principled confidence estimate for the system's final answer. We formulate this as a confidence composition problem where combining confidence across agents and reasoning stages while preserving both selective utility and probabilistic reliability. We study confidence-aware routing and log-odds pooling protocols that select among candidate answers and output a system-level confidence. Across five benchmarks, 30 heterogeneous and homogeneous model pairs, and two confidence estimators, our gated-fusion methods improve AUARC and reduce Brier score over single agent, standard debate, and selective debate baselines, while retaining competitive weighted F1-score as a correctness metric. We further show that our log-odds fusion is overconfident due to correlated intermediate signals. We propose a shared dependence discount that substantially improves reliability while preserving predictions.