arXiv · 2610.05198
Same Predictions, Different Harms: Causal Auditing of Patient World Models
Abstract
Patient world models used for clinical trial simulation can agree on transition kernels and arm-specific risks, yet disagree on the fraction of patients harmed by switching treatment---the counterfactual quantity that matters for intervention-aware reasoning. We audit this reliability gap in a two-stage shared-response SCM: a categorical intermediate health state is followed by common terminal care. Under independent stages, the sharp harm interval has closed-form endpoints for at most three intermediate states, with an exactness boundary at four states. Declared dependence and response-mismatch budgets yield calibrated outer bounds when stage independence or complete mediation is relaxed; in a symmetric three-state model the entire sensitivity frontier is sharp, $[0,\min\{1/2,1/3+(ρ+δ)/2\}]$, and shows exactly how budgets erase the gain over endpoint-only bounds. Two eight-variable response LPs propagate interventional uncertainty for finite-sample audits. Exact witnesses verify attainability. On public clinical simulators (EpiCare; sepsis), native configurations show little resolved stage dependence and no additional joint-compatibility gain over pairwise transport---honest negative results for reliability claims. All experiments are locally reproducible; guarantees remain conditional on the stated causal model. The results provide a concrete protocol for deciding when a patient world model is safe to trust for counterfactual harm.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yicheng Qi, Xiyi Xiong. 2026-10-04. Same Predictions, Different Harms: Causal Auditing of Patient World Models. https://arxiv.org/abs/2610.05198
Cite the original work for its findings. Save a collection to share your selection of sources.