arXiv Science⌕ Search

arXiv · 2610.06377

Calculating the Expected Value of Sample Information for Observational Studies affected by Confounding

Abstract

Background: The Expected Value of Sample Information (EVSI) quantifies the value of collecting additional evidence to inform a health economic model. Existing EVSI methods typically assume idealized data collection mechanisms, most commonly randomized controlled trials (RCTs). However, in many realistic contexts, additional evidence comes from observational studies, which are subject to confounding or other bias. In this work, we define a methodology to calculate EVSI when additional data are observational and confounded. Methods: First, we define a simulation-based framework in which confounded observational data are generated through Inverse Target Trial Emulation (ITTE), a methodology that generates observational data with controlled levels of confounding starting from initial level data or prior information on population structure. Then, we apply inverse probability weighting (IPW) to obtain an adjusted summary statistic of the data targeting the corresponding randomized estimand. EVSI is finally computed using a regression based approach. Moreover, we propose a computationally efficient method to determine the sample size required to recover the EVSI achievable under an idealized randomized design. Results: We apply the methodology to two health economic models: a Normal Normal conjugate model and a chemotherapy treatment model combining a decision tree and Markov structure. We show that EVSI computed from observational data, even after adjustment, is lower than EVSI based on randomized data. The loss in EVSI increases with the level of induced confounding. Conclusions: This methodology extends EVSI when future evidence is expected to be observational and affected by confounding, enabling value-of-information analysis in more realistic and feasible data collection scenarios. It also provides a principled approach to sample size planning when observational evidence is anticipated.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Luca Benetti, Gianluca Baio, Anna Heath. 2026-10-05. Calculating the Expected Value of Sample Information for Observational Studies affected by Confounding. https://arxiv.org/abs/2610.06377

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Robust Bayesian methods using amortized simulation-based inference

Bayesian simulation-based inference (SBI) methods are used in statistical models where simulation is feasible but the likelihood is intractable. Standard SBI methods can perform poorly in cases of model misspecification, and there has been much recent work on modified SBI approaches which are robust to misspecified likelihoods. However, less attention has been given to the issue of inappropriate prior specification, which is the focus of this work. In conventional Bayesian modelling, there will often be a wide range of prior distributions consistent with limited prior knowledge expressed by an expert. Choosing a single prior can lead to an inappropriate choice, possibly conflicting with the likelihood information. Robust Bayesian methods, where a class of priors is considered instead of a single prior, can address this issue. For each density in the prior class, a posterior can be computed, and the range of the resulting inferences is informative about posterior sensitivity to the prior imprecision. We consider density ratio classes for the prior and implement robust Bayesian SBI using amortized neural methods developed recently in the literature. We also discuss methods for checking for conflict between a density ratio class of priors and the likelihood, and sequential updating methods for examining conflict between different groups of summary statistics. The methods are illustrated for several simulated and real examples.

stat.ME↗

On the permutation equivariance principle for causal estimands

In many causal inference problems, multiple action variables share a common causal role yet lack a natural ordering. \revblue{We consider $K\geq2$ action variables, each evaluated under treatment or control conditions,} and formalize permutation equivariance, the principle that permuting the variables permutes the corresponding estimands in a trackable manner, hence preserving their scientific meaning. We characterize this principle algebraically and present a complete class of weighted permutation equivariant estimands capturing main effects and interactions of all orders. We discuss the interpretation and choice of weights and characterize residual-free estimands, whose inclusion--exclusion sum recovers the endpoint contrast between the all-treated and all-control configurations. \revblue{Applying our general framework to network interference yields a new hierarchy of direct, indirect, and overall effects, whose first-order aggregates recover the average effects of \citet{hu2022average}. We also identify the overall effects as mixed derivatives of expected welfare under independent Bernoulli assignment.} We illustrate the framework through the contexts of factorial studies, causal mediation, and network interference.

stat.ME↗

Testing the Deviation of a High-dimensional Mean

This paper investigates the high-dimensional one-sample equivalence testing problem for a reference mean vector $\bbmu_0$: $H_0:\|\bbmu-\bbmu_0\|>d_0$ versus $H_1: \|\bbmu-\bbmu_0\|\leq d_0$ for a pre-specified length threshold $d_0 > 0$. In high-dimensional regimes, classical equivalence tests suffer from severe asymptotic breakdowns. To resolve this difficulty, we propose a novel test statistic based on a two-armed bandit (TAB) U-process, which leverages positive/negative feedback mechanisms from control theory. We rigorously establish the weak convergence of TAB U-process to a novel two-dimensional stochastic differential equation (SDE), termed the U-bandit SDE. By employing in-depth analysis of this SDE, we characterize the asymptotic size and power of the resulting test statistic. The theoretical framework is extended to the two-sample setting and validated through finite-sample simulations and an empirical analysis of microbiome compositions.

stat.ME↗