arXiv ScienceSearch

arXiv subjects

Luke Hagar

Publications and source records attributed to Luke Hagar.

14 recordsLinked to original sources

COBRA-DOSE: Copula-based Bayesian Model Averaging for Dose Selection

Early-phase clinical trials for dose selection typically enrol few patients and aim to identify doses that are both safe and promising for further study. While traditional approaches identify the maximum tolerated dose, modern trials for targeted therapies often seek the optimal biological dose, defined as the lowest dose achieving sufficient biological activity with acceptable safety. In immunology settings, assessment of biological activity is based on multiple biomarkers or clinical endpoints. Clinicians leading dose-selection efforts would thus benefit from transparent summaries of the probabilities of observing combinations of biomarker outcomes across doses. However, such inference is challenging in small samples where complex modelling assumptions are difficult to verify. To address this limitation, we propose COBRA-DOSE, a framework for posterior predictive inference based on two endpoints that models dependence via copulas and accounts for uncertainty in both marginal distributions and dependence structures through Bayesian model averaging. This approach avoids reliance on a single model and yields interpretable quantities for clinical decision making. We demonstrate the performance of COBRA-DOSE using DEN-181, a phase I immunology trial in rheumatoid arthritis. We also provide a general implementation of our approach through the CobraDose package in R.

stat.ME

Computationally Efficient Experimental Design with Generalized Posteriors

The hybrid approach to experimental design aims to control frequentist operating characteristics of Bayesian decision procedures. These operating characteristics are assessed by simulating sampling distributions of posterior summaries under assumed data-generation processes that also define posterior distributions. Model misspecification can distort effect estimation and compromise control over operating characteristics. Generalized posterior distributions are defined using generalized likelihoods that characterize data generation under fewer assumptions, enhancing the robustness of Bayesian analysis and study design. However, widely applicable and computationally efficient design methodology with generalized posteriors is lacking. We propose an efficient method to determine suitable sample sizes and decision criteria associated with generalized posteriors under the hybrid approach. Using theoretical results to model posterior summaries as linear functions of the sample size, we efficiently assess operating characteristics throughout the sample size space given simulations conducted at only two sample sizes. While the benefits of the proposed methodology are illustrated by redesigning an adaptive clinical trial with time-to-event outcomes, we overview our framework's broader applicability to experiments involving Bayesian analogues to M-estimation.

stat.ME

An Efficient Framework for Robust Sample Size Determination

In many settings, robust data analysis involves computational methods for uncertainty quantification and statistical inference. To design frequentist studies that leverage robust analysis methods, suitable sample sizes to achieve desired power are often found by estimating sampling distributions of p-values via intensive simulation. Moreover, most sample size recommendations rely heavily on assumptions about a single data-generating process. Consequently, robustness in data analysis does not by itself imply robustness in study design, as examining sample size sensitivity to data-generating assumptions typically requires further simulations. We propose an economical alternative for determining sample sizes that are robust to multiple data-generating mechanisms. Applying our theoretical results that model p-values as a function of the sample size, we assess power across the sample size space using simulations conducted at only two sample sizes for each data-generating mechanism. We demonstrate the broad applicability of our methodology to study design based on M-estimators in both experimental and observational settings through a varied set of clinical examples.

stat.ME

Bayesian Design of Experiments in the Presence of Nuisance Parameters

Design of experiments has traditionally relied on the frequentist hypothesis testing framework where the optimal size of the experiment is specified as the minimum sample size that guarantees a required level of power. Sample size determination may be performed analytically when the test statistic has a known asymptotic sampling distribution and, therefore, the power function is available in analytic form. Bayesian methods have gained popularity in all stages of discovery, namely, design, analysis and decision making. Bayesian decision procedures rely on posterior summaries whose sampling distributions are commonly estimated via Monte Carlo simulations. In the design of scientific studies, the Bayesian approach incorporates uncertainty about the design value(s) instead of conditioning on a single value of the model parameter(s). Accounting for uncertainties in the design value(s) is particularly critical when the model includes nuisance parameters. In this manuscript, we propose methodology that utilizes the large-sample properties of the posterior distribution together with Bayesian additive regression trees (BART) to efficiently obtain the optimal sample size and decision criteria in fixed and adaptive designs. We introduce a fully Bayesian procedure that incorporates the uncertainty associated with the model parameters including the nuisance parameters at the design stage. The proposed approach significantly reduces the computational burden associated with Bayesian design and enables the wide adoption of Bayesian operating characteristics.

stat.ME

An Efficient Approach to Design Bayesian Platform Trials

Platform trials evaluate multiple experimental treatments against a common control group (and/or against each other), which often reduces the trial duration and sample size. Bayesian platform designs offer several practical advantages, including the flexible addition or removal of experimental arms using posterior probabilities and the incorporation of prior/external information. Regulatory agencies require that the operating characteristics of Bayesian designs are assessed by estimating the sampling distribution of posterior probabilities via Monte Carlo simulation. It is computationally intensive to repeat this simulation process for all design configurations considered, particularly for platform trials with complex interim decision procedures. In this paper, we propose an efficient method to assess operating characteristics and determine sample sizes as well as other design parameters for Bayesian platform trials. We prove theoretical results that allow us to model the joint sampling distribution of posterior probabilities across multiple endpoints and trial stages using simulations conducted at only two sample sizes. This work is motivated by design complexities in the SSTARLET trial, an ongoing Bayesian adaptive platform trial for tuberculosis preventive therapies (ClinicalTrials.gov ID: NCT06498414). Our proposed design method is not only computationally efficient but also capable of accommodating intricate, real-world trial constraints like those encountered in SSTARLET.

stat.ME

Group Sequential Design with Posterior and Posterior Predictive Probabilities

Group sequential designs drive innovation in clinical, industrial, and corporate settings. Early stopping for failure in sequential designs conserves experimental resources, whereas early stopping for success accelerates access to improved interventions. Bayesian decision procedures provide a formal and intuitive framework for early stopping using posterior and posterior predictive probabilities. Design parameters including decision thresholds and sample sizes are chosen to control the error probabilities associated with the sequential decision process. These choices are routinely made based on estimating the sampling distribution of posterior summaries via intensive Monte Carlo simulations for each sample size and design scenario considered. In this paper, we propose an efficient method to calibrate decision thresholds to pre-specified alpha- and beta-spending functions and determine minimum sample sizes for Bayesian group sequential designs. We prove theoretical results that enable posterior and posterior predictive probabilities to be modeled as a function of the sample size. Using these functions, we assess error probabilities at a range of sample sizes given simulations conducted at only two sample sizes. The effectiveness of our methodology is highlighted using several substantive examples.

stat.ME

Design of Bayesian Clinical Trials with Clustered Data

In the design of clinical trials, it is essential to assess the design operating characteristics (e.g., power and the type I error rate). Common practice for the evaluation of operating characteristics in Bayesian clinical trials relies on estimating the sampling distribution of posterior summaries via Monte Carlo simulation. It is computationally intensive to repeat this estimation process for each design configuration considered, particularly for clustered data that are analyzed using complex, high-dimensional models. In this paper, we propose an efficient method to assess operating characteristics and determine sample sizes for Bayesian trials with clustered data. We prove theoretical results that enable posterior probabilities to be modeled as a function of the number of clusters. Using these functions, we assess operating characteristics at a range of sample sizes given simulations conducted at only two cluster counts. These theoretical results are also leveraged to quantify the impact of simulation variability on our sample size recommendations. The applicability of our methodology is illustrated using an example cluster-randomized Bayesian clinical trial.

stat.ME

An Economical Approach to Design Posterior Analyses

To design Bayesian studies, criteria for the operating characteristics of posterior analyses - such as power and the type I error rate - are often assessed by estimating sampling distributions of posterior probabilities via simulation. In this paper, we propose an economical method to determine optimal sample sizes and decision criteria for such studies. Using our theoretical results that model posterior probabilities as a function of the sample size, we assess operating characteristics throughout the sample size space given simulations conducted at only two sample sizes. These theoretical results are used to construct bootstrap confidence intervals for the optimal sample sizes and decision criteria that reflect the stochastic nature of simulation-based design. We also repurpose the simulations conducted in our approach to efficiently investigate various sample sizes and decision criteria using contour plots. The broad applicability and wide impact of our methodology is illustrated using two clinical examples.

stat.ME

Design of Bayesian A/B Tests Controlling False Discovery Rates and Power

Businesses frequently run online controlled experiments (i.e., A/B tests) to learn about the effect of an intervention on multiple business metrics. To account for multiple hypothesis testing, multiple metrics are commonly aggregated into a single composite measure, losing valuable information, or strict family-wise error rate adjustments are imposed, leading to reduced power. In this paper, we propose an economical framework to design Bayesian A/B tests while controlling both power and the false discovery rate (FDR). Selecting optimal decision thresholds to control power and the FDR typically relies on intensive simulation at each sample size considered. Our framework efficiently recommends optimal sample sizes and decision thresholds for Bayesian A/B tests that satisfy criteria for the FDR and average power. Our approach is efficient because we leverage new theoretical results to obtain these recommendations using simulations conducted at only two sample sizes. Our methodology is illustrated using an example based on a real A/B test involving several metrics.

stat.ME

Bioequivalence Design with Sampling Distribution Segments

In bioequivalence design, power analyses dictate how much data must be collected to detect the absence of clinically important effects. Power is computed as a tail probability in the sampling distribution of the pertinent test statistics. When these test statistics cannot be constructed from pivotal quantities, their sampling distributions are approximated via repetitive, time-intensive computer simulation. We propose a novel simulation-based method to quickly approximate the power curve for many such bioequivalence tests by efficiently exploring segments (as opposed to the entirety) of the relevant sampling distributions. Despite not estimating the entire sampling distribution, this approach prompts unbiased sample size recommendations. We illustrate this method using two-group bioequivalence tests with unequal variances and overview its broader applicability in clinical design. All methods proposed in this work can be implemented using the developed dent package in R.

stat.ME

Posterior Ramifications of Prior Dependence Structures

Prior elicitation methods for Bayesian analyses transfigure prior information into quantifiable prior distributions. Recently, methods that leverage copulas have been proposed to accommodate more flexible dependence structures when eliciting multivariate priors. We show that the posterior cannot retain many of these flexible prior dependence structures in large-sample settings, and we emphasize that it is our responsibility as statisticians to communicate this to practitioners. We therefore overview objectives for prior specification that guide conversations between statisticians and practitioners to promote alignment between the flexibility in the prior dependence structure and the objectives for posterior analysis. Because correctly specifying the dependence structure a priori can be difficult, we consider how the choice of prior copula impacts the posterior distribution in terms of asymptotic convergence of the posterior mode. Our resulting recommendations clarify when it is useful to elicit intricate prior dependence structures and when it is not.

stat.ME

From Augmentation to Decomposition: A New Look at CUPED in 2023

Ten years ago, CUPED (Controlled Experiments Utilizing Pre-Experiment Data) mainstreamed the idea of variance reduction leveraging pre-experiment covariates. Since its introduction, it has been implemented, extended, and modernized by major online experimentation platforms. Many researchers and practitioners often interpret CUPED as a regression adjustment. In this article, we clarify its similarities and differences to regression adjustment and present CUPED as a more general augmentation framework which is closer to the spirit of the 2013 paper. We show that the augmentation view naturally leads to cleaner developments of variance reduction beyond simple average metrics, including ratio metrics and percentile metrics. Moreover, the augmentation view can go beyond using pre-experiment data and leverage in-experiment data, leading to significantly larger variance reduction. We further introduce metric decomposition using approximate null augmentation (ANA) as a mental model for in-experiment variance reduction. We study it under both a Bayesian framework and a frequentist optimal proxy metric framework. Metric decomposition arises naturally in conversion funnels, so this work has broad applicability.

stat.AP

Fast Power Curve Approximation for Posterior Analyses

Bayesian hypothesis tests leverage posterior probabilities, Bayes factors, or credible intervals to inform data-driven decision making. We propose a framework for power curve approximation with such hypothesis tests. We present a fast approach to explore the approximate sampling distribution of posterior probabilities when the conditions for the Bernstein-von Mises theorem are satisfied. We extend that approach to consider segments of such sampling distributions in a targeted manner for each sample size explored. These sampling distribution segments are used to construct power curves for various types of posterior analyses. Our resulting method for power curve approximation is orders of magnitude faster than conventional power curve estimation for Bayesian hypothesis tests. We also prove the consistency of the corresponding power estimates and sample size recommendations under certain conditions.

stat.ME

An Economical Approach to Design with Precision Criteria

Estimation frameworks for statistical inference are preferred to hypothesis testing when quantifying uncertainty and precise estimation are more valuable than binary decisions about statistical significance. Study design for estimation-based investigations often uses precision criteria to select sample sizes that control the length of interval estimates with respect to a sampling distribution. In this paper, we formally define the length probability distribution which characterizes the probability of obtaining a sufficiently narrow interval estimate as a function of the sample size. This distribution can then be used to determine the smallest sample size needed to ensure an interval estimate is sufficiently narrow. We prove that this distribution is approximately normal in large-sample settings for many data generation processes. However, this approximate normality may not hold for studies with moderate sample sizes, particularly when incorporating prior information or obtaining asymmetric interval estimates. Thus, we also propose an efficient simulation-based approach that estimates the sampling distribution of interval estimate lengths at only two sample sizes. Our methodology provides a unified framework for design with precision criteria in Bayesian and frequentist settings with parametric, semiparametric, and nonparametric inference. We illustrate the broad applicability of this framework with various examples.

stat.ME