arXiv ScienceSearch

arXiv subjects

Tyrel Stokes

Publications and source records attributed to Tyrel Stokes.

7 recordsLinked to original sources

Domain Adaptation Under MNAR Missingness

Current domain adaptation methods under missingness shift are restricted to Missing At Random (MAR) missingness mechanisms. However, in many real-world examples, the MAR assumption may be too restrictive. When covariates are Missing Not At Random (MNAR) in both source and target data, the common covariate shift solutions, including importance weighting, are not directly applicable. We show that under reasonable assumptions, the problem of MNAR missingness shift can be reduced to an imputation problem. This allows us to leverage recent methodological developments in both the traditional statistics and machine/deep-learning literature for MNAR imputation to develop a novel domain adaptation procedure for MNAR missingness shift. We further show that our proposed procedure can be extended to handle simultaneous MNAR missingness and covariate shifts. We apply our procedure to Electronic Health Record (EHR) data from two hospitals in south and northeast regions of the US. In this setting we expect different hospital networks and regions to serve different populations and to have different procedures, practices, and software for inputting and recording data, causing simultaneous missingness and covariate shifts.

stat.ME

A generative approach to frame-level multi-competitor races

Multi-competitor races often feature complicated within-race strategies that are difficult to capture when training data on race outcome level data. Further, models which do not account for such strategic effects may suffer from confounded inferences and predictions. In this work we develop a general generative model for multi-competitor races which allows analysts to explicitly model certain strategic effects such as changing lanes or drafting and separate these impacts from competitor ability. The generative model allows one to simulate full races from any real or created starting position which opens new avenues for attributing value to within-race actions and to perform counter-factual analyses. This methodology is sufficiently general to apply to any track based multi-competitor races where both tracking data is available and competitor movement is well described by simultaneous forward and lateral movements. We apply this methodology to one-mile horse races using data provided by the New York Racing Association (NYRA) and the New York Thoroughbred Horsemen's Association (NYTHA) for the Big Data Derby 2022 Kaggle Competition. This data features granular tracking data for all horses at the frame-level (occurring at approximately 4hz). We demonstrate how this model can yield new inferences, such as the estimation of horse-specific speed profiles which vary over phases of the race, and examples of posterior predictive counterfactual simulations to answer questions of interest such as starting lane impacts on race outcomes.

stat.ME

Simulation Experiments as a Causal Problem

Simulation methods are among the most ubiquitous methodological tools in statistical science. In particular, statisticians often is simulation to explore properties of statistical functionals in models for which developed statistical theory is insufficient or to assess finite sample properties of theoretical results. We show that the design of simulation experiments can be viewed from the perspective of causal intervention on a data generating mechanism. We then demonstrate the use of causal tools and frameworks in this context. Our perspective is agnostic to the particular domain of the simulation experiment which increases the potential impact of our proposed approach. In this paper, we consider two illustrative examples. First, we re-examine a predictive machine learning example from a popular textbook designed to assess the relationship between mean function complexity and the mean-squared error. Second, we discuss a traditional causal inference method problem, simulating the effect of unmeasured confounding on estimation, specifically to illustrate bias amplification. In both cases, applying causal principles and using graphical models with parameters and distributions as nodes in the spirit of influence diagrams can 1) make precise which estimand the simulation targets , 2) suggest modifications to better attain the simulation goals, and 3) provide scaffolding to discuss performance criteria for a particular simulation design.

stat.ME

Injury risk increases minimally over a large range of changes in activity level in children

Background: Limited research exists on the association between changes in physical activity levels and injury in children. Objective: To assess how well different variations of the acute:chronic workload ratio (ACWR), a measure of change in activity, predict injury in children. Methods: We conducted a prospective cohort study using data from 1670 Danish schoolchildren measured over 5.5 years (2008 to 2014). Coupled 4-week, uncoupled 4-week, and uncoupled 5-week ACWRs were calculated using activity frequency in the past week as the acute load (numerator), and average weekly activity frequency in the past 4 or 5 weeks as the chronic load (denominator). We modelled the relationship between different ACWR variations and injury using generalized linear and generalized additive models, with and without accounting for repeated measures. Results: The prognostic relationship between the ACWR and injury risk was best represented using a generalized additive mixed model for the uncoupled 5-week ACWR. It predicted an injury risk of ~3% for ACWRs between 0.8 (activity level decreased by 20%) and 1.5 (activity level increased by 50%). When activity decreased by more than 20% (ACWR< 0.8), injury risk was lower (minimum of 1.5% at ACWR=0). When activity increased by more than 50% (ACWR > 1.5), injury risk was higher (maximum of 6% at ACWR = 5). Girls were at significantly higher risk of injury than boys. Conclusion: Increases in physical activity in children are associated with much lower injury risks compared to previous results in adults.

q-bio.QM

Implementing multiple imputation for missing data in longitudinal studies when models are not feasible: A tutorial on the random hot deck approach

Objective: Researchers often use model-based multiple imputation to handle missing at random data to minimize bias while making the best use of all available data. However, there are sometimes constraints within the data that make model-based imputation difficult and may result in implausible values. In these contexts, we describe how to use random hot deck imputation to allow for plausible multiple imputation in longitudinal studies. Study Design and Setting: We illustrate random hot deck multiple imputation using The Childhood Health, Activity, and Motor Performance School Study Denmark (CHAMPS-DK), a prospective cohort study that measured weekly sports participation for 1700 Danish schoolchildren. We matched records with missing data to several observed records, generated probabilities for matched records using observed data, and sampled from these records based on the probability of each occurring. Because imputed values are generated randomly, multiple complete datasets can be created and analyzed similar to model-based multiple imputation. Conclusion: Multiple imputation using random hot deck imputation is an alternative method when model-based approaches are infeasible, specifically where there are constraints within and between covariates.

stat.ME

Causal Simulation Experiments: Lessons from Bias Amplification

Recent theoretical work in causal inference has explored an important class of variables which, when conditioned on, may further amplify existing unmeasured confounding bias (bias amplification). Despite this theoretical work, existing simulations of bias amplification in clinical settings have suggested bias amplification may not be as important in many practical cases as suggested in the theoretical literature.We resolve this tension by using tools from the semi-parametric regression literature leading to a general characterization in terms of the geometry of OLS estimators which allows us to extend current results to a larger class of DAGs, functional forms, and distributional assumptions. We further use these results to understand the limitations of current simulation approaches and to propose a new framework for performing causal simulation experiments to compare estimators. We then evaluate the challenges and benefits of extending this simulation approach to the context of a real clinical data set with a binary treatment, laying the groundwork for a principled approach to sensitivity analysis for bias amplification in the presence of unmeasured confounding.

stat.ME

The acute:chronic workload ratio: challenges and prospects for improvement

Injuries occur when an athlete performs a greater amount of activity (workload) than what their body can absorb. To maximize the positive effects of training while avoiding injuries, athletes and coaches need to determine safe workload levels. The International Olympic Committee has recommended using the acute:chronic workload ratio (ACRatio) to monitor injury risk, and has provided thresholds to minimize risk. However, there are several limitations to the ACRatio which may impact the validity of current recommendations. In this review, we discuss previously published and novel challenges with the ACRatio, and possible strategies to address them. These challenges include 1) formulating the ACRatio as a proportion rather than a measure of change, 2) its use of unweighted averages to measure activity loads, 3) inapplicability of the ACRatio to sports where athletes taper their activity, 4) discretization of the ACRatio prior to model selection, 5) the establishment of the model using sparse data, 6) potential bias in the ACRatio of injured athletes, 7) unmeasured confounding, and 8) application of the ACRatio to subsequent injuries.

stat.AP