arXiv Science⌕ Search

arXiv subjects

Lauren A. Castro

Publications and source records attributed to Lauren A. Castro.

4 recordsLinked to original sources

Tree statistics improve neural simulation-based inference of epidemiological parameters over case counts alone

During an infectious disease outbreak, accurate estimates of epidemiological parameters such as the reproduction number and infection duration can inform the public health response. Recently, neural simulation-based inference (NSBI) methods have become popular for estimating these parameters due to lower computational costs and more input data flexibility than traditional methods. Input data modalities that have been evaluated independently in an NSBI framework include case counts and transmission trees. Here, we evaluate whether combining information from these two data types outperforms using information from each individually with NSBI. To do so, we simulated stochastic susceptible-infectious-recovered (SIR) outbreaks across a range of basic reproduction numbers (R0) and average infection durations (D) to generate training, validation, and testing datasets. We then calculated, for each simulated outbreak, summary statistics from case count trajectories and from transmission trees. Using these features, individually and in combination, we performed a type of NSBI called neural ratio estimation (NRE) to estimate R0 and D. We also calculated maximum likelihood estimates of the parameters by fitting a deterministic SIR model using the case count trajectory. All three NRE models outperformed the baseline method. Furthermore, incorporating tree summary statistics improved the estimation of both R0 and D over case count summaries alone. This suggests that trees provide valuable additional information for estimating epidemiological parameters above and beyond case counts alone.

stat.AP↗

Leveraging Synthetic and Genetic Data to Improve Epidemic Forecasting

Forecasting infectious disease outbreaks is hard. Forecasting emerging infectious diseases with limited historical data is even harder. In this paper, we investigate ways to improve emerging infectious disease forecasting under operational constraints. Specifically, we explore two options likely to be available near the start of an emerging disease outbreak: synthetic data and genetic information. For this investigation, we conducted an experiment where we trained deep learning models on different combinations of real and synthetic data, both with and without genetic information, to explore how these models compare when forecasting COVID-19 cases for US states. All models are developed with an eye towards forecasting the next pandemic. We find that models trained with synthetic data have better forecast accuracy than models trained on real data alone, and models that use genetic variants have better forecast accuracy compared to those that do not. All models outperformed a baseline persistence model (a feat only accomplished by 7 out of 22 real-time COVID-19 cases forecasting models as reported in [38]) and multiple models outperformed the COVIDHub-4_week_ensemble. This paper demonstrates the value of these underutilized sources of information and provides a blueprint for forecasting future pandemics.

stat.AP↗

Ionospheric Observations from the ISS: Overcoming Noise Challenges in Signal Extraction

The Electric Propulsion Electrostatic Analyzer Experiment (ÈPÈE) is a compact ion energy bandpass filter deployed on the International Space Station (ISS) in March 2023 and providing continuous measurements through April 2024. This period coincides with the Solar Cycle 25 maximum, capturing unique observations of solar activity extremes in the mid- to low-latitude regions of the topside ionosphere. From these in situ spectra we derive plasma parameters that inform space-weather impacts on satellite navigation and radio communication. We present a statistical processing pipeline for ÈPÈE that (i) estimates the instrument noise floor, (ii) accounts for irregular temporal sampling, and (iii) extracts ionospheric signals. Rather than discarding noisy data, the method learns a baseline noise model and fits the measurement surface using a scaled Vecchia Gaussian process approximation, recovering values typically rejected by thresholding. The resulting products increase data coverage and enable noise-assisted monitoring of ionospheric variability.

physics.space-ph↗

Mapping Incidence and Prevalence Peak Data for SIR Forecasting Applications

Infectious disease modeling and forecasting have played a key role in helping assess and respond to epidemics and pandemics. Recent work has leveraged data on disease peak infection and peak hospital incidence to fit compartmental models for the purpose of forecasting and describing the dynamics of a disease outbreak. Incorporating these data can greatly stabilize a compartmental model fit on early observations, where slight perturbations in the data may lead to model fits that project wildly unrealistic peak infection. We introduce a new method for incorporating historic data on the value and time of peak incidence of hospitalization into the fit for a Susceptible-Infectious-Recovered (SIR) model by formulating the relationship between an SIR model's starting parameters and peak incidence as a system of two equations that can be solved computationally. This approach is assessed for practicality in terms of accuracy and speed of computation via simulation. To exhibit the modeling potential, we update the Dirichlet-Beta State Space modeling framework to use hospital incidence data, as this framework was previously formulated to incorporate only data on total infections.

stat.ME↗