arXiv ScienceSearch

arXiv subjects

Christopher Jackson

Publications and source records attributed to Christopher Jackson.

At least 19 recordsLinked to original sources

Stable and practical semi-Markov modelling of intermittently-observed data

Multi-state models are commonly used for intermittent observations of a state over time, but these are generally based on the Markov assumption, that transition rates are independent of the time spent in current and previous states. In a semi-Markov model, the rates can depend on the time spent in the current state, though available methods for this are either restricted to specific state structures or lack general software. This paper develops the approach of using a "phase-type" distribution for the sojourn time in a state, which expresses a semi-Markov model as a hidden Markov model, allowing the likelihood to be calculated easily for any state structure. While this approach involves a proliferation of latent parameters, identifiability can be improved by restricting the phase-type family to one which approximates a simpler distribution such as the Gamma or Weibull. This paper proposes a moment-matching method to obtain this approximation, making general semi-Markov models for intermittent data accessible in software for the first time. The method is implemented in a new R package, "msmbayes", which implements Bayesian or maximum likelihood estimation for multi-state models with general state structures and covariates. The software is tested using simulation-based calibration, and an application to cognitive function decline illustrates the use of the method in a typical modelling workflow.

stat.ME

Voice EHR: Introducing Multimodal Audio Data for Health

Artificial intelligence (AI) models trained on audio data may have the potential to rapidly perform clinical tasks, enhancing medical decision-making and potentially improving outcomes through early detection. Existing technologies depend on limited datasets collected with expensive recording equipment in high-income countries, which challenges deployment in resource-constrained, high-volume settings where audio data may have a profound impact on health equity. This report introduces a novel data type and a corresponding collection system that captures health data through guided questions using only a mobile/web application. The app facilitates the collection of an audio electronic health record (Voice EHR) which may contain complex biomarkers of health from conventional voice/respiratory features, speech patterns, and spoken language with semantic meaning and longitudinal context, potentially compensating for the typical limitations of unimodal clinical datasets. This report presents the application used for data collection, initial experiments on data quality, and case studies which demonstrate the potential of voice EHR to advance the scalability/diversity of audio AI.

cs.SD

Bayesian blockwise inference for joint models of longitudinal and multistate processes

Joint models (JM) for longitudinal and survival data have gained increasing interest and found applications in a wide range of clinical and biomedical settings. These models facilitate the understanding of the relationship between outcomes and enable individualized predictions. In many applications, more complex event processes arise, necessitating joint longitudinal and multistate models. However, their practical application can be hindered by computational challenges due to increased model complexity and large sample sizes. Motivated by a longitudinal multimorbidity analysis of large UK health records, we have developed a scalable Bayesian methodology for such joint multistate models that is capable of handling complex event processes and large datasets, with straightforward implementation. We propose two blockwise inference approaches for different inferential purposes based on different levels of decomposition of the multistate processes. These approaches leverage parallel computing, ease the specification of different models for different transitions, and model/variable selection can be performed within a Bayesian framework using Bayesian leave-one-out cross-validation. Using a simulation study, we show that the proposed approaches achieve satisfactory performance regarding posterior point and interval estimation, with notable gains in sampling efficiency compared to the standard estimation strategy. We illustrate our approaches using a large UK electronic health record dataset where we analysed the coevolution of routinely measured systolic blood pressure (SBP) and the progression of multimorbidity, defined as the combinations of three chronic conditions. Our analysis identified distinct association structures between SBP and different disease transitions.

stat.ME

survextrap: a package for flexible and transparent survival extrapolation

Health policy decisions are often informed by estimates of long-term survival based primarily on short-term data. A range of methods are available to include longer-term information, but there has previously been no comprehensive and accessible tool for implementing these. This paper introduces a novel model and software package for parametric survival modelling of individual-level, right-censored data, optionally combined with summary survival data on one or more time periods. It could be used to estimate long-term survival based on short-term data from a clinical trial, combined with longer-term disease registry or population data, or elicited judgements. All data sources are represented jointly in a Bayesian model. The hazard is modelled as an M-spline function, which can represent potential changes in the hazard trajectory at any time. Through Bayesian estimation, the model automatically adapts to fit the available data, and acknowledges uncertainty where the data are weak. Therefore long-term estimates are only confident if there are strong long-term data, and inferences do not rely on extrapolating parametric functions learned from short-term data. The effects of treatment or other explanatory variables can be estimated through proportional hazards or with a flexible non-proportional hazards model. Some commonly-used mechanisms for survival can also be assumed: cure models, additive hazards models with known background mortality, and models where the effect of a treatment wanes over time. All of these features are provided for the first time in an R package, $\texttt{survextrap}$, in which models can be fitted using standard R survival modelling syntax. This paper explains the model, and demonstrates the use of the package to fit a range of models to common forms of survival data used in health technology assessments.

stat.ME

SNOWMASS Neutrino Frontier NF10 Topical Group Report: Neutrino Detectors

We discuss here future neutrino detectors with physics goals ranging from the eV to the EeV scale. The focus is on future enabling technologies for such detectors, rather than existing detectors or those under construction. The report includes methodologies across the broad spectrum of neutrino physics: liquid noble and other cryogenic detectors, includin LAr and LXe TPCs; photon-based detectors including technologies enabling hybrid Cherenkov/scintillation detectors; low-threshold detectors which use a wide variety of technologies to probe physics like coherent neutrino-nucleus scattering or detection of cosmic background neutrinos; and ultra-high energy detectors including optical and radio detectors, as well as tracking detectors for use at the forward physics facility of the LHC

hep-ex

A comparison of two frameworks for multi-state modelling, applied to outcomes after hospital admissions with COVID-19

We compare two multi-state modelling frameworks that can be used to represent dates of events following hospital admission for people infected during an epidemic. The methods are applied to data from people admitted to hospital with COVID-19, to estimate the probability of admission to ICU, the probability of death in hospital for patients before and after ICU admission, the lengths of stay in hospital, and how all these vary with age and gender. One modelling framework is based on defining transition-specific hazard functions for competing risks. A less commonly used framework defines partially-latent subpopulations who will experience each subsequent event, and uses a mixture model to estimate the probability that an individual will experience each event, and the distribution of the time to the event given that it occurs. We compare the advantages and disadvantages of these two frameworks, in the context of the COVID-19 example. The issues include the interpretation of the model parameters, the computational efficiency of estimating the quantities of interest, implementation in software and assessing goodness of fit. In the example, we find that some groups appear to be at very low risk of some events, in particular ICU admission, and these are best represented by using "cure-rate" models to define transition-specific hazards. We provide general-purpose software to implement all the models we describe in the "flexsurv" R package, which allows arbitrarily-flexible distributions to be used to represent the cause-specific hazards or times to events.

stat.ME

A Facility for Low-Radioactivity Underground Argon

The DarkSide-50 experiment demonstrated the ability to extract and purify argon from deep underground sources and showed that the concentration of $^{39}$Ar in that argon was greatly reduced from the level found in argon derived from the atmosphere. That discovery broadened the physics reach of argon-based detector and created a demand for low-radioactivity underground argon (UAr) in high-energy physics, nuclear physics, and in environmental and allied sciences. The Global Argon Dark Matter Collaboration (GADMC) is preparing to produce UAr for DarkSide-20k, but a general UAr supply for the community does not exist. With the proper resources, those plants could be operated as a facility to supply UAr for most of the experiments after the DarkSide 20k production. However, if the current source becomes unavailable, or UAr masses greater than what is available from the current source is needed, then a new source must be found. To find a new source will require understanding the production of the radioactive argon isotopes underground in a gas field, and the ability to measure $^{37}$Ar, $^{39}$Ar, and $^{42}$Ar to ultra-low levels. The operation of a facility creates a need for ancillary systems to monitor for $^{37}$Ar, $^{39}$Ar, or $^{42}$Ar infiltration either directly or indirectly, which can also be used to vet the $^{37}$Ar, $^{39}$Ar, and $^{42}$Ar levels in a new UAr source, but requires the ability to separate UAr from the matrix well gas. Finding methods to work with industry to find gas streams enriched in UAr, or to commercialize a UAr facility, are highly desirable.

physics.ins-det

Bayesian multistate modelling of incomplete chronic disease burden data

A widely-used model for determining the long-term health impacts of public health interventions, often called a "multistate lifetable", requires estimates of incidence, case fatality, and sometimes also remission rates, for multiple diseases by age and gender. Generally, direct data on both incidence and case fatality are not available in every disease and setting. For example, we may know population mortality and prevalence rather than case fatality and incidence. This paper presents Bayesian continuous-time multistate models for estimating transition rates between disease states based on incomplete data. This builds on previous methods by using a formal statistical model with transparent data-generating assumptions, while providing accessible software as an R package. Rates for people of different ages and areas can be related flexibly through splines or hierarchical models. Previous methods are also extended to allow age-specific trends through calendar time. The model is used to estimate case fatality for multiple diseases in the city regions of England, based on incidence, prevalence and mortality data from the Global Burden of Disease study. The estimates can be used to inform health impact models relating to those diseases and areas. Different assumptions about rates are compared, and we check the influence of different data sources.

stat.AP

Dark Matter Detection Capabilities of a Large Multipurpose Liquid Argon Time Projection Chamber

Liquid Argon Time Projection Chambers are planned to comprise a central role in the future of the U.S. High Energy Physics neutrino program. In particular, this detector technology will form the basis for the 40 kton Deep Underground Neutrino Experiment (DUNE). In this paper we take as a starting point the dual phase far detector design proposed by the DUNE experiment and ask what changes are necessary to allow one of the four 10 kt modules to be sensitive to heavy Weakly Interacting Massive Particle (WIMP) dark matter. We show that with control over backgrounds and the use of low radioactivity argon, which may be commercially available on that timescale, along with a significant increase in light detection, one DUNE-like module gives a competitive WIMP detection sensitivity, particularly above a dark matter mass of 100 GeV.

physics.ins-det

The Game of Cycles

The Game of Cycles, introduced by Su (2020), is played on a simple connected planar graph together with its bounded cells, and players take turns marking edges with arrows according to a sink-source rule that gives the game a topological flavor. The object of the game is to produce a cycle cell---a cell surrounded by arrows all cycling in one direction---or to make the last possible move. We analyze the two-player game for various classes of graphs and determine who has a winning strategy. We also establish a topological property of the game: that a board with every edge marked must have a cycle cell.

math.CO

Supervised learning of photoelectron counting in scintillator-based dark matter experiments

Many scintillator based detectors employ a set of photomultiplier tubes (PMT) to observe the scintillation light from potential signal and background events. It is important to be able to count the number of photoelectrons (PE) in the pulses observed in the PMTs, because the position and energy reconstruction of the events is directly related to how well the spatial distribution of the PEs in the PMTs as well as their total number might be measured. This task is challenging for fast scintillators, since the PEs often overlap each other in time. Standard Bayesian statistics methods are often used and this has been the method employed in analyzing the data from liquid argon experiments such as MiniCLEAN and DEAP. In this work, we show that for the MiniCLEAN detector it is possible to use a multi-layer perceptron to learn the number of PEs using only raw pulse features with better accuracy and precision than existing methods. This can even help to perform position reconstruction with better accuracy and precision, at least in some generic cases.

physics.ins-det

Calculating the Expected Value of Sample Information in Practice: Considerations from Three Case Studies

Investing efficiently in future research to improve policy decisions is an important goal. Expected Value of Sample Information (EVSI) can be used to select the specific design and sample size of a proposed study by assessing the benefit of a range of different studies. Estimating EVSI with the standard nested Monte Carlo algorithm has a notoriously high computational burden, especially when using a complex decision model or when optimizing over study sample sizes and designs. Therefore, a number of more efficient EVSI approximation methods have been developed. However, these approximation methods have not been compared and therefore their relative advantages and disadvantages are not clear. A consortium of EVSI researchers, including the developers of several approximation methods, compared four EVSI methods using three previously published health economic models. The examples were chosen to represent a range of real-world contexts, including situations with multiple study outcomes, missing data, and data from an observational rather than a randomized study. The computational speed and accuracy of each method were compared, and the relative advantages and implementation challenges of the methods were highlighted. In each example, the approximation methods took minutes or hours to achieve reasonably accurate EVSI estimates, whereas the traditional Monte Carlo method took weeks. Specific methods are particularly suited to problems where we wish to compare multiple proposed sample sizes, when the proposed sample size is large, or when the health economic model is computationally expensive. All the evaluated methods gave estimates similar to those given by traditional Monte Carlo, suggesting that EVSI can now be efficiently computed with confidence in realistic examples.

stat.AP

A guide to Value of Information methods for prioritising research in health impact modelling

Health impact simulation models are used to predict how a proposed intervention or scenario will affect public health outcomes, based on available data and knowledge of the process. The outputs of these models are uncertain due to uncertainty in the structure and inputs to the model. In order to assess the extent of uncertainty in the outcome we must quantify all potentially relevant uncertainties. Then to reduce uncertainty we should obtain and analyse new data, but it may be unclear which parts of the model would benefit from such extra research. This paper presents methods for uncertainty quantification and research prioritisation in health impact models based on Value of Information (VoI) analysis. Specifically, we 1. discuss statistical methods for quantifying uncertainty in this type of model, given the typical kinds of data that are available, which are often weaker than the ideal data that are desired; 2. show how the expected value of partial perfect information (EVPPI) can be calculated to compare how uncertainty in each model parameter influences uncertainty in the output; 3. show how research time can be prioritised efficiently, in the light of which components contribute most to outcome uncertainty. The same methods can be used whether the purpose of the model is to estimate quantities of interest to a policy maker, or to explicitly decide between policies. We demonstrate how these methods might be used in a model of the impact of air pollution on health outcomes.

stat.ME

The Low-Radioactivity Underground Argon Workshop: A workshop synopsis

In response to the growing need for low-radioactivity argon, community experts and interested parties came together for a 2-day workshop to discuss the worldwide low-radioactivity argon needs and the challenges associated with its production and characterization. Several topics were covered: experimental needs and requirements for low-radioactivity argon, the sources of low-radioactivity argon and its production, how long-lived argon radionuclides are created in nature, measuring argon radionuclides, and other applicable topics. The Low-Radioactivity Underground Argon (LRUA) workshop took place on March 19-20, 2018 at Pacific Northwest National Laboratory in Richland Washington, USA. This paper is a synopsis of the workshop with the associated abstracts from the talks.

physics.ins-det

Value of Information: Sensitivity Analysis and Research Design in Bayesian Evidence Synthesis

Suppose we have a Bayesian model which combines evidence from several different sources. We want to know which model parameters most affect the estimate or decision from the model, or which of the parameter uncertainties drive the decision uncertainty. Furthermore we want to prioritise what further data should be collected. These questions can be addressed by Value of Information (VoI) analysis, in which we estimate expected reductions in loss from learning specific parameters or collecting data of a given design. We describe the theory and practice of VoI for Bayesian evidence synthesis, using and extending ideas from health economics, computer modelling and Bayesian design. The methods are general to a range of decision problems including point estimation and choices between discrete actions. We apply them to a model for estimating prevalence of HIV infection, combining indirect information from several surveys, registers and expert beliefs. This analysis shows which parameters contribute most of the uncertainty about each prevalence estimate, and provides the expected improvements in precision from collecting specific amounts of additional data.

stat.AP

Non-holonomic tomography II: Detecting correlations in multiqudit systems

In the context of quantum tomography, quantities called a partial determinants\cite{jackson2015detecting} were recently introduced. PDs (partial determinants) are explicit functions of the collected data which are sensitive to the presence of state-preparation-and-measurement (SPAM) correlations. In this paper, we demonstrate further applications of the PD and its generalizations. In particular we construct methods for detecting various types of SPAM correlation in multiqudit systems | e.g. measurement-measurement correlations. The relationship between the PDs of each method and the correlations they are sensitive to is topological. We give a complete classification scheme for all such methods but focus on the explicit details of only the most scalable methods, for which the number of settings scale as $\mathcal{O}(d^4)$. This paper is the second of a two part series where the first paper[2] is about theoretical perspectives of the PD and its interpretation as a holonomy.

quant-ph

Non-holonomic tomography I: The Born rule as a connection between experiments

In the context of quantum tomography, we recently introduced a quantity called a partial determinant \cite{jackson2015detecting}. PDs (partial determinants) are explicit functions of the collected data which are sensitive to the presence of state-preparation-and-measurment (SPAM) correlated errors. As such, PDs bypass the need to estimate state-preparation or measurement parameters individually. In the present work, we suggest a theoretical perspective for the PD. We show that the PD is a holonomy and that the notions of state, measurement, and tomography can be generalized to non-holonomic constraints. To illustrate and clarify these abstract concepts, direct analogies are made to parallel transport, thermodynamics, and gauge field theory. This paper is the first of a two part series where the second paper [2] is about scalable applications of the PD to multiqudit systems.

quant-ph

The Born rule as a parallel transport equation: Detecting multiqudit state-preparation and measurement correlations

In the context of quantum tomography, we recently introduced a quantity called a partial determinant \cite{jackson2015detecting}. PDs (partial determinants) are explicit functions of the collected data which are sensitive to the presence of state-preparation-and-measurement (SPAM) correlations. Importantly, this is done without any need to estimate state-preparation or measurement parameters. In the present work, we wish to better explain our theoretical perspective behind the PD. Further, we would like to demonstrate that there is an overwhelming variety of applications and generalizations of the PD. In particular we will construct methods for detecting SPAM correlations in multiqudit systems. The relationship between the PDs of each method and the correlations they are sensitive to is topological. We give a classification of all such methods but focus on explicitly detailing only the most scalable methods, $\mathcal{O}(d^4)$.

quant-ph