arXiv ScienceSearch

arXiv subjects

Thomas House

Publications and source records attributed to Thomas House.

At least 19 recordsLinked to original sources

Symptom clusters in Long COVID in the UK: prospective community-based cohort study using unsupervised machine learning

Long COVID is a condition usually defined by persisting symptoms following infection by the SARS-CoV-2 virus beyond the acute phase of infection. The condition has a significant impact on healthcare systems, the economy, and the individuals living with it. Due to the diverse and extensive symptomatology of long COVID, symptom co-occurrence tracking can be used to capture patient experiences and improve understanding, diagnosis, and management. Here, we leverage the UK Office for National Statistics (ONS) COVID-19 Infection Survey (CIS). The CIS was run between April 2020 and March 2023. On February 3, 2021, the ONS launched a new CIS question evaluating self-reported symptom persistence of 23 long COVID symptoms post self-reported COVID-19 infection. We use Jaccard symptom-by symptom distance matrices, derived from the binary survey responses of the presence of each symptom. Three methods are then used to visualise symptom co-occurrence: heatmaps, non-metric multidimensional scaling, and agglomerative hierarchical clustering with complete linkage. We split our analysis into two parts, first looking at all long COVID survey responses (n = 207,319) and then looking at responses stratified by time-since-onset of long COVID up to 24 months (n = 30,224 at zero months-since-onset). We find higher symptom co-occurrence in prolonged long COVID. We also find clusters of neurological/systemic symptoms, gastrointestinal symptoms, and respiratory symptoms, with separability in symptom co-occurrence by organ system becoming less pronounced as time-since-onset of long COVID increases. Shortness of breath, weakness/tiredness, and muscle ache present as core symptoms long COVID. We demonstrate that the symptom experience of long COVID evolves from early stages to late stages with higher symptom burden and multi-systemic presentation. We also find three possible phenotypes of early long COVID.

stat.AP

A framework for combined epidemiological-genomic inference to improve estimation of household model parameters

Models incorporating household structure, with different rates of transmission within and between households, are widely used in infectious disease epidemiology. These models can be calibrated using final-size data in which transmission ordering is ignored because it does not affect the distribution of final outbreak sizes. In particular, many distinct transmission histories produce identical final epidemiological outcomes, making it difficult to distinguish internal (within-household) from external (between-household) transmission and limiting parameter identifiability. Here, we develop a continuous-time Markov chain formulation for household transmission dynamics in which the model state space is expanded to include transmission graphs describing infection direction and order, with idealised pathogen genomic data used to identify the transmission histories compatible with observations. We conduct simulation studies which show that incorporating genetic information substantially concentrates the regions of high likelihood compared with models based on epidemiological data alone. In particular, genomic data reduces the dependence between internal and external transmission parameters, removing the characteristic ridge associated with their weak identifiability. These results demonstrate that graph-resolved household models enable improved transmission inference while maintaining analytical and computational tractability.

q-bio.PE

How complex behavioural contagion can prevent infectious diseases from becoming endemic

Infectious disease transmission in human populations has a complex two-way interaction with changes in host behaviour. It is increasingly recognised that incorporating adaptive behavioural change into epidemic models is important for improving understanding of infectious disease dynamics and developing policy-relevant modelling tools. An important aspect of behavioural dynamics is social contagion, where people tend to adopt behaviours exhibited by others around them. In a simple behavioural contagion model, the behaviour uptake rate increases linearly with the number of contacts who have adopted a given behaviour. Here, we explore an epidemic model with complex behavioural contagion, where the behaviour uptake rate is a nonlinear function of the number of behaving contacts. We identify key bifurcation parameters of the model, which include the basic reproduction number $R_0$, the strength of the behavioural effect on disease transmission, and the speed of behaviour uptake relative to behaviour abandonment. We show that, in some regions of parameter space, the model has multiple disease-free equilibria. In this situation, the occurrence of an epidemic in a population with an initially low level of behaviour practice can trigger a self-sustaining increase in behaviour, which then causes the disease to be eliminated. In some cases, while moderate values of $R_0$ lead to the disease becoming endemic, higher values of $R_0$ may lead to behaviour-driven disease elimination. We demonstrate that this mechanism of epidemic-triggered uptake of behaviour leading to disease elimination can occur in the presence and absence of temporary post-infection immunity.

q-bio.PE

Bayesian Inference for Epidemic Final Size Datasets with Hidden Underlying Household Structure

Households represent a key unit of interest in infectious disease epidemiology, in both empirical studies and mathematical modelling. The within-household transmission potential of a disease is often summarised by a secondary attack ratio (SAR). Despite its widespread use, the SAR depends on the household size distribution (HHSD) seen during the study period, making it difficult to generalise to new contexts. Extending estimates of transmission potential to new populations instead requires estimates of person-to-person transmission rates which can be convoluted with data on population structure to parametrise mechanistic transmission models. In this study we present a new Bayesian inference method which uses an MCMC algorithm to infer the transmission intensity by imputing the unreported household structure underlying the epidemic. This method can be run on household epidemiological data reported at varying levels of resolution. For synthetic data from a realistic underlying HHSD, we were able to achieve over 90% coverage in our estimates of transmission rate consistently. We were also able to consistently achieve over 90% coverage for data generated with a pathological underlying HHSD, given strong information about the HHSD. Using an existing dataset which recorded micro-scale household epidemiological outcomes during the COVID-19 pandemic, we show that stratifying observed SARs by household size substantially reduces the uncertainty in estimates. Our findings suggest that researchers conducting household epidemiological studies can improve the utility of results for infectious disease modellers by reporting household-stratified estimates. These results aim to encourage the reporting of higher resolution outputs in epidemiological field work as, in the absence of strong priors, transmission parameters were not easily identifiable from low resolution datasets, which are often reported.

stat.AP

Inference for Within- and Between-Partnership Transmission Rates for HIV Infection

HIV transmission within serodiscordant couples remains a significant public health challenge, particularly in sub-Saharan Africa. Estimating the rate of such infection, alongside the rates of introduction of infection from outside the partnership, is a special case of the more general epidemiological challenge of inferring intensities of within- and between-group intensities of transmission. This study presents a stochastic susceptible-infected (SI) pair model for estimating key epidemiological parameters governing HIV transmission within and between couples, which we further extend to account for gender-specific differences in infection dynamics. Using a likelihood-based inference approach, we estimate transmission parameters and associated uncertainty from observed data. These values can be used to inform infection prevention strategies for HIV, and the methodology proposed can be generalised to other epidemiological settings.

stat.AP

A Behaviour and Disease Model of Testing and Isolation

There has been interest in the interactions between infectious disease dynamics and behaviour for most of the history of mathematical epidemiology. This has included consideration of which mathematical models best capture each phenomenon, as well as their interaction, but typically in a manner that is agnostic to the exact behaviour in question. Here, we investigate interacting behaviour and disease dynamics specifically related to decisions around testing and isolation. To carry out our investigation we extend an existing "behaviour and disease" (BaD) model by incorporating the dynamics of symptomatic testing and isolation, including the influence of positive tests on perception of infection risk. We provide a dynamical systems analysis of the ordinary differential equations that define this model, providing theoretical results on its behaviour early in a new outbreak (particularly its basic reproduction number) and endemicity of the system (its steady states and associated stability criteria). We then supplement these findings with a numerical analysis to inform how temporal and cumulative outbreak metrics depend on the model parameter values for epidemic and endemic regimes. We observe novel model outputs such as epidemics that have more observed cases detected through increased testing, but are less objectively severe in terms of total number of infections.

q-bio.PE

Estimating the duration of RT-PCR positivity for SARS-CoV-2 from doubly interval censored data with undetected infections

Monitoring the incidence of new infections during a pandemic is critical for an effective public health response. General population prevalence surveys for SARS-CoV-2 can provide high-quality data to estimate incidence. However, estimation relies on understanding the distribution of the duration that infections remain detectable. This study addresses this need using data from the Coronavirus Infection Survey (CIS), a long-term, longitudinal, general population survey conducted in the UK. Analyzing these data presents unique challenges, such as doubly interval censoring, undetected infections, and false negatives. We propose a Bayesian nonparametric survival analysis approach, estimating a discrete-time distribution of durations and integrating prior information derived from a complementary study. Our methodology is validated through a simulation study, including its resilience to model misspecification, and then applied to the CIS dataset. This results in the first estimate of the full duration distribution in a general population, as well as methodology that could be transferred to new contexts.

stat.ME

Investigating the Trade-off between Infections and Social Interactions Using a Compact Model of Endemic Infections on Networks

In many epidemiological and ecological contexts, there is a trade-off between infections and interactions. This arises because the links between individuals capable of spreading infections are also often associated with beneficial activities. Here, we consider how the presence of explicit network structure changes the optimal solution of a class of infection-interaction trade-offs. In order to do this, we develop and analyse a low-dimensional dynamical system approximating the network SIS epidemic. We find that network structure in the form of heterogeneous numbers of contacts can have a significant impact on the optimal number of contacts that comes out of a trade-off model.

physics.soc-ph

Tracking the Structure and Sentiment of Vaccination Discussions on Mumsnet

Vaccination is one of the most impactful healthcare interventions in terms of lives saved at a given cost, leading the anti-vaccination movement to be identified as one of the top 10 threats to global health in 2019 by the World Health Organization. This issue increased in importance during the COVID-19 pandemic where, despite good overall adherence to vaccination, specific communities still showed high rates of refusal. Online social media has been identified as a breeding ground for anti-vaccination discussions. In this work, we study how vaccination discussions are conducted in the discussion forum of Mumsnet, a United Kingdom based website aimed at parents. By representing vaccination discussions as networks of social interactions, we can apply techniques from network analysis to characterize these discussions, namely network comparison, a task aimed at quantifying similarities and differences between networks. Using network comparison based on graphlets -- small connected network subgraphs -- we show how the topological structure vaccination discussions on Mumsnet differs over time, in particular before and after COVID-19. We also perform sentiment analysis on the content of the discussions and show how the sentiment towards vaccinations changes over time. Our results highlight an association between differences in network structure and changes to sentiment, demonstrating how network comparison can be used as a tool to guide and enhance the conclusions from sentiment analysis.

cs.SI

Calculation of Epidemic First Passage and Peak Time Probability Distributions

Understanding the timing of the peak of a disease outbreak forms an important part of epidemic forecasting. In many cases, such information is essential for planning increased hospital bed demand and for designing of public health interventions. The time taken for an outbreak to become large is inherently stochastic, and therefore uncertain, but after a sufficient number of infections has been reached the subsequent dynamics can be modelled accurately using ordinary differential equations. Here, we present analytical and numerical methods for approximating the time at which a stochastic model of a disease outbreak reaches a large number of cases and for quantifying the uncertainty arising from demographic stochasticity around that time. We then project this uncertainty forwards in time using an ordinary differential equation model in order to obtain a distribution for the peak timing of the epidemic that agrees closely with large simulations but that, for error tolerances relevant to most realistic applications, requires a fraction of the computational cost of full Monte Carlo approaches.

q-bio.PE

Modelling and classifying joint trajectories of self-reported mood and pain in a large cohort study

It is well-known that mood and pain interact with each other, however individual-level variability in this relationship has been less well quantified than overall associations between low mood and pain. Here, we leverage the possibilities presented by mobile health data, in particular the "Cloudy with a Chance of Pain" study, which collected longitudinal data from the residents of the UK with chronic pain conditions. Participants used an App to record self-reported measures of factors including mood, pain and sleep quality. The richness of these data allows us to perform model-based clustering of the data as a mixture of Markov processes. Through this analysis we discover four endotypes with distinct patterns of co-evolution of mood and pain over time. The differences between endotypes are sufficiently large to play a role in clinical hypothesis generation for personalised treatments of comorbid pain and low mood.

stat.AP

Comparing directed networks via denoising graphlet distributions

Network comparison is a widely-used tool for analyzing complex systems, with applications in varied domains including comparison of protein interactions or highlighting changes in structure of trade networks. In recent years, a number of network comparison methodologies based on the distribution of graphlets (small connected network subgraphs) have been introduced. In particular, NetEmd has recently achieved state of the art performance in undirected networks. In this work, we propose an extension of NetEmd to directed networks and deal with the significant increase in complexity of graphlet structure in the directed case by denoising through linear projections. Simulation results show that our framework is able to improve on the performance of a simple translation of the undirected NetEmd algorithm to the directed case, especially when networks differ in size and density.

cs.SI

The role of regular asymptomatic testing in reducing the impact of a COVID-19 wave

Testing for infection with SARS-CoV-2 is an important intervention in reducing onwards transmission of COVID-19, particularly when combined with the isolation and contact-tracing of positive cases. Many countries with the capacity to do so have made use of lab-processed Polymerase Chain Reaction (PCR) testing targeted at individuals with symptoms and the contacts of confirmed cases. Alternatively, Lateral Flow Tests (LFTs) are able to deliver a result quickly, without lab-processing and at a relatively low cost. Their adoption can support regular mass asymptomatic testing, allowing earlier detection of infection and isolation of infectious individuals. In this paper we extend and apply the agent-based epidemic modelling framework Covasim to explore the impact of regular asymptomatic testing on the peak and total number of infections in an emerging COVID-19 wave. We explore testing with LFTs at different frequency levels within a population with high levels of immunity and with background symptomatic PCR testing, case isolation and contact tracing for testing. The effectiveness of regular asymptomatic testing was compared with `lockdown' interventions seeking to reduce the number of non-household contacts across the whole population through measures such as mandating working from home and restrictions on gatherings. Since regular asymptomatic testing requires only those with a positive result to reduce contact, while lockdown measures require the whole population to reduce contact, any policy decision that seeks to trade off harms from infection against other harms will not automatically favour one over the other. Our results demonstrate that, where such a trade off is being made, at moderate rates of early exponential growth regular asymptomatic testing has the potential to achieve significant infection control without the wider harms associated with additional lockdown measures.

q-bio.PE

A computational framework for modelling infectious disease policy based on age and household structure with applications to the COVID-19 pandemic

The widespread, and in many countries unprecedented, use of non-pharmaceutical interventions (NPIs) during the COVID-19 pandemic has highlighted the need for mathematical models which can estimate the impact of these measures while accounting for the highly heterogeneous risk profile of COVID-19. Models accounting either for age structure or the household structure necessary to explicitly model many NPIs are commonly used in infectious disease modelling, but models incorporating both levels of structure present substantial computational and mathematical challenges due to their high dimensionality. Here we present a modelling framework for the spread of an epidemic that includes explicit representation of age structure and household structure. Our model is formulated in terms of tractable systems of ordinary differential equations for which we provide an open-source Python implementation. Such tractability leads to significant benefits for model calibration, exhaustive evaluation of possible parameter values, and interpretability of results. We demonstrate the flexibility of our model through four policy case studies, where we quantify the likely benefits of the following measures which were either considered or implemented in the UK during the current COVID-19 pandemic: control of within- and between-household mixing through NPIs; formation of support bubbles during lockdown periods; out-of-household isolation (OOHI); and temporary relaxation of NPIs during holiday periods. Our ordinary differential equation formulation and associated analysis demonstrate that multiple dimensions of risk stratification and social structure can be incorporated into infectious disease models without sacrificing mathematical tractability. This model and its software implementation expand the range of tools available to infectious disease policy analysts.

q-bio.PE

Diversity of symptom phenotypes in SARS-CoV-2 community infections observed in multiple large datasets

Through the use of cutting-edge unsupervised classification techniques from statistics and machine learning, we characterise symptom phenotypes among symptomatic SARS-CoV-2 PCR-positive community cases. We first analyse each dataset in isolation and across age bands, before using methods that allow us to compare multiple datasets. While we observe separation due to the total number of symptoms experienced by cases, we also see a separation of symptoms into gastrointestinal, respiratory and other types, and different symptom co-occurrence patterns at the extremes of age. In this way, we are able to demonstrate the deep structure of symptoms of COVID-19 without usual biases due to study design. This is expected to have implications for the identification and management of community SARS-CoV-2 cases and could be further applied to symptom-based management of other diseases and syndromes.

stat.AP

Total Effect Analysis of Vaccination on Household Transmission in the Office for National Statistics COVID-19 Infection Survey

We investigate the distribution of numbers of secondary cases in households in the Office for National Statistics COVID-19 Infection Survey (ONS CIS), stratified by timing of vaccination and infection in the households. This shows a total effect of a statistically significant approximate halving of the secondary attack rate in households following vaccination.

q-bio.PE

Inferring Risks of Coronavirus Transmission from Community Household Data

The response of many governments to the COVID-19 pandemic has involved measures to control within- and between-household transmission, providing motivation to improve understanding of the absolute and relative risks in these contexts. Here, we perform exploratory, residual-based, and transmission-dynamic household analysis of the Office for National Statistics (ONS) COVID-19 Infection Survey (CIS) data from 26 April 2020 to 15 July 2021 in England. This provides evidence for: (i) temporally varying rates of introduction of infection into households broadly following the trajectory of the overall epidemic and vaccination programme; (ii) Susceptible-Infectious Transmission Probabilities (SITPs) of within-household transmission in the 15-35% range; (iii) the emergence of the Alpha and Delta variants, with the former being around 50% more infectious than wildtype and 35% less infectious than Delta within households; (iv) significantly (in the range 25-300%) more risk of bringing infection into the household for workers in patient-facing roles pre-vaccine; (v) increased risk for secondary school-age children of bringing the infection into the household when schools are open; (vi) increased risk for primary school-age children of bringing the infection into the household when schools were open since the emergence of new variants.

stat.AP

Fast Approximate Bayesian Contextual Cold Start Learning (FAB-COST)

Cold-start is a notoriously difficult problem which can occur in recommendation systems, and arises when there is insufficient information to draw inferences for users or items. To address this challenge, a contextual bandit algorithm -- the Fast Approximate Bayesian Contextual Cold Start Learning algorithm (FAB-COST) -- is proposed, which is designed to provide improved accuracy compared to the traditionally used Laplace approximation in the logistic contextual bandit, while controlling both algorithmic complexity and computational cost. To this end, FAB-COST uses a combination of two moment projection variational methods: Expectation Propagation (EP), which performs well at the cold start, but becomes slow as the amount of data increases; and Assumed Density Filtering (ADF), which has slower growth of computational cost with data size but requires more data to obtain an acceptable level of accuracy. By switching from EP to ADF when the dataset becomes large, it is able to exploit their complementary strengths. The empirical justification for FAB-COST is presented, and systematically compared to other approaches on simulated data. In a benchmark against the Laplace approximation on real data consisting of over $670,000$ impressions from autotrader.co.uk, FAB-COST demonstrates at one point increase of over $16\%$ in user clicks. On the basis of these results, it is argued that FAB-COST is likely to be an attractive approach to cold-start recommendation systems in a variety of contexts.

stat.ML