arXiv Science⌕ Search

arXiv · 1606.07769

How to estimate epidemic risk from incomplete contact diaries data?

Abstract

Social interactions shape the patterns of spreading processes in a population. Techniques such as diaries or proximity sensors allow to collect data about encounters and to build networks of contacts between individuals. The contact networks obtained from these different techniques are however quantitatively different. Here, we first show how these discrepancies affect the prediction of the epidemic risk when these data are fed to numerical models of epidemic spread: low participation rate, under-reporting of contacts and overestimation of contact durations in contact diaries with respect to sensor data determine indeed important differences in the outcomes of the corresponding simulations {with for instance an enhanced sensitivity to initial conditions}. Most importantly, we investigate if and how information gathered from contact diaries can be used in such simulations in order to yield an accurate description of the epidemic risk, assuming that data from sensors represent the ground truth. The contact networks built from contact sensors and diaries present indeed several structural similarities: this suggests the possibility to construct, using only the contact diary network information, a surrogate contact network such that simulations using this surrogate network give the same estimation of the epidemic risk as simulations using the contact sensor network. We present and {compare} several methods to build such surrogate data, and show that it is indeed possible to obtain a good agreement between the outcomes of simulations using surrogate and sensor data, as long as the contact diary information is complemented by publicly available data describing the heterogeneity of the durations of human contacts.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rossana Mastrandrea, Alain Barrat. 2016-06-27. How to estimate epidemic risk from incomplete contact diaries data?. https://doi.org/10.1371/journal.pcbi.1005002

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Optimality as a Generative Principle for Network Structure

Most real-world networks have evolved or been engineered to optimize some function, yet a unified framework for studying optimal networks across domains is lacking. We introduce GradNet, an AI-enabled framework that treats network topology as a continuously differentiable object, enabling the design and study of networks that optimize structural and dynamical objectives amenable to automatic differentiation under realistic constraints. We derive general optimality conditions, including an equimarginal principle for linear budgets, that make optimized networks analytically tractable. Canonical network features emerge spontaneously from constrained optimization: maximizing Kuramoto synchronization under coupling budgets yields sparse, bipartite, frequency-disassortative networks; minimizing social tension in opinion dynamics reproduces the factional split in Zachary's karate club; and maximizing communication capacity in spatial quantum networks under distance-dependent costs recovers minimum spanning trees. GradNet thus serves both as a network design tool scalable beyond $10^5$ nodes and as a scientific probe of structure-function relationships.

physics.soc-ph↗

Community-Centers Identify Robust Biomarker in High-Dimensional, Low-Sample-Size Gene Expression Data

High-dimensional, low-sample-size bulk gene expression data poses a fundamental challenge in transcriptomics, which typically includes tens of thousands of genes but relatively few samples, leading to overfitting and unstable selection for key features as biomarkers. We propose a regression by community centers (RCC) framework tailored for such data. RCC converts gene expression data into a feature proximity network and leverages the phase transition of the giant connected component to identify a critical distance threshold that yields a sparse yet maximally informative network representation, indicated by a minimum normalized shortest compression length. Intuitively, at the criticality, the network balances fragmentation and over-connectivity, allowing meaningful gene-gene associations communities to emerge while filtering noises, such that community-centers capture the most informative and non-redundant signals. These representative genes are then fed to a simple ordinary least squares (OLS) model for downstream predictions. Across five bulk gene expression datasets, RCC displays exceptional prediction performance, clearly outperforming benchmark methods, remains robust to missing data and noises, and does not rely on external biological knowledge. These results suggest that network-based representations provide an effective and general framework for high-dimensional, low-sample-size prediction tasks beyond transcriptomics and illustrate how network-science ideas can support robust learning in data-scarce regimes.

physics.soc-ph↗

The efficacy-competitiveness frontier in the Swiss system

An ideal tournament is both efficacious and attractive: it selects the best player as the winner, but only at the end of the tournament in order to maintain suspense as long as possible. We identify and explore the trade-off between these criteria in Swiss-system chess team tournaments via simulations. Organising more rounds is shown to improve efficacy but increase the probability of a stakeless last round. The 2008 reform in the Chess Olympiad ranking is found to be a questionable move, as the previous rule based on board points determines the efficacy-competitiveness frontier, regardless of whether two or three points are awarded for winning a match. Combining board points and match points is better than using match points only. Our results can help optimise the design of the Swiss system, which, besides chess, is also widely used in e-sports.

physics.soc-ph↗