arXiv ScienceSearch

arXiv subjects

Elizabeth Juarez-Colunga

Publications and source records attributed to Elizabeth Juarez-Colunga.

5 recordsLinked to original sources

Generalized Bayesian Clustering with Regression for Unaligned Longitudinal Binary Data

We propose a generalized Bayesian clustering with regression model for unaligned longitudinal binary outcomes, motivated by seizure diary data from the Human Epilepsy Project. Seizure diaries are sparse, irregularly observed, and vary enormously across patients. A single fully-specified generative model tends to be either misspecified or computationally inefficient. We address the challenge by two strategies. We set up a regression by way of clustering as model-based clustering using a mixture model. For the latter, we take a generalized Bayesian perspective which replaces the full likelihood with a loss-based update using a generalized likelihood. We combine a trajectory similarity loss and a regression loss, so that clustering is informed by both trajectory similarity and the prediction of outcomes The trajectory similarity loss is constructed by representing each trajectory as an (empirical) distribution of subsequences, called reads, and then is defined based on the sliced Wasserstein distance between these empirical distributions. This loss allows alignment-free comparison of sequences that are irregularly observed or temporally misaligned, and it scales quasi-linearly in trajectory length. The regression loss is the negative log-likelihood of a probit regression. A prior on the cluster-specific parameters is defined by way of a Dirichlet process prior on the mixing measure.

stat.ME

Distributional Determinantal Point Process for Repulsive Clustering of Distributions

We introduce the distributional determinantal point process (dDPP) as a novel repulsive point process whose atoms are probability distributions rather than points in a real space. The dDPP is constructed via an L-ensemble with a sliced Wasserstein (SW) kernel between distributions. We show its validity as a well-defined point process. In the discrete setting, we derive concentration results for plug-in estimators of the L-ensemble, the correlation kernel, and their determinants given i.i.d. samples from the distributional atoms. Leveraging this framework, we propose a distribution-valued random partition model by way of a repulsive generalized Bayesian mixture model. The model places a dDPP prior over the atoms of the mixing measure and defines a generalized likelihood based on SW distance. To summarize posterior inference, we develop a decision-theoretic approach to report a point estimate of the mixing measure as a Bayes rule under a hierarchical optimal transport utility function. The latter is a natural choice given that the mixing measure is itself a distribution over distributions. We use the proposed framework for inference with single-cell gene expression data and human epilepsy data, producing interpretable and well-separated clusters that reflect meaningful structure in the data.

stat.ME

A Mixed Self-Exciting Process to Model Epileptic Seizures

Epilepsy is a neurological disorder characterized by recurrent seizures affecting more than 70 million people worldwide. Often, an individual with epilepsy is more likely to experience subsequent seizures following an initial seizure, a process we call seizure clustering. Motivated by seizure diary data collected over three years from 407 individuals newly diagnosed with focal epilepsy in the Human Epilepsy Project (HEP), we propose a Bayesian mixed Hawkes process model that addresses seizure clustering and heterogeneity between individuals. In the Hawkes process, the intensity is accelerated each time an event occurs, through the composition of background and excitation intensity functions. The proposed model incorporates a Weibull baseline intensity to model a trend in background seizure rates over time, while the excitation process accounts for seizure clustering within individuals. We model heterogeneity among individuals by including covariates and random effects in both the background and excitation intensities. In the HEP study, the average time between primary and secondary seizures within an individual is 1.57 (95\% CrI: 1.43, 1.70) days, with an average of 2.20 (1.96, 2.47) seizures per cluster. We demonstrate that omitting random effects in the presence of heterogeneity leads to underestimation of the background intensity and overestimation of excitation rates.

stat.ME

Borrowing strength between unaligned binary time-series via Bayesian nonparametric rescaling of Unified Skewed Normal priors

We define a Bayesian semi-parametric model to effectively conduct inference with unaligned longitudinal binary data. The proposed strategy is motivated by data from the Human Epilepsy Project (HEP), which collects seizure occurrence data for epilepsy patients, together with relevant covariates. The model is designed to flexibly accommodate the particular challenges that arise with such data. First, epilepsy data require models that can allow for extensive heterogeneity, across both patients and time. With this regard, state space models offer a flexible, yet still analytically amenable class of models. Nevertheless, seizure time-series might share similar behavioral patterns, such as local prolonged periods of elevated seizure presence, which we refer to as "clumping". Such similarities can be used to share strength across patients and define subgroups. However, due to the lack of alignment, straightforward hierarchical modeling of latent state space parameters is not practicable. To overcome this constraint, we construct a strategy that preserves the flexibility of individual trajectories while also exploiting similarities across individuals to borrow information through a nonparametric prior. On the one hand, heterogeneity is ensured by (almost) subject-specific state-space submodels. On the other, borrowing of information is obtained by introducing a Pitman-Yor prior on group-specific probabilities for patterns of clinical interest. We design a posterior sampling strategy that leverages recent developments of binary state space models using the Unified Skewed Normal family (SUN). The model, which allows the sharing of information across individuals with similar disease traits over time, can more generally be adapted to any setting characterized by unaligned binary longitudinal data.

stat.ME

A Framework for Covariate Balance using Bregman Distances

A common goal in observational research is to estimate marginal causal effects in the presence of confounding variables. One solution to this problem is to use the covariate distribution to weight the outcomes such that the data appear randomized. The propensity score is a natural quantity that arises in this setting. Propensity score weights have desirable asymptotic properties, but they often fail to adequately balance covariate data in finite samples. Empirical covariate balancing methods pose as an appealing alternative by exactly balancing the sample moments of the covariate distribution. With this objective in mind, we propose a framework for estimating balancing weights by solving a constrained convex program where the criterion function to be optimized is a Bregman distance. We then show that the different distances in this class render identical weights to those of other covariate balancing methods. A series of numerical studies are presented to demonstrate these similarities.

stat.ME