arXiv ScienceSearch

arXiv subjects

Mayetri Gupta

Publications and source records attributed to Mayetri Gupta.

6 recordsLinked to original sources

A Bayesian Bi-Directional Splitting Framework for Variable Selection in Large Datasets

Modern tabular datasets are becoming increasingly large, both in the number of samples and covariates, posing significant challenges for Bayesian variable selection due to the resulting computational burden. While there is extensive literature on scaling Bayesian inference to large numbers of observations or high-dimensional covariate spaces, comparatively little work addresses scalable Bayesian variable selection when both dimensions are large simultaneously in a practical setting. This paper presents a novel Bayesian variable selection framework for efficiently analysing data with a large number of both rows and columns. The proposed framework operates via a divide-and-conquer approach, splitting data into batches along both directions and analysing each batch independently in parallel, after which the results are combined together in a two-phase consensus procedure. Experiments demonstrate the computational gains while retaining strong variable selection performance, successfully identifying relevant signals in noisy, high-dimensional settings despite the loss of information induced by data splitting. Practical guidelines are also provided, including recommendations for tuning parameter choices and effective data partitioning strategies, followed by a real application to an H3N2 influenza dataset.

stat.ME

ZINBGT: Exploratory Data Analysis of Single-Cell Transcriptomic Expression Using Mixture Models

Single-cell transcriptomic data approximates the abundance of proteins at a high resolution, but its noisiness necessitates transformation by a pipeline of methods before analysis and inference. In the absence of robust validation of these pipelines and methods, it remains unclear how best to process any particular dataset. To compensate for this, popular visualisation methods, e.g., t-SNE and UMAP, are commonly used to produce descriptions of datasets. Such visualisations are incomplete and provide subjective descriptions of samples rather than statistically meaningful statements about technical noise or biology. In this paper, we introduce the Zero-Inflated Negative-Binomial with Geometric Tail (ZINBGT), a mixture-model-based strategy for producing interpretable visualisations of each gene's expression across cells, along with diagnostic summaries that use Wasserstein distance to highlight outlier genes. These diagnostics are used to reveal an outlier gene within a T. brucei sample. This method is applied to a human immune-cell dataset, highlighting the relationship between sparsity, mean, and spread across genes, as well as revealing an issue with the use of zero-inflated negative-binomial distributions to model single-cell RNA data. An investigation of simulated datasets intended to replicate the immune-cell data revealed discrepancies with the ground truth, establishing purposes for which these simulated datasets are unsuitable. Finally, we list a number of different domains to which this method can be applied.

stat.AP

On the identifiability of Dirichlet mixture models

We study identifiability of finite mixtures of Dirichlet distributions on the interior of the simplex. We first prove a shift identity showing that every Dirichlet density can be written as a mixture of $J$ shifted Dirichlet densities, where $J-1$ is the dimension of the simplex support, which yields non-identifiability on the full parameter space. We then show that identifiability is recovered on a fixed-total parameter slice and on restricted box-type regions. On the full parameter space, we prove that any nontrivial linear relation among Dirichlet kernels must involve at least $J$ coefficients sharing a common sign, and deduce that mixtures with fewer than $J$ atoms are identifiable. We further report direct non-identifiability implications for unrestricted finite mixtures of generalized Dirichlet, Dirichlet-multinomial, fixed-topic-matrix latent Dirichlet allocation, Beta-Liouville, and inverted Beta-Liouville models.

math.ST

On the large-sample limits of some Bayesian model evaluation statistics

Model selection and order selection problems frequently arise in statistical practice. A popular approach to addressing these problems in the frequentist setting involves information criteria based on penalised maxima of log-likelihoods for competing models. In the Bayesian context, similar criteria are employed, replacing the maximised log-likelihoods with posterior expectations of the log-likelihood. Despite their popularity in applications, the large-sample behaviour of these criteria -- such as the deviance information criterion (DIC), Bayesian predictive information criterion (BPIC), and widely applicable Bayesian information criterion (WBIC) -- has received relatively little attention. In this work, we investigate the almost-sure limits of these criteria and establish novel results on posterior and generalised posterior consistency, which are of independent interest. The utility of our theoretical findings is demonstrated via illustrative technical and numerical examples.

math.ST

Finite sample inference for empirical Bayesian methods

In recent years, empirical Bayesian (EB) inference has become an attractive approach for estimation in parametric models arising in a variety of real-life problems, especially in complex and high-dimensional scientific applications. However, compared to the relative abundance of available general methods for computing point estimators in the EB framework, the construction of confidence sets and hypothesis tests with good theoretical properties remains difficult and problem specific. Motivated by the universal inference framework of Wasserman et al. (2020), we propose a general and universal method, based on holdout likelihood ratios, and utilizing the hierarchical structure of the specified Bayesian model for constructing confidence sets and hypothesis tests that are finite sample valid. We illustrate our method through a range of numerical studies and real data applications, which demonstrate that the approach is able to generate useful and meaningful inferential statements in the relevant contexts.

stat.ME

Model selection and sensitivity analysis for sequence pattern models

In this article we propose a maximal a posteriori (MAP) criterion for model selection in the motif discovery problem and investigate conditions under which the MAP asymptotically gives a correct prediction of model size. We also investigate robustness of the MAP to prior specification and provide guidelines for choosing prior hyper-parameters for motif models based on sensitivity considerations.

math.ST