arXiv ScienceSearch

arXiv subjects

Alessandro Colombi

Publications and source records attributed to Alessandro Colombi.

7 recordsLinked to original sources

Bayesian nonparametric inference for modal missing species and features

Species and feature sampling problems arise naturally whenever each observed unit is associated with one or more labels from a countable alphabet, and inference focuses on the unobserved portion of the distribution over such an alphabet. Within this framework, recent contributions have shifted attention from inference on the total probability mass over unseen labels to distribution-free confidence intervals for the largest unobserved label probability (i.e., the modal missing probability), thereby providing more refined information on whether the missing mass is concentrated on a few high-prevalence unseen labels or distributed across many negligible ones. Besides lacking a unified framework for species and features, these approaches employ a worst-case perspective that causes substantial information loss and produces overly-conservative intervals. We address these limitations through a unified model-based framework for inference on modal missing probabilities in both species and feature settings, which leverages a flexible Bayesian nonparametric formulation to localize uncertainty around models compatible with the observed data. This leads to sharper closed-form credible intervals that effectively exploit prior information and observed data, while preserving the theoretical frequentist properties and robustness of distribution-free intervals. Simulation studies confirm these improvements, while an organized crime application illustrates how our contribution has the potential to reshape law-enforcement decision-making in investigations.

stat.ME

Confidence intervals for maximum unseen probabilities, with application to sequential sampling design

Discovery problems often require deciding whether additional sampling is needed to detect all categories whose prevalence exceeds a prespecified threshold. We study this question under a Bernoulli product (incidence) model, where categories are observed only through presence--absence across sampling units. Our inferential target is the \emph{maximum unseen probability}, the largest prevalence among categories not yet observed. We develop nonasymptotic, distribution-free upper confidence bounds for this quantity in two regimes: bounded alphabets (finite and known number of categories) and unbounded alphabets (countably infinite under a mild summability condition). We characterise the limits of data-independent worst-case bounds, showing that in the unbounded regime no nontrivial data-independent procedure can be uniformly valid. We then propose data-dependent bounds in both regimes and establish matching lower bounds demonstrating their near-optimality. We compare empirically the resulting procedures in both simulated and real datasets. Finally, we use these bounds to construct sequential stopping rules with finite-sample guarantees, and demonstrate robustness to contamination that introduces spurious low-prevalence categories.

stat.ME

Bayesian discovery of species in multiple areas

In ecology, the description of species composition and biodiversity calls for statistical methods that involve estimating features of interest in unobserved samples based on an observed one. In the last decade, the Bayesian nonparametrics literature has thoroughly investigated the case where data arise from a homogeneous population. In this work, we propose a novel framework to address heterogeneous populations, specifically dealing with scenarios where data arise from two areas. This setting significantly increases the mathematical complexity of the problem and, as a consequence, it has received limited attention in the literature. While early approaches leverage computational methods, we provide a distributional theory for the in-sample analysis of any observed sample and enable out-of-sample prediction for the number of unseen distinct and shared species in additional samples of arbitrary sizes. The latter also extends the frequentist estimators, which solely deal with one-step-ahead prediction. Furthermore, our results can be applied to address sample size determination in sampling problems aimed at detecting distinct and shared species. Our results are illustrated in a real-world dataset concerning a population of ants in the city of Trieste.

stat.ME

Hierarchical Mixture of Finite Mixtures

Statistical modelling in the presence of data organized in groups is a crucial task in Bayesian statistics. The present paper conceives a mixture model based on a novel family of Bayesian priors designed for multilevel data and obtained by normalizing a finite point process. In particular, the work extends the popular Mixture of Finite Mixture model to the hierarchical framework to capture heterogeneity within and between groups. A full distribution theory for this new family and the induced clustering is developed, including the marginal, posterior, and predictive distributions. Efficient marginal and conditional Gibbs samplers are designed to provide posterior inference. The proposed mixture model overcomes the Hierarchical Dirichlet Process, the utmost tool for handling multilevel data, in terms of analytical feasibility, clustering discovery, and computational time. The motivating application comes from the analysis of shot put data, which contains performance measurements of athletes across different seasons. In this setting, the proposed model is exploited to induce clustering of the observations across seasons and athletes. By linking clusters across seasons, similarities and differences in athletes' performances are identified.

stat.ME

Learning block structured graphs in Gaussian graphical models

Within the framework of Gaussian graphical models, a prior distribution for the underlying graph is introduced to induce a block structure in the adjacency matrix of the graph and learning relationships between fixed groups of variables. A novel sampling strategy named Double Reversible Jumps Markov chain Monte Carlo is developed for block structural learning, under the conjugate G-Wishart prior. The algorithm proposes moves that add or remove not just a single link but an entire group of edges. The method is then applied to smooth functional data. The classical smoothing procedure is improved by placing a graphical model on the basis expansion coefficients, providing an estimate of their conditional independence structure. Since the elements of a B-Spline basis have compact support, the independence structure is reflected on well-defined portions of the domain. A known partition of the functional domain is exploited to investigate relationships among the substances within the compound.

stat.ME

Gaussian graphical modeling for spectrometric data analysis

Motivated by the analysis of spectrometric data, we introduce a Gaussian graphical model for learning the dependence structure among frequency bands of the infrared absorbance spectrum. The spectra are modeled as continuous functional data through a B-spline basis expansion and a Gaussian graphical model is assumed as a prior specification for the smoothing coefficients to induce sparsity in their precision matrix. Bayesian inference is carried out to simultaneously smooth the curves and to estimate the conditional independence structure between portions of the functional domain. The proposed model is applied to the analysis of infrared absorbance spectra of strawberry purees.

stat.ME

Exotic atoms at extremely high magnetic fields: the case of neutron star atmosphere

The presence of exotic states of matter in neutron stars (NSs) is currently an open issue in physics. The appearance of muons, kaons, hyperons, and other exotic particles in the inner regions of the NS, favored by energetic considerations, is considered to be an effective mechanism to soften the equation of state (EoS). In the so-called two-families scenario, the softening of the EoS allows for NSs characterized by very small radii, which become unstable and convert into a quark stars (QSs). In the process of conversion of a NS into a QS material can be ablated by neutrinos from the surface of the star. Not only neutron-rich nuclei, but also more exotic material, such as hypernuclei or deconfined quarks, could be ejected into the atmosphere. In the NS atmosphere, atoms like H, He, and C should exist, and attempts to model the NS thermal emission taking into account their presence, with spectra modified by the extreme magnetic fields, have been done. However, exotic atoms, like muonic hydrogen $(p\,\mu^-)$ or the so-called Sigmium $(\Sigma^+\,e^-)$, could also be present during the conversion process or in its immediate aftermath. At present, analytical expressions of the wave functions and eigenvalues for these atoms have been calculated only for H. In this work, we extend the existing solutions and parametrizations to the exotic atoms $(p\,\mu^-)$ and $(\Sigma^+\,e^-)$, making some predictions on possible transitions. Their detection in the spectra of NS would provide experimental evidence for the existence of hyperons in the interior of these stars.

nucl-th