arXiv ScienceSearch

arXiv subjects

Stefano Favaro

Publications and source records attributed to Stefano Favaro.

At least 19 recordsLinked to original sources

Online activity prediction via generalized Indian buffet process models

Online A/B tests are the standard tool for data-driven decision-making at scale. Among the design choices with the largest impact on statistical power is the triggering mechanism: how many users to expose and for how long. This often requires forecasting user engagement, i.e., whether enough users will trigger, and when a target participation level will be reached, from limited pilot data. We introduce a Bayesian nonparametric model for predicting both new-user counts and total triggers, accommodating the heavy-tailed engagement patterns typical of web experiments. All predictive quantities can be computed without intensive numerical procedures such as Markov chain Monte Carlo (MCMC) or variational inference. We evaluate on three public datasets (over 450 public benchmark evaluations) and a proprietary benchmark drawn from 759 production A/B tests comprising 1,774 arms. Across the benchmark analyses, our models are competitive and frequently improve accuracy in forecasting new users, total triggers, and time to reach a target sample size compared with state-of-the-art competitors, especially when only a few pilot days are observed.

stat.AP

Bayesian nonparametric inference for modal missing species and features

Species and feature sampling problems arise naturally whenever each observed unit is associated with one or more labels from a countable alphabet, and inference focuses on the unobserved portion of the distribution over such an alphabet. Within this framework, recent contributions have shifted attention from inference on the total probability mass over unseen labels to distribution-free confidence intervals for the largest unobserved label probability (i.e., the modal missing probability), thereby providing more refined information on whether the missing mass is concentrated on a few high-prevalence unseen labels or distributed across many negligible ones. Besides lacking a unified framework for species and features, these approaches employ a worst-case perspective that causes substantial information loss and produces overly-conservative intervals. We address these limitations through a unified model-based framework for inference on modal missing probabilities in both species and feature settings, which leverages a flexible Bayesian nonparametric formulation to localize uncertainty around models compatible with the observed data. This leads to sharper closed-form credible intervals that effectively exploit prior information and observed data, while preserving the theoretical frequentist properties and robustness of distribution-free intervals. Simulation studies confirm these improvements, while an organized crime application illustrates how our contribution has the potential to reshape law-enforcement decision-making in investigations.

stat.ME

Asymptotic behavior of clusters in hierarchical species sampling models

Consider a random sample of size $N$ from a hierarchical species sampling model. In this paper, we study the large $N$ asymptotic behavior of the number ${\bf K}_N$ of clusters in the random sample and the number ${\bf \widetilde M}_{\ell,N}$ of clusters represented by exactly $\ell$ latent first-level clusters in the first level of the hierarchical model. In particular, we establish almost sure and $L^p$ convergence for ${\bf \widetilde M}_{\ell,N}$, Gaussian fluctuations and the law of the iterated logarithm for ${\bf K}_N$, and large deviation principles for both ${\bf K}_N$ and ${\bf \widetilde M}_{\ell,N}$. Our approach relies on a random sample size (or random index) representation of the number of clusters through the corresponding non-hierarchical species sampling model.

math.PR

Elements of Conformal Prediction

Predictive inference is a fundamental task in statistics, traditionally addressed using parametric assumptions about the data distribution and detailed analyses of how models learn from data. In recent years, conformal prediction has emerged as an alternative framework that is well suited to modern applications involving high-dimensional data and complex machine learning models. Its appeal stems from being both distribution-free---relying mainly on symmetry assumptions such as exchangeability---and model-agnostic, treating the learning algorithm as a black box. Even under such limited assumptions, conformal prediction provides exact finite-sample guarantees, although these are typically marginal and require careful interpretation. This paper explains the core ideas of conformal prediction and reviews selected methods. Rather than offering an exhaustive survey, it aims to provide a clear conceptual entry point and a pedagogical overview of the field.

stat.ME

Posterior contraction rates in Sobolev norms and Bayesian derivative estimation for infinite-dimensional exponential families

We study posterior contraction in positive-order Sobolev norms and Bayesian derivative estimation for infinite-dimensional exponential families. We embed the natural parameter in a Hilbert scale and model it via a standard Gaussian series prior expanded in the eigenbasis generating the scale. Under a two-sided link condition on the Fisher information and suitable local regularity assumptions, we show that smoothness-matching priors achieve minimax-optimal posterior contraction rates in any Hilbert scale norm up to the regularity of the ground truth. Our analysis builds on the novel approach to posterior contraction based on the Wasserstein distance recently introduced by Dolera et al. (2024). It combines refined Laplace-type estimates for infinite-dimensional integrals associated to the posterior kernels with a mixed-geometry estimate controlling their stability under fluctuations in the data, itself resting on a tailored Poincaré inequality for posterior distributions conditioned on neighbourhoods of the truth. We apply the general theory to density estimation with a logistic parametrisation, Poisson intensity estimation with an exponential link, and the Gaussian white-noise model, yielding minimax contraction rates in Sobolev norms across all three settings. In particular, these yield optimal recovery of density score functions and derivatives of Poisson intensities.

math.ST

The variance of the Pitman--Yor process: a Cifarelli--Regazzini identity and inversion formula

The celebrated Cifarelli--Regazzini identity for the Dirichlet process and its analytic inversion lie at the foundation of an elegant distributional theory for linear functionals of random probability measures, providing, in particular, an explicit formula for the density function of the Dirichlet mean. Nonlinear functionals of the Dirichlet process, by contrast, remain substantially less understood in the literature. In this paper, we develop a transform-and-inversion strategy beyond linear functionals by considering the variance of the Pitman--Yor process, which generalizes the Dirichlet process. We establish a Cifarelli--Regazzini identity for the Pitman--Yor variance, whose direct analytic inversion yields an explicit formula for its density function. The corresponding results for the Dirichlet variance are recovered as a limiting case.

math.PR

Quasi-Bayesian sequential deconvolution

Density deconvolution is the inverse problem of estimating a probability density from observations contaminated by additive noise. Traditionally studied in static or batch settings, it increasingly arises with streaming data, where existing frequentist and Bayesian procedures face substantial computational bottlenecks. We develop a quasi-Bayesian nonparametric method for sequential density deconvolution based on Newton's recursive algorithm. The resulting estimate is straightforward to evaluate and scalable to massive datasets, as its per-observation computational cost remains constant as new data arrive. The quasi-Bayesian interpretation enables uncertainty quantification: local and uniform central limit theorems yield asymptotic credible intervals and bands, respectively. Under a frequentist data-generating model, we establish $L^1$-consistency for the proposed estimate and show that it asymptotically agrees with the estimate that would be obtained if the uncontaminated variables were directly observed. Further, under additional regularity conditions, we derive an $L^1$-Wasserstein convergence rate and establish merging, at an explicit rate, with the Bayesian nonparametric posterior mean estimate under a Dirichlet process mixture model. Synthetic-data experiments, together with an acquisition-ordered flow-cytometry application, demonstrate accuracy comparable to Bayesian nonparametric and Fourier deconvolution methods, while offering a substantial computational advantage.

stat.ME

Merging of Bayes and quasi-Bayes empirical Bayes procedures for Poisson compound decisions

The Poisson compound decision problem is a long-standing problem in statistics, in which empirical Bayes methods are used to estimate Poisson means under a mixture model. We study this problem from the viewpoint of $g$-modeling, comparing two nonparametric strategies for estimating the unknown mixing distribution: a Bayesian empirical Bayes strategy, based on the Dirichlet process posterior, and a quasi-Bayesian empirical Bayes strategy, based on Newton's algorithm. The latter is computationally attractive, but its relationship with the Bayesian strategy requires theoretical justification. Under a Poisson mixture model with a ``true'', or oracle, mixing distribution, we establish concentration rates for the marginal probability mass functions induced by the Bayesian and quasi-Bayesian estimates. These rates are then translated into rates of decay for the corresponding regrets, interpreted as excess Bayes risks, and used to prove a frequentist merging result between the Bayesian and quasi-Bayesian empirical Bayes strategies. We also extend the analysis to the multidimensional Poisson compound decision problem. Numerical experiments on synthetic data illustrate that the quasi-Bayesian strategy achieves accuracy comparable to the Bayesian strategy, while requiring substantially fewer computational resources, especially in the multidimensional setting.

stat.ME

Bayesian Nonparametric Privacy-Preserving Synthetic Data Generation: I. Discrete Data

Synthetic data generation is a powerful approach to privacy-preserving statistical analysis, where data-release mechanisms are governed by a privacy-utility tradeoff: they should provide privacy guarantees while preserving the statistical utility of confidential data. We develop a Bayesian nonparametric framework for private synthetic data generation tailored to discrete data. Specifically, the confidential data are modeled as a random sample from an unknown discrete distribution endowed with a Pitman-Yor process prior, and synthetic data are generated from the corresponding posterior-predictive distribution. Since the Pitman-Yor process defines an almost surely discrete random probability measure, the resulting mechanism is naturally suited to data with ties and settings involving a potentially large, unknown, or growing number of categories. We study differential privacy guarantees of the Pitman-Yor posterior-predictive mechanism across the three regimes of the discount parameter $σ\in(-\infty,1)$. For $σ\in(0,1)$, we establish an instance-level $(\varepsilon,δ)$-differential privacy guarantee. For $σ=0$ and $σ<0$, corresponding respectively to the Dirichlet process prior and to a parametric Dirichlet-Multinomial model, stronger guarantees are obtained, under suitable conditions on the released sample size. We also investigate statistical utility, or informativity, of the released data via the expected $1$-Wasserstein distance between the empirical distribution of the synthetic data and the "true" data-generating distribution. For $σ<0$ and $σ=0$, we prove consistency of the empirical distribution in this metric and derive explicit convergence rates, making precise the privacy-utility tradeoff: stronger privacy guarantees impose more restrictive choices of the released sample size, slowing down convergence to the "true" data-generating distribution.

math.ST

Quasi-Bayes empirical Bayes estimation of sums of random variables

The estimation of sums of functions of observable and unobservable variables is a long-standing problem in statistics with applications across many domains. Empirical Bayes methods provide a natural framework for this task under mixture models, but existing approaches often rely on restrictive parametric assumptions or apply only to limited classes of functionals in nonparametric settings. We propose a nonparametric methodology, referred to as quasi-Bayes empirical Bayes, that addresses these limitations through a recursive estimation of the mixing distribution based on Newton's algorithm. The resulting plug-in estimate of the target sum is computationally efficient, scalable, and applicable to a broad class of utility functions, while enabling uncertainty quantification via asymptotic credible intervals derived from a Gaussian central limit theorem. We establish large sample asymptotic theoretical guarantees by proving a merging between the quasi-Bayes and Bayes estimates and by showing consistency under a correctly specified frequentist model. Synthetic-data and real-data analyses demonstrate the practical accuracy and stability of the method, with performance comparable to, and in some cases better than, existing empirical Bayes procedures.

stat.ME

Quasi-Bayes empirical Bayes: a sequential approach to the Poisson compound decision problem

The Poisson compound decision problem is a long-standing problem is statistics, for which empirical Bayes methods are commonly used to estimate Poisson means in static or batch settings. We consider this problem in a streaming, or online, framework. Building on a quasi-Bayesian approach based on Newton's algorithm, we develop a sequential estimate that is easy to evaluate, computationally efficient, and has constant per-observation cost as the data accrue. We establish frequentist guarantees for the proposed estimate, including consistency and asymptotic optimality, with optimality understood as vanishing excess Bayes risk, or regret. Empirical performance is assessed through simulation studies and comparisons with benchmark procedures.

stat.ME

Asymptotic regimes for maximum likelihood estimation in the Ewens--Pitman model: When the strength parameter matters

We study the large sample asymptotic behaviour of the Maximum Likelihood Estimator of the discount and strength parameters $(α,θ)$ in the Ewens--Pitman model for random partitions, under mild assumptions on the data-generating mechanism. We show that four distinct regimes arise, depending on the limiting behaviour of the frequency spectrum. In particular, in contrast with previous work, we find that $θ$ may play a crucial role asymptotically. We further show that the existing literature implicitly focuses on only two of these regimes, and we relate this restriction to the constraints imposed by infinite exchangeability. Under the latter, indeed, the number of distinct blocks and the frequency spectrum are necessarily tied by a rigid structural relation. We prove that this lack of flexibility can be overcome through what we call the scaled Ewens--Pitman model, in which $θ$ is allowed to grow with the sample size $n$. Finally, we provide empirical evidence from real-world data showing that such extensions are needed to capture frequency spectra that fall outside the classical Ewens--Pitman framework.

math.ST

Confidence intervals for maximum unseen probabilities, with application to sequential sampling design

Discovery problems often require deciding whether additional sampling is needed to detect all categories whose prevalence exceeds a prespecified threshold. We study this question under a Bernoulli product (incidence) model, where categories are observed only through presence--absence across sampling units. Our inferential target is the \emph{maximum unseen probability}, the largest prevalence among categories not yet observed. We develop nonasymptotic, distribution-free upper confidence bounds for this quantity in two regimes: bounded alphabets (finite and known number of categories) and unbounded alphabets (countably infinite under a mild summability condition). We characterise the limits of data-independent worst-case bounds, showing that in the unbounded regime no nontrivial data-independent procedure can be uniformly valid. We then propose data-dependent bounds in both regimes and establish matching lower bounds demonstrating their near-optimality. We compare empirically the resulting procedures in both simulated and real datasets. Finally, we use these bounds to construct sequential stopping rules with finite-sample guarantees, and demonstrate robustness to contamination that introduces spurious low-prevalence categories.

stat.ME

A Central Limit Theorem for the Ewens-Pitman random partition in the large-$θ$ regime via a martingale approach

The Ewens-Pitman model defines a distribution on random partitions of $\{1,\ldots,n\}$, with parameters $α\in [0,1)$ and $θ> -α$; the case $α=0$ reduces to the classical Ewens model from population genetics. We investigate the large-$n$ asymptotic behaviour of the Ewens-Pitman random partition in the nonstandard regime $θ=λn$ with $λ>0$, establishing joint fluctuation results for the total number of blocks $K_n^{\{n\}}$ and the counts $K_{r,n}^{\{n\}}$ of blocks of sizes $r=1,\dots,d$, for fixed $d\in\mathbb{N}$. In particular, for $α\in[0,1)$ and $θ=λn$, our main result provides a strong law of large numbers and a central limit theorem for the $(d+1)$-dimensional vector $\mathbf{K}_{d,n}^{\{n\}} = \bigl(K_n^{\{n\}}, K_{1,n}^{\{n\}}, \dots, K_{d,n}^{\{n\}}\bigr)^T$ as $n \to \infty$. The proof exploits the Chinese restaurant sequential construction under $θ=λn$ and a central limit theorem for triangular arrays of martingales, extending techniques previously developed for the classical regime with fixed $θ$. As corollaries of our results, we recover known asymptotics for $K_n^{\{n\}}$ and derive new strong laws and central limit theorems for each fixed $K_{r,n}^{\{n\}}$, thereby completing earlier weak-law results and providing a comprehensive asymptotic description of the Ewens-Pitman partition structure in the large-$θ$ setting.

math.PR

A Gaussian process limit for the self-normalized Ewens-Pitman process

For an integer $n\geq1$, consider a random partition $Π_{n}$ of $\{1,\ldots,n\}$ into $K_{n}$ partition sets with $K_{r,n}$ partition subsets of size $r=1,\ldots,n$, and assume $Π_{n}$ distributed according to the Ewens-Pitman model with parameters $α\in]0,1[$ and $θ>-α$. Although the large-$n$ asymptotic behaviors of $K_{n}$ and $K_{r,n}$ are well understood in terms of almost sure convergence and Gaussian fluctuations, much less is known about the asymptotic behavior of $P_{r,n}=K_{r,n}/K_n$ and of the self-normalized Ewens-Pitman process $(P_{1,n},P_{2,n},\dots)$. Motivated by the almost sure convergence of $(P_{1,n},P_{2,n},\dots)$ to the Sibuya distribution $p_α=(p_α(1),p_α(2),\ldots)$, where $p_α(r)$ is the probability mass at $r=1,2,\ldots$, we establish the $\ell^{2}$ distributional convergence \begin{displaymath} \sqrt{K_{n}}((P_{1,n},\,P_{2,n},\ldots)-p_α)\underset{n\rightarrow+\infty}{\overset{\cL}{\longrightarrow}}\mathcal{G}(Γ_α), \end{displaymath} where $\mathcal{G}(Γ_α)$ stands for a centered Gaussian process with covariance matrix $Γ_α=diag(p_α) - p_α p_α^T$. We apply our result to the estimation of the parameter

math.PR

Conformal Inference for Open-Set and Imbalanced Classification

This paper presents a conformal prediction method for classification in highly imbalanced and open-set settings, where there are many possible classes and not all may be represented in the data. Existing approaches require a finite, known label space and typically involve random sample splitting, which works well when there is a sufficient number of observations from each class. Consequently, they have two limitations: (i) they fail to provide adequate coverage when encountering new labels at test time, and (ii) they may become overly conservative when predicting previously seen labels. To obtain valid prediction sets in the presence of unseen labels, we compute and integrate into our predictions a new family of conformal p-values that can test whether a new data point belongs to a previously unseen class. We study these p-values theoretically, establishing their optimality, and uncover an intriguing connection with the classical Good--Turing estimator for the probability of observing a new species. To make more efficient use of imbalanced data, we also develop a selective sample splitting algorithm that partitions training and calibration data based on label frequency, leading to more informative predictions. Despite breaking exchangeability, this allows maintaining finite-sample guarantees through suitable re-weighting. With both simulated and real data, we demonstrate our method leads to prediction sets with valid coverage even in challenging open-set scenarios with infinite numbers of possible labels, and produces more informative predictions under extreme class imbalance.

stat.ML

Large-scale entity resolution via microclustering Ewens--Pitman random partitions

We introduce the microclustering Ewens--Pitman model for random partitions, obtained by scaling the strength parameter of the Ewens--Pitman model linearly with the sample size. The resulting random partition is shown to have the microclustering property, namely: the size of the largest cluster grows sub-linearly with the sample size, while the number of clusters grows linearly. By leveraging the interplay between the Ewens--Pitman random partition with the Pitman--Yor process, we develop efficient variational inference schemes for posterior computation in entity resolution. Our approach achieves a speed-up of three orders of magnitude over existing Bayesian methods for entity resolution, while maintaining competitive empirical performance.

stat.ME

Non-asymptotic approximations of Gaussian neural networks via second-order Poincaré inequalities

There is a recent and growing literature on large-width asymptotic and non-asymptotic properties of deep Gaussian neural networks (NNs), namely NNs with weights initialized as Gaussian distributions. For a Gaussian NN of depth $L\geq1$ and width $n\geq1$, it is well-known that, as $n\rightarrow+\infty$, the NN's output converges (in distribution) to a Gaussian process. Recently, some quantitative versions of this result, also known as quantitative central limit theorems (QCLTs), have been obtained, showing that the rate of convergence is $n^{-1}$, in the $2$-Wasserstein distance, and that such a rate is optimal. In this paper, we investigate the use of second-order Poincaré inequalities as an alternative approach to establish QCLTs for the NN's output. Previous approaches consist of a careful analysis of the NN, by combining non-trivial probabilistic tools with ad-hoc techniques that rely on the recursive definition of the network, typically by means of an induction argument over the layers, and it is unclear if and how they still apply to other NN's architectures. Instead, the use of second-order Poincaré inequalities rely only on the fact that the NN is a functional of a Gaussian process, reducing the problem of establishing QCLTs to the algebraic problem of computing the gradient and Hessian of the NN's output, which still applies to other NN's architectures. We show how our approach is effective in establishing QCLTs for the NN's output, though it leads to suboptimal rates of convergence. We argue that such a worsening in the rates is peculiar to second-order Poincaré inequalities, and it should be interpreted as the "cost" for having a straightforward, and general, procedure for obtaining QCLTs.

cs.LG