arXiv ScienceSearch

arXiv subjects

Fabrizio Leisen

Publications and source records attributed to Fabrizio Leisen.

At least 19 recordsLinked to original sources

Some cautionary tales about Bayesian predictive inference

Two misunderstandings, frequently arising in Bayesian predictive inference, are discussed. The first deals with the data generating mechanism, while the second consists in overestimating the role played by asymptotic exchangeability. Some consequences of such misunderstandings are highlighted through examples.

math.ST

Bayesian model selection of vine copulas: a loss-based perspective

The growing popularity of vine copulas in multivariate statistical analysis is largely driven by their ability to capture complex dependence structures. However, this flexibility comes at a cost, as the number of possible vine models grows rapidly and becomes intractable even in moderately low-dimensional settings. These limitations affect the practical applicability of current Bayesian inference and model selection approaches, effectively restricting it to problems of relatively small-dimension due to their high computational cost. This paper addresses the still open challenge of efficient model selection and estimation in Bayesian vine methodology. We propose a novel framework for Bayesian vine copula model selection that combines loss-based model priors with the shotgun stochastic search strategy. The strength of the proposed approach is twofold: it promotes sparsity and enables fast and effective structure selection. Furthermore, our comprehensive framework jointly identifies the vine structure, selects the copula families, and estimates the model parameters. The power of the proposed approach is demonstrated via simulation studies and an application to a real dataset of EFT portfolio asset returns.

stat.ME

Conformalized Super Learner

The Super Learner (SL) is a widely used ensemble method that combines point predictions from a library of learners based on their predictive performance. Interval predictions are of considerable practical interest because they allow uncertainty in predictions produced by an individual learner or an ensemble to be quantified. Several methods have been proposed for constructing interval predictions based on the SL, however, these approaches are typically justified using asymptotic arguments or rely on computationally intensive procedures such as the bootstrap. Conformal prediction (CP) is a machine learning framework for constructing prediction intervals with finite-sample and asymptotic coverage guarantees under mild conditions. We propose coupling CP with the SL through a natural construction that mirrors the original SL framework, using individual learner weights and combining learner-specific conformity scores via a weighted majority vote. We characterize the properties of the resulting SL-based prediction intervals for continuous outcomes. We cover settings under exchangeability, potential violations of exchangeability, and data-generating mechanisms exhibiting heteroscedasticity, sparsity, and other forms of distributional heterogeneity. A comprehensive simulation study shows that the conformalized SL achieves valid finite-sample coverage with competitive performance relative to the true data-generating mechanism. A central contribution of this work is an application to predicting creatinine levels using socio-demographic, biometric, and laboratory measurements. This example demonstrates the benefits of an ensemble with carefully selected learners designed to capture key aspects of complex regression functions, including non-linear effects, interactions, sparsity, heteroscedasticity, and robustness to outliers.

stat.ML

Conditional Copula models using loss-based Bayesian Additive Regression Trees

The study of dependence between random variables under external influences is a challenging problem in multivariate analysis. We address this by proposing a novel semi-parametric approach for conditional copula models using Bayesian additive regression trees (BART) models. BART is becoming a popular approach in statistical modelling due to its simple ensemble type formulation complemented by its ability to provide inferential insights. Although BART allows us to model complex functional relationships, it tends to suffer from overfitting. In this article, we exploit a loss-based prior for the tree topology that is designed to reduce the tree complexity. In addition, we propose a novel adaptive Reversible Jump Markov Chain Monte Carlo algorithm that is ergodic in nature and requires very few assumptions allowing us to model complex and non-smooth likelihood functions with ease. Moreover, we show that our method can efficiently recover the true tree structure and approximate a complex conditional copula parameter, and that our adaptive routine can explore the true likelihood region under a sub-optimal proposal variance. Lastly, we provide case studies concerning the effect of gross domestic product on the dependence between the life expectancies and literacy rates of the male and female populations of different countries.

stat.ME

Weak convergence of predictive distributions

Let $(X_n)$ be a sequence of random variables with values in a standard Borel space $S$. We investigate the condition \begin{gather}\label{x56w1q} E\bigl\{f(X_{n+1})\mid X_1,\ldots,X_n\bigr\}\,\quad\text{converges in probability,}\tag{*} \\\text{as }n\rightarrow\infty,\text{ for each bounded Borel function }f:S\rightarrow\mathbb{R}.\notag \end{gather} Some consequences of \eqref{x56w1q} are highlighted and various sufficient conditions for it are obtained. In particular, \eqref{x56w1q} is characterized in terms of stable convergence. Since \eqref{x56w1q} holds whenever $(X_n)$ is conditionally identically distributed, three weak versions of the latter condition are investigated as well. For each of such versions, our main goal is proving (or disproving) that \eqref{x56w1q} holds. Several counterexamples are given.

math.PR

Conformalized Regression for Continuous Bounded Outcomes

Regression problems with bounded continuous outcomes frequently arise in statistical and machine learning applications, such as the analysis of rates and proportions. A central challenge in this setting is predicting the response at a new covariate value. Most of the existing literature has focused either on point prediction or on interval prediction based on asymptotic approximations. We develop conformal prediction intervals for bounded outcomes within the framework of transformation regression models, encompassing widely used models such as beta regression and logit-normal regression. We construct non-conformity scores based on model-aligned residuals and identify a quantile-residual score that is particularly well suited to bounded outcomes, bridging normalized conformal prediction and distributional conformal prediction. This score accounts for both the heteroscedasticity inherent in such data and the asymmetry that emerges near the boundaries of the response space. We establish marginal validity and asymptotic conditional validity for both full and split conformal prediction, holding under model misspecification. A comprehensive simulation study confirms that both methods empirically attain valid finite-sample coverage, including cases under model misspecification. A real-data application demonstrates their practical performance against bootstrap-based alternatives.

stat.ML

Restricted mean survival times for comparing grouped survival data: a Bayesian nonparametric approach

Comparing survival experiences of different groups of data is an important issue in several applied problems. A typical example is where one wishes to investigate treatment effects. Here we propose a new Bayesian approach based on restricted mean survival times (RMST). A nonparametric prior is specified for the underlying survival functions: this extends the standard univariate neutral to the right processes to a multivariate setting and induces a prior for the RMST's. We rely on a representation as exponential functionals of compound subordinators to determine closed form expressions of prior and posterior mixed moments of RMST's. These results are used to approximate functionals of the posterior distribution of RMST's and are essential for comparing time--to--event data arising from different samples.

stat.ME

Asymptotics of predictive distributions driven by sample means and variances

Let $\alpha_n(\cdot)=P\bigl(X_{n+1}\in\cdot\mid X_1,\ldots,X_n\bigr)$ be the predictive distributions of a sequence $(X_1,X_2,\ldots)$ of $p$-dimensional random vectors. Suppose $$\alpha_n= \mathcal{N} _p (M_n,Q_n)$$ where $M_n=\frac{1}{n}\sum_{i=1}^nX_i$ and $Q_n=\frac{1}{n}\sum_{i=1}^n(X_i-M_n)(X_i-M_n)^t$. Then, there is a random probability measure $\alpha$ on the Borel subsets of $\mathbb{R}^p$ such that $\lVert\alpha_n-\alpha\rVert\overset{a.s.}\longrightarrow 0$ where $\lVert\cdot\rVert$ is total variation distance. An explicit expression for $\alpha$ is provided and the convergence rate of $\lVert\alpha_n-\alpha\rVert$ is shown to be arbitrarily close to $n^{-1/2}$. Moreover, it is still true that $\lVert\alpha_n-\alpha\rVert\overset{a.s.}\longrightarrow 0$ even if $\alpha_n=\mathcal{L}(M_n,Q_n)$ where $\mathcal{L}$ belongs to a class of distributions much larger than the normal. The predictives $\alpha_n$ are useful in various frameworks, including Bayesian predictive inference and predictive resampling. Finally, the asymptotic behavior of copula-based predictive distributions (introduced in [13]) is investigated and a numerical experiment is performed.

math.ST

Generating knockoffs via conditional independence

Let $X$ be a $p$-variate random vector and $\widetilde{X}$ a knockoff copy of $X$ (in the sense of \cite{CFJL18}). A new approach for constructing $\widetilde{X}$ (henceforth, NA) has been introduced in \cite{JSPI}. NA has essentially three advantages: (i) To build $\widetilde{X}$ is straightforward; (ii) The joint distribution of $(X,\widetilde{X})$ can be written in closed form; (iii) $\widetilde{X}$ is often optimal under various criteria. However, for NA to apply, $X_1,\ldots, X_p$ should be conditionally independent given some random element $Z$. Our first result is that any probability measure $\mu$ on $\mathbb{R}^p$ can be approximated by a probability measure $\mu_0$ of the form $$\mu_0\bigl(A_1\times\ldots\times A_p\bigr)=E\Bigl\{\prod_{i=1}^p P(X_i\in A_i\mid Z)\Bigr\}.$$ The approximation is in total variation distance when $\mu$ is absolutely continuous, and an explicit formula for $\mu_0$ is provided. If $X\sim\mu_0$, then $X_1,\ldots,X_p$ are conditionally independent. Hence, with a negligible error, one can assume $X\sim\mu_0$ and build $\widetilde{X}$ through NA. Our second result is a characterization of the knockoffs $\widetilde{X}$ obtained via NA. It is shown that $\widetilde{X}$ is of this type if and only if the pair $(X,\widetilde{X})$ can be extended to an infinite sequence so as to satisfy certain invariance conditions. The basic tool for proving this fact is de Finetti's theorem for partially exchangeable sequences. In addition to the quoted results, an explicit formula for the conditional distribution of $\widetilde{X}$ given $X$ is obtained in a few cases. In one of such cases, it is assumed $X_i\in\{0,1\}$ for all $i$.

math.ST

A probabilistic view on predictive constructions for Bayesian learning

Given a sequence $X=(X_1,X_2,\ldots)$ of random observations, a Bayesian forecaster aims to predict $X_{n+1}$ based on $(X_1,\ldots,X_n)$ for each $n\ge 0$. To this end, in principle, she only needs to select a collection $\sigma=(\sigma_0,\sigma_1,\ldots)$, called ``strategy" in what follows, where $\sigma_0(\cdot)=P(X_1\in\cdot)$ is the marginal distribution of $X_1$ and $\sigma_n(\cdot)=P(X_{n+1}\in\cdot\mid X_1,\ldots,X_n)$ the $n$-th predictive distribution. Because of the Ionescu-Tulcea theorem, $\sigma$ can be assigned directly, without passing through the usual prior/posterior scheme. One main advantage is that no prior probability is to be selected. In a nutshell, this is the predictive approach to Bayesian learning. A concise review of the latter is provided in this paper. We try to put such an approach in the right framework, to make clear a few misunderstandings, and to provide a unifying view. Some recent results are discussed as well. In addition, some new strategies are introduced and the corresponding distribution of the data sequence $X$ is determined. The strategies concern generalized P\'olya urns, random change points, covariates and stationary sequences.

stat.ME

Kernel based Dirichlet sequences

Let $X=(X_1,X_2,\ldots)$ be a sequence of random variables with values in a standard space $(S,\mathcal{B})$. Suppose \begin{gather*} X_1\sim\nu\quad\text{and}\quad P\bigl(X_{n+1}\in\cdot\mid X_1,\ldots,X_n\bigr)=\frac{\theta\nu(\cdot)+\sum_{i=1}^nK(X_i)(\cdot)}{n+\theta}\quad\quad\text{a.s.} \end{gather*} where $\theta>0$ is a constant, $\nu$ a probability measure on $\mathcal{B}$, and $K$ a random probability measure on $\mathcal{B}$. Then, $X$ is exchangeable whenever $K$ is a regular conditional distribution for $\nu$ given any sub-$\sigma$-field of $\mathcal{B}$. Under this assumption, $X$ enjoys all the main properties of classical Dirichlet sequences, including Sethuraman's representation, conjugacy property, and convergence in total variation of predictive distributions. If $\mu$ is the weak limit of the empirical measures, conditions for $\mu$ to be a.s. discrete, or a.s. non-atomic, or $\mu\ll\nu$ a.s., are provided. Two CLT's are proved as well. The first deals with stable convergence while the second concerns total variation distance.

math.PR

Bayesian predictive inference without a prior

Let $(X_n:n\ge 1)$ be a sequence of random observations. Let $\sigma_n(\cdot)=P\bigl(X_{n+1}\in\cdot\mid X_1,\ldots,X_n\bigr)$ be the $n$-th predictive distribution and $\sigma_0(\cdot)=P(X_1\in\cdot)$ the marginal distribution of $X_1$. In a Bayesian framework, to make predictions on $(X_n)$, one only needs the collection $\sigma=(\sigma_n:n\ge 0)$. Because of the Ionescu-Tulcea theorem, $\sigma$ can be assigned directly, without passing through the usual prior/posterior scheme. One main advantage is that no prior probability has to be selected. In this paper, $\sigma$ is subjected to two requirements: (i) The resulting sequence $(X_n)$ is conditionally identically distributed, in the sense of Berti, Pratelli and Rigo (2004); (ii) Each $\sigma_{n+1}$ is a simple recursive update of $\sigma_n$. Various new $\sigma$ satisfying (i)-(ii) are introduced and investigated. For such $\sigma$, the asymptotics of $\sigma_n$, as $n\rightarrow\infty$, is determined. In some cases, the probability distribution of $(X_n)$ is also evaluated.

math.ST

New perspectives on knockoffs construction

Let $\Lambda$ be the collection of all probability distributions for $(X,\widetilde{X})$, where $X$ is a fixed random vector and $\widetilde{X}$ ranges over all possible knockoff copies of $X$ (in the sense of \cite{CFJL18}). Three topics are developed in this paper: (i) A new characterization of $\Lambda$ is proved; (ii) A certain subclass of $\Lambda$, defined in terms of copulas, is introduced; (iii) The (meaningful) special case where the components of $X$ are conditionally independent is treated in depth. In real problems, after observing $X=x$, each of points (i)-(ii)-(iii) may be useful to generate a value $\widetilde{x}$ for $\widetilde{X}$ conditionally on $X=x$.

math.ST

A Copula-based Fully Bayesian Nonparametric Evaluation of Cardiovascular Risk Markers in the Mexico City Diabetes Study

Cardiovascular disease lead the cause of death world wide and several studies have been carried out to understand and explore cardiovascular risk markers in normoglycemic and diabetic populations. In this work, we explore the association structure between hyperglycemic markers and cardiovascular risk markers controlled by triglycerides, body mass index, age and gender, for the normoglycemic population in The Mexico City Diabetes Study. Understanding the association structure could contribute to the assessment of additional cardiovascular risk markers in this low income urban population with a high prevalence of classic cardiovascular risk biomarkers. The association structure is measured by conditional Kendall's tau, defined through conditional copula functions. The latter are in turn modeled under a fully Bayesian nonparametric approach, which allows the complete shape of the copula function to vary for different values of the controlled covariates.

stat.AP

Completely Random Measures and L\'evy Bases in Free probability

This paper develops a theory for completely random measures in the framework of free probability. A general existence result for free completely random measures is established, and in analogy to the classical work of Kingman it is proved that such random measures can be decomposed into the sum of a purely atomic part and a (freely) infinitely divisible part. The latter part (termed a free L\'evy basis) is studied in detail in terms of the free L\'evy-Khintchine representation and a theory parallel to the classical work of Rajput and Rosinski is developed. Finally a L\'evy-It\^o type decomposition for general free L\'evy bases is established.

math.PR

Compound vectors of subordinators and their associated positive L\'evy copulas

L\'evy copulas are an important tool which can be used to build dependent L\'evy processes. In a classical setting, they have been used to model financial applications. In a Bayesian framework they have been employed to introduce dependent nonparametric priors which allow to model heterogeneous data. This paper focuses on introducing a new class of L\'evy copulas based on a class of subordinators recently appeared in the literature, called \textit{Compound Random Measures}. The well-known Clayton L\'evy copula is a special case of this new class. Furthermore, we provide some novel results about the underlying vector of subordinators such as a series representation and relevant moments. The article concludes with an application to a Danish fire dataset.

stat.ME

A P\'olya-Gamma Sampler for a Generalized Logistic Regression

In this paper we introduce a novel Bayesian data augmentation approach for estimating the parameters of the generalised logistic regression model. We propose a P\'olya-Gamma sampler algorithm that allows us to sample from the exact posterior distribution, rather than relying on approximations. A simulation study illustrates the flexibility and accuracy of the proposed approach to capture heavy and light tails in binary response data of different dimensions. The methodology is applied to two different real datasets, where we demonstrate that the P\'olya-Gamma sampler provides more precise estimates than the empirical likelihood method, outperforming approximate approaches.

stat.ME

On a flexible construction of a negative binomial model

This work presents a construction of stationary Markov models with negative-binomial marginal distributions. A simple closed form expression for the corresponding transition probabilities is given, linking the proposal to well-known classes of birth and death processes and thus revealing interesting characterizations. The advantage of having such closed form expressions is tested on simulated and real data.

stat.ME