arXiv ScienceSearch

arXiv subjects

Luis Pericchi

Publications and source records attributed to Luis Pericchi.

7 recordsLinked to original sources

Bounds on Intrinsic Bayes Factors and Least Favorable Intrinsic Priors for General Statistical Hypothesis Testing

Hypothesis Testing is the most contentious procedure in statistical Methodology. P values rejects Null Hypotheses far too easily, specially for large samples. On the other hand, Bayes Factors depends on assumptions, for example regarding Intrinsic Bayes Factors, which average? Arithmetic, Geometric, Median? Our bound is the infimum over all the averages. We develop a lower bound on Intrinsic Bayes Factors that adjust authomatically with the sample size. Furthermore, we introduce the new idea of {\it{\textbf{Least Favorable Intrinsic Prior}}}, which corresponds to the least favourable possible training samples. The bound sets a bridge between Intrinsic Bayes Factors and Adrian Smith and David Spiegelhalter methodology.

math.ST

Objective Model Prior Probabilities in Variable Selection

For many years it was routine to use equal model prior probabilities in Bayesian model uncertainty analysis. At least twenty years ago it became clear that this was problematic, leading to support of much too large models in the increasingly huge model spaces being considered in genomics and other fields. A popular replacement was to adopt a suggestion of Harold Jeffreys for the variable selection problem in which a total of $k$ possible variables are being considered for inclusion in the model: give the collection of all models containing $d$ variables ($d = 0, . . . , k$) prior probability $1/(k + 1)$ and then divide this prior probability equally among the models in the collection. Many other choices of model prior probabilities that impose severe parsimony have also been introduced. We begin by reviewing the problems with using equal model prior probabilities and then discuss some serious problems with the Jeffreys choice. Finally, we introduce and study a number of objective alternative choices of model prior probabilities, from both numerical and theoretical perspectives.

stat.ME

A Bridge between Cross-validation Bayes Factors and Geometric Intrinsic Bayes Factors

Model Selections in Bayesian Statistics are primarily made with statistics known as Bayes Factors, which are directly related to Posterior Probabilities of models. Bayes Factors require a careful assessment of prior distributions as in the Intrinsic Priors of Berger and Pericchi (1996a) and integration over the parameter space, which may be highly dimensional. Recently researchers have been proposing alternatives to Bayes Factors that require neither integration nor specification of priors. These developments are still in a very early stage and are known as Prior-free Bayes Factors, Cross-Validation Bayes Factors (CVBF), and Bayesian "Stacking." This kind of method and Intrinsic Bayes Factor (IBF) both avoid the specification of prior. However, this Prior-free Bayes factor might need a careful choice of a training sample size. In this article, a way of choosing training sample sizes for the Prior-free Bayes factor based on Geometric Intrinsic Bayes Factors (GIBFs) is proposed and studied. We present essential examples with a different number of parameters and study the statistical behavior both numerically and theoretically to explain the ideas for choosing a feasible training sample size for Prior-free Bayes Factors. We put forward the "Bridge Rule" as an assignment of a training sample size for CVBF's that makes them close to Geometric IBFs. We conclude that even though tractable Geometric IBFs are preferable, CVBF's, using the Bridge Rule, are useful and economical approximations to Bayes Factors.

stat.ME

Changing the paradigm of fixed significance levels: Testing Hypothesis by Minimizing Sum of Errors Type I and Type II

Our purpose, is to put forward a change in the paradigm of testing by generalizing a very natural idea exposed by Morris DeGroot (1975) aiming to an approach that is attractive to all schools of statistics, in a procedure better suited for the needs of science. DeGroot's seminal idea is to base testing statistical hypothesis on minimizing the weighted sum of type I plus type II error instead of of the prevailing paradigm which is fixing type I error and minimizing type II error. DeGroot's result is that in simple vs simple hypothesis the optimal criterion is to reject, according to the likelihood ratio as the evidence (ordering) statistics using a fixed threshold value, instead of a fixed tail probability. By defining expected type I and type II errors, we generalize DeGroot's approach and find that the optimal region is defined by the ratio of evidences, that is, averaged likelihoods (with respect to a prior measure) and a threshold fixed. This approach yields an optimal theory in complete generality, which the Classical Theory of Testing does not. This can be seen as a Bayes-Non-Bayes compromise: the criteria (weighted sum of type I and type II errors) is Frequentist, but the test criterion is the ratio of marginalized likelihood, which is Bayesian. We give arguments, to push the theory still further, so that the weighting measures (priors)of the likelihoods does not have to be proper and highly informative, but just predictively matched, that is that predictively matched priors, give rise to the same evidence (marginal likelihoods) using minimal (smallest) training samples. The theory that emerges, similar to the theories based on Objective Bayes approaches, is a powerful response to criticisms of the prevailing approach of hypothesis testing, see for example Ioannidis (2005) and Siegfried (2010) among many others.

stat.ME

A Robust Bayesian Dynamic Linear Model for Latin-American Economic Time Series: "The Mexico and Puerto Rico Cases"

The traditional time series methodology requires at least a preliminary transformation of the data to get stationarity. On the other hand, Robust Bayesian Dynamic Models (RBDMs) do not assume a regular pattern or stability of the underlying system but can include points of statement breaks. In this paper we use RBDMs in order to account possible outliers and structural breaks in Latin-American economic time series. We work with important economic time series from Puerto Rico and Mexico. We show by using a random walk model how RBDMs can be applied for detecting historic changes in the economic inflation of Mexico. Also, we model the Consumer Price Index (CPI), the Economic Activity Index (EAI) and the total number of employments (TNE) economic time series in Puerto Rico using local linear trend and seasonal RBDMs with observational and states variances. The results illustrate how the model accounts the structural breaks for the historic recession periods in Puerto Rico.

stat.ME

Quick Anomaly Detection by the Newcomb--Benford Law, with Applications to Electoral Processes Data from the USA, Puerto Rico and Venezuela

A simple and quick general test to screen for numerical anomalies is presented. It can be applied, for example, to electoral processes, both electronic and manual. It uses vote counts in officially published voting units, which are typically widely available and institutionally backed. The test examines the frequencies of digits on voting counts and rests on the First (NBL1) and Second Digit Newcomb--Benford Law (NBL2), and in a novel generalization of the law under restrictions of the maximum number of voters per unit (RNBL2). We apply the test to the 2004 USA presidential elections, the Puerto Rico (1996, 2000 and 2004) governor elections, the 2004 Venezuelan presidential recall referendum (RRP) and the previous 2000 Venezuelan Presidential election. The NBL2 is compellingly rejected only in the Venezuelan referendum and only for electronic voting units. Our original suggestion on the RRP (Pericchi and Torres, 2004) was criticized by The Carter Center report (2005). Acknowledging this, Mebane (2006) and The Economist (US) (2007) presented voting models and case studies in favor of NBL2. Further evidence is presented here. Moreover, under the RNBL2, Mebane's voting models are valid under wider conditions. The adequacy of the law is assessed through Bayes Factors (and corrections of $p$-values) instead of significance testing, since for large sample sizes and fixed $α$ levels the null hypothesis is over rejected. Our tests are extremely simple and can become a standard screening that a fair electoral process should pass.

stat.ME