arXiv ScienceSearch

arXiv subjects

Roman Guchenko

Publications and source records attributed to Roman Guchenko.

6 recordsLinked to original sources

Priority-preserving augmentation of goodness-of-fit tests by conditional calibration

An omnibus goodness-of-fit statistic can fail to exploit simple diagnostic evidence efficiently. For example, under a standard normal null, an exceptionally large observation is direct evidence of a scale or tail departure even when the primary omnibus statistic does not cross its critical value. This motivates a simple augmentation principle: retain an established primary test, but reserve a small part of its null rejection budget for secondary statistics that encode natural features such as variation, asymmetry, or tail behavior. We implement this principle by calibrating each secondary acceptance region under the null conditional on acceptance at all preceding stages. Once calibrated, the resulting test is a fixed rectangular acceptance rule; the ordering is a mechanism for choosing its boundaries and assigning ordered first-rejection contributions, not a sequential-sampling scheme. An unconditional stage-budget parameterization makes the central trade-off explicit: additional sensitivity is purchased by removing a prespecified, usually small, amount of rejection probability from the primary stage. We establish strong consistency of the quantile-based Monte Carlo calibration and give an observation-specific pooled-rank version with exact randomized finite-m null size. In an experiment under a standard normal null with n=10, we augment the Kolmogorov--Smirnov statistic with sample variance and sample skewness. Assigning only 0.75% of the total Type I error budget to the two secondary statistics changes power against N(0.5,1) from 0.2742 to 0.2733, while increasing power against N(0,0.2^2) from 0.4693 to 0.9827. This focused experiment is a proof of principle rather than an exhaustive comparison of normality tests: a small unconditional allocation to prespecified diagnostics can greatly broaden power while preserving almost all of the primary test's power.

stat.ME

Goodness of Fit Tests Based on Joint Densities of Multiple Sample Statistics

We propose goodness-of-fit tests based on simulated confidence sets for joint distributions of multiple sample statistics, focusing on absolutely continuous null distributions with known parameters. One class of tests uses hyperrectangular confidence sets for principal components of order statistics and related statistic vectors. Extending earlier work on horizontal and vertical confidence bands for cumulative distribution functions, these tests are compared with some classical, Zhang, and related graphical tests. Simulations show that the proposed procedures are competitive with, and often more powerful than, existing methods. We also study the geometry of principal-component-based statistics; under a normal null distribution, the first principal component corresponds to the sample mean, while the second is related to a linear analogue of variance. A second class of tests uses confidence sets of arbitrary shape constructed through highest density regions. Unlike earlier kernel-density-based approaches, we use a k-nearest-neighbor method for detecting highest density regions, which is better suited to higher-dimensional statistic vectors. We study tests based on order statistics, empirical distribution function values, moments, and combinations of classical goodness-of-fit statistics. The resulting procedures are powerful against a wide range of alternatives. We also outline a two-sample extension via permutation tests based on joint distributions of several statistics and compare moment-based versions with energy-distance permutation tests. Finally, we discuss transformations other than the probability integral transform, showing that mapping data to another target distribution, such as the standard normal, can be advantageous when powerful tests are available for that distribution.

stat.ME

Statistical Modelling for Improving Efficiency of Online Advertising

Real-time bidding has transformed the digital advertising landscape, allowing companies to buy website advertising space in a matter of milliseconds in the time it takes a webpage to load. Joint research between Cardiff University and Crimtan has employed statistical modelling in conjunction with machine-learning techniques on big data to develop computer algorithms that can select the most appropriate person to which an ad should be shown. These algorithms have been used to identify suitable bidding strategies for that particular advert in order to make the whole process as profitable as possible for businesses. Crimtan's use of the algorithms have enabled them to improve the service that they offer to clients, save money, make significant efficiency gains and attract new business. This has had a knock-on effect with the clients themselves, who have reported an increase in conversion rates as a result of more targeted, accurate and informed advertising. We have also used mixed Poisson processes for modelling for analysing repeat-buying behaviour of online customers. To make numerical comparisons, we use real data collected by Crimtan in the process of running several recent ad campaigns.

stat.AP

Optimal discrimination designs for semi-parametric models

Much of the work in the literature on optimal discrimination designs assumes that the models of interest are fully specified, apart from unknown parameters in some models. Recent work allows errors in the models to be non-normally distributed but still requires the specification of the mean structures. This research is motivated by the interesting work of Otsu (2008) to discriminate among semi-parametric models by generalizing the KL-optimality criterion proposed by L\'opez-Fidalgo et al. (2007) and Tommasi and L\'opez-Fidalgo (2010). In our work we provide further important insights in this interesting optimality criterion. In particular, we propose a practical strategy for finding optimal discrimination designs among semi-parametric models that can also be verified using an equivalence theorem. In addition, we study properties of such optimal designs and identify important cases where the proposed semi-parametric optimal discrimination designs coincide with the celebrated T -optimal designs.

stat.ME

Efficient computation of Bayesian optimal discriminating designs

An efficient algorithm for the determination of Bayesian optimal discriminating designs for competing regression models is developed, where the main focus is on models with general distributional assumptions beyond the "classical" case of normally distributed homoscedastic errors. For this purpose we consider a Bayesian version of the Kullback- Leibler (KL) optimality criterion introduced by L\'opez-Fidalgo et al. (2007). Discretizing the prior distribution leads to local KL-optimal discriminating design problems for a large number of competing models. All currently available methods either require a large computation time or fail to calculate the optimal discriminating design, because they can only deal efficiently with a few model comparisons. In this paper we develop a new algorithm for the determination of Bayesian optimal discriminating designs with respect to the Kullback-Leibler criterion. It is demonstrated that the new algorithm is able to calculate the optimal discriminating designs with reasonable accuracy and computational time in situations where all currently available procedures are either slow or fail.

stat.CO

Bayesian T-optimal discriminating designs

The problem of constructing Bayesian optimal discriminating designs for a class of regression models with respect to the T-optimality criterion introduced by Atkinson and Fedorov (1975a) is considered. It is demonstrated that the discretization of the integral with respect to the prior distribution leads to locally T-optimal discrimination designs can only deal with a few comparisons, but the discretization of the Bayesian prior easily yields to discrimination design problems for more than 100 competing models. A new efficient method is developed to deal with problems of this type. It combines some features of the classical exchange type algorithm with the gradient methods. Convergence is proved and it is demonstrated that the new method can find Bayesian optimal discriminating designs in situations where all currently available procedures fail.

stat.ME