arXiv ScienceSearch

arXiv subjects

Mikhail Chebunin

Publications and source records attributed to Mikhail Chebunin.

13 recordsLinked to original sources

Limit theorems for a class of random outer measures in infinite urn schemes

An urn scheme is a probabilistic model in which balls are placed into urns sequentially and independently of each other. All balls share the same probability distribution for hitting the urns. In the simplest case, there is a finite number of urns and the probabilities of hitting each urn are equal. In an infinite urn scheme, there is a countable number of urns, and the hitting probabilities form a probability mass function on the set of urn labels, so they depend on the urn number. The statistics of interest is the number of urns with at least k balls after throwing n balls. Thus, we assume that there is a countable family of urns, and we fix the probabilities for a ball to hit each urn (the same for all balls). For an arbitrary subset A of the unit interval [0, 1], we do not consider all ball indices from 1 to n, but only those that belong to the set nA, and we study the number of urns with at least k balls after throwing the balls with indices in nA. This number is non-negative, and if the set A is empty, this number is equal to zero. Moreover, if k=1, then it satisfies the property of countable subadditivity. Hence, the number of non-empty urns for ball indices in nA, where A is an arbitrary subset of the unit interval, satisfies all the axioms of an outer measure on the unit interval. We study the properties of the statistics of interest. Our main result is a functional central limit theorem for sets A consisting of finite unions of intervals and parameterized by their boundary points. We consider applications of this theorem to elementary probabilistic models of text.

math.PR

On strong sharp phase transition in the random connection model

We consider a random connection model (RCM) $\xi$ driven by a Poisson process $\eta$. We derive exponential moment bounds for an arbitrary cluster, provided that the intensity $t$ of $\eta$ is below a certain critical intensity $t_T$. The associated subcritical regime is characterized by a finite mean cluster size, uniformly in space. Under an exponential decay assumption on the connection function, we also show that the cluster diameters are exponentially small as well. In the important stationary marked case and under a uniform moment bound on the connection function, we show that $t_T$ coincides with $t_c$, the largest $t$ for which $\xi$ does not percolate. In this case, we also derive some percolation mean field bounds. These findings generalize some of the recent results. Even in the classical unmarked case, our results are more general than what has been previously known. Our proofs are partially based on some stochastic monotonicity properties, which might be of interest in their own right.

math.PR

On the uniqueness of the infinite cluster and the cluster density in the Poisson driven random connection model

We consider a random connection model (RCM) on a general space driven by a Poisson process whose intensity measure is scaled by a parameter $t\ge 0$. We say that the infinite clusters are deletion stable if the removal of a Poisson point cannot split a cluster in two or more infinite clusters. We prove that this stability together with a natural irreducibility assumption implies uniqueness of the infinite cluster. Conversely, if the infinite cluster is unique then this stability property holds. Several criteria for irreducibility will be established. We also study the analytic properties of expectations of functions of clusters as a function of $t$. In particular we show that the position dependent cluster density is differentiable. A significant part of this paper is devoted to the important case of a stationary marked RCM (in Euclidean space), containing the Boolean model with general compact grains and the so-called weighted RCM as special cases. In this case we establish differentiability and a convexity property of the cluster density $\kappa(t)$. These properties are crucial for our proof of deletion stability of the infinite clusters but are also of interest in their own right. It then follows that an irreducible stationary marked RCM can have at most one infinite cluster. This extends and unifies several results in the literature.

math.PR

Limit theorems for forward and backward processes of numbers of non-empty urns in infinite urn schemes

We study the joint asymptotics of forward and backward processes of numbers of non-empty urns in an infinite urn scheme. The probabilities of balls hitting the urns are assumed to satisfy the conditions of regular decrease. We prove weak convergence to a two-dimensional Gaussian process. Its covariance function depends only on exponent of regular decrease of probabilities. We obtain parameter estimates that have a normal asymototics for its joint distribution together with forward and backward processes. We use these estimates to construct statistical tests for the homogeneity of the urn scheme on the number of thrown balls.

math.PR

Asymptotics of joint orderings of compound Poisson fields

We are developing a new method for the analysis of queuing systems with heterogeneous in time and space compound (marked) Poisson input flow. The state space of the input flow is embedded in a higher-dimensional space with a homogeneous marked Poisson field on it. We prove limit theorems for partial sums of marks under the ordering of field points by coordinates.

math.PR

Asymptotics of sums of regression residuals under multiple ordering of regressors

We prove theorems about the Gaussian asymptotics of an empirical bridge built from linear model regressors with multiple regressor ordering. We study the testing of the hypothesis of a linear model for the components of a random vector: one of the components is a linear combination of the others up to an error that does not depend on the other components of the random vector. The results of observations of independent copies of a random vector are sequentially ordered in ascending order of several of its components. The result is a sequence of vectors of higher dimension, consisting of induced order statistics (concomitants) corresponding to different orderings. For this sequence of vectors, without the assumption of a linear model for the components, we prove a lemma of weak convergence of the distributions of an appropriately centered and normalized process to a centered Gaussian process with almost surely continuous trajectories. Assuming a linear relationship of the components, standard least squares estimates are used to compute regression residuals -- the differences between response values and the predicted ones by the linear model. We prove a theorem of weak convergence of the process of regression residuals under the necessary normalization to a centered Gaussian process.

math.ST

Modifications of Simon text model

We discuss probability text models and their modifications. We construct processes of different and unique words in a text. The models are to correspond to the real text statistics. The infinite urn model (Karlin model) and the Simon model are the most known models of texts, but they do not give the ability to simulate the number of unique words correctly. The infinite urn model give sometimes the incorrect limit of the relative number of unique and different words. The Simon model states a linear growth of the numbers of different and unique words. We propose three modifications of the Karlin and Simon models. The first one is the offline variant, the Simon model starts after the completion of the infinite urn scheme. We prove limit theorems for this modification in embedded times only. The second modification involves the compound Poisson process in the infinite urn model. We prove limit theorems for it. The third modification is the online variant, the Simon redistribution works at any toss of the Karlin model. In contrast to the compound Poisson model, we have no analytics for this modification. We test all the modifications by the simulation and have a good correspondence to the real texts.

math.ST

A statistical test for correspondence of texts to the Zipf-Mandelbrot law

We analyse correspondence of a text to a simple probabilistic model. The model assumes that the words are selected independently from an infinite dictionary. The probability distribution correspond to the Zipf---Mandelbrot law. We count sequentially the numbers of different words in the text and get the process of the numbers of different words. Then we estimate Zipf---Mandelbrot law parameters using the same sequence and construct an estimate of the expectation of the number of different words in the text. Then we subtract the corresponding values of the estimate from the sequence and normalize along the coordinate axes, obtaining a random process on a segment from 0 to 1. We prove that this process (the empirical text bridge) converges weakly in the uniform metric on $C (0,1)$ to a centered Gaussian process with continuous a.s. paths. We develop and implement an algorithm for approximate calculation of eigenvalues of the covariance function of the limit Gaussian process, and then an algorithm for calculating the probability distribution of the integral of the square of this process. We use the algorithm to analyze uniformity of texts in English, French, Russian and Chinese.

math.ST

Functional central limit theorems for occupancies and missing mass process in infinite urn models

We study the infinite urn scheme when the balls are sequentially distributed over an infinite number of urns labelled 1,2,... so that the urn $j$ at every draw gets a ball with probability $p_j$, $\sum_j p_j=1$. We prove functional central limit theorems for discrete time and the poissonised version for the urn occupancies process, for the odd-occupancy and for the missing mass processes extending the known non-functional central limit theorems.

math.PR

A statistical test for the Zipf's law by deviations from the Heaps' law

We explore a probabilistic model of an artistic text: words of the text are chosen independently of each other in accordance with a discrete probability distribution on an infinite dictionary. The words are enumerated 1, 2, $\ldots$, and the probability of appearing the $i$'th word is asymptotically a power function. Bahadur proved that in this case the number of different words depends on the length of the text is asymptotically a power function, too. On the other hand, in the applied statistics community, there exist statements supported by empirical observations, the Zipf's and the Heaps' laws. We highlight the links between Bahadur results and Zipf's/Heaps' laws, and introduce and analyse a corresponding statistical test.

math.ST

Asymptotically normal estimators for Zipf's law

Zipf's law states that sequential frequencies of words in a text correspond to a power function. Its probabilistic model is an infinite urn scheme with asymptotically power distribution. The exponent of this distribution must be estimated. We use the number of different words in a text and similar statistics to construct asymptotically normal estimators of the exponent.

math.ST