arXiv ScienceSearch

arXiv subjects

Ariel Avital

Publications and source records attributed to Ariel Avital.

4 recordsLinked to original sources

Total Variation Distance between Product Distributions: an Analytic Proxy

We characterize, up to universal constants, the total variation distance between finite products of arbitrary probability measures by a simple formula. A theorem of Latala reduces this expression to one scalar equation. As an application, we characterize the sample complexity of equal-prior binary hypothesis testing uniformly in the weak-detection regime, where squared Hellinger distance alone does not determine the answer, and also characterize the TV distance between multinomial distributions.

math.PR

TV over Bernoulli products: the small parameter regime

We study the total variation distance (TV) between two $n$-fold Bernoulli product measures parametrized by $\vec p=(p_1,\ldots,p_n)$ and $\vec q=(q_1,\ldots,q_n)$, respectively, in the \emph{tiny} and \emph{small} regimes. In the tiny regime, we have $p_i,q_i\lesssim 1/n^2$, and in the small regime, $p_i,q_i\lesssim 1/n$. We discover that in the tiny regime, the TV distance behaves as $\|\vec p-\vec q\|_1$, while in the small regime, it behaves as \[ \sum_{i=1}^n \Big| p_i\prod_{j\neq i}(1-p_j) - q_i\prod_{j\neq i}(1-q_j) \Big|, \] both up to absolute constants. Along the way we discover some identities of possible independent interest.

math.PR

Sharp bounds on aggregate expert error

We revisit the classic problem of aggregating binary advice from conditionally independent experts, also known as the Naive Bayes setting. Our quantity of interest is the error probability of the optimal decision rule. In the case of symmetric errors (sensitivity = specificity), reasonably tight bounds on the optimal error probability are known. In the general asymmetric case, we are not aware of any nontrivial estimates on this quantity. Our contribution consists of sharp upper and lower bounds on the optimal error probability in the general case, which recover and sharpen the best known results in the symmetric special case. Since this turns out to be equivalent to estimating the total variation distance between two product distributions, our results also have bearing on this important and challenging problem.

math.PR

Non-parametric Binary regression in metric spaces with KL loss

We propose a non-parametric variant of binary regression, where the hypothesis is regularized to be a Lipschitz function taking a metric space to [0,1] and the loss is logarithmic. This setting presents novel computational and statistical challenges. On the computational front, we derive a novel efficient optimization algorithm based on interior point methods; an attractive feature is that it is parameter-free (i.e., does not require tuning an update step size). On the statistical front, the unbounded loss function presents a problem for classic generalization bounds, based on covering-number and Rademacher techniques. We get around this challenge via an adaptive truncation approach, and also present a lower bound indicating that the truncation is, in some sense, necessary.

cs.LG