arXiv Science⌕ Search

arXiv · 2609.38412

Local polynomial density ratio estimation

Abstract

We propose a novel local-polynomial estimator of the ratio $r=f/g$ of two $d$-dimensional densities $f$ and $g$, from which independent samples are available. The estimator is shown to achieve pointwise minimax optimal rates over Hölder classes of arbitrary smoothness index without additional logarithmic factors, and with smoothness being only assumed of $r$ but not of $f$ nor $g$. Our analysis remains valid for points on the boundary of the support of the distribution associated to $g$ under a mild geometric assumption on the (unknown) boundary. We also derive a concentration inequality for the estimator, which can be useful in applications to classification, and give a rate in the supremum norm, where the supremum is also taken over boundary support points. The smoothness class over which the rate is obtained is sufficiently large that the individual densities $f$ and $g$ cannot be consistently estimated uniformly over this class. Direct estimators of the partial derivatives of $r$ together with rates of convergence are also provided. We also obtain asymptotic normality of the estimators under quite general assumptions and again including boundary points, with a consistent variance estimator allowing for data-driven studentization. Moreover, we show how our estimator can be used to estimate the Kullback-Leibler information, in a construction which additionally uses debiasing. We provide the parametric rate together with asymptotic normality under sufficient smoothness of $r$ relative to the dimension $d$.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hajo Holzmann, Alexander Meister. 2026-09-29. Local polynomial density ratio estimation. https://arxiv.org/abs/2609.38412

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

KL Convergence Guarantees for Score diffusion models under minimal data assumptions

Diffusion models are a new class of generative models that revolve around the estimation of the score function associated with a stochastic differential equation. Subsequent to its acquisition, the approximated score function is then harnessed to simulate the corresponding time-reversal process, ultimately enabling the generation of approximate data samples. Despite their evident practical significance these models carry, a notable challenge persists in the form of a lack of comprehensive quantitative results, especially in scenarios involving non-regular scores and estimators. In almost all reported bounds in Kullback Leibler (KL) divergence, it is assumed that either the score function or its approximation is Lipschitz uniformly in time. However, this condition is very restrictive in practice or appears to be difficult to establish. To circumvent this issue, previous works mainly focused on establishing convergence bounds in KL for an early stopped version of the diffusion model and a smoothed version of the data distribution, or assuming that the data distribution is supported on a compact manifold. These explorations have led to interesting bounds in either Wasserstein or Fortet-Mourier metrics. However, the question remains about the relevance of such early-stopping procedure or compactness conditions. In particular, if there exist a natural and mild condition ensuring explicit and sharp convergence bounds in KL. In this article, we tackle the aforementioned limitations by focusing on score diffusion models with fixed step size stemming from the Ornstein-Uhlenbeck semigroup and its kinetic counterpart. Our study provides a rigorous analysis, yielding simple, improved and sharp convergence bounds in KL applicable to any data distribution with finite Fisher information with respect to the standard Gaussian distribution.

math.ST↗

Estimation of conditional inequality curves and measures via estimating the conditional quantile function

In the paper conditional inequality curves and measures are proposed which allow us to describe the inequality/concentration of the conditional distribution of the feature we are interested in with respect to certain continuous variables. Moreover, for a graphical illustration of the change in values of the proposed conditional indices, a curve of conditional inequality measures is introduced. To estimate the curves and measures, a new method is proposed to estimate the conditional quantile function. This method uses quantile regression estimates for a given set of quantile orders, followed by isotonic regression on the estimated regression coefficients to ensure that the estimated conditional quantile function is nondecreasing. The consistency of the proposed estimators is proved while their finite sample performance is evaluated through simulation studies and compared with existing approaches. Finally, practical application of conditional curves and measures is demonstrated by determining estimated curves of conditional salary inequalities with respect to years of experience in different employee tenure groups, based on some real data. The code used to prepare the simulation results presented in this paper is available in a dedicated GitHub repository.

math.ST↗

Tracy-Widom, Gaussian, and Bootstrap: Approximations for Leading Eigenvalues in High-Dimensional PCA

Under certain conditions, the largest eigenvalue of a sample covariance matrix undergoes a well-known phase transition when the sample size $n$ and data dimension $p$ diverge proportionally. In the subcritical regime, this eigenvalue has fluctuations of order $n^{-2/3}$ that can be approximated by a Tracy-Widom distribution, while in the supercritical regime, it has fluctuations of order $n^{-1/2}$ that can be approximated with a Gaussian distribution. However, the statistical problem of determining which regime underlies a given dataset is far from resolved. We develop a new testing framework and procedure to address this problem. In particular, we demonstrate that the procedure has an asymptotically controlled level, and that it is power consistent for certain alternatives. Also, this testing procedure enables the design a new bootstrap method for approximating the distributions of functionals of the leading sample eigenvalues within the subcritical regime -- which is the first such method that is supported by theoretical guarantees.

math.ST↗