arXiv ScienceSearch

arXiv · 2503.13986

Stratified Permutational Berry--Esseen Bounds and Their Applications to Statistics

Abstract

The stratified linear permutation statistic arises in various statistics problems, including stratified and post-stratified survey sampling, stratified and post-stratified experiments, conditional permutation tests, etc. Although we can derive the Berry--Esseen bounds for the stratified linear permutation statistic based on existing bounds for the non-stratified statistics, those bounds are not sharp, and moreover, this strategy does not work in general settings with heterogeneous strata with varying sizes. We first use Stein's method to obtain a unified stratified permutational Berry--Esseen bound that can accommodate heterogeneous strata. We then apply the bound to various statistics problems, leading to stronger theoretical quantifications and thereby facilitating statistical inference in those problems.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Pengfei Tian, Fan Yang, Peng Ding. 2026-03-17. Stratified Permutational Berry--Esseen Bounds and Their Applications to Statistics. https://arxiv.org/abs/2503.13986

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A Note on Inferential Decisions, Errors and Path-Dependency

Consider sequential binary testing under model uncertainty or misspecification in an otherwise 'ideal' setting: the a posteriori belief process and its objective conditional probability counterpart may differ but converge to the same correct outcome. We show that under common conditions (defined) unless the two are 'essentially identical', differing only by a priori factors, time-homogeneous continuous decisions based on one must fail to be path-independent with respect to state-variables based on the other or any non-essentially-identical process. The difference between them, inferential error, decomposes uniquely into two independent components: a path-independent systematic bias and a path-dependent, not necessarily systematic, error.

math.ST

Multiscale Localized Inference for Networks with a Measured Vertex Coordinate

In many networks each vertex has a position measured from outside the network: a neuron's location along the body axis, a residue's index along a protein sequence, a genomic bin's position in base pairs. The chance that two vertices connect is then a surface over pairs of positions, and questions about the network become questions about small regions of that surface. A region on the diagonal covers pairs inside one stretch of the axis, and a departure there is a community with a boundary. A region off the diagonal covers pairs spanning two separated stretches, and a departure there is a bridge. Existing methods address one part of this at a time: community detection returns groups without placing them on the axis, block models fix the width in advance, and scan statistics test one window at one scale. We expand the surface in a wavelet dictionary whose elements are exactly such regions, at every position and width. We derive in closed form the modularity, edge length, transitivity and degree spread each element produces, and the generating element is recovered from those summaries. A scan estimates every coefficient against a background fitted on separate edges and controls the error rate across all positions and widths at once. Applied to a connectome, a protein contact map and a chromatin contact map, it recovers known anatomy in the first and is calibrated in all three. When positions are estimated with error near the width sought, the location of a departure is not identified at any signal strength, and a check decides this before any analysis.

math.ST

Algorithms for adaptive and heteroskedastic linear regression at the computational threshold

We study finite-sample linear regression in the presence of varied and unknown label noise, focusing on the heteroskedastic and adaptive linear regression models. Heteroskedastic linear regression models settings where the labels are of varying quality. We receive $n$ pairs $(X_i,Y_i)$ with labels $Y_i=X_i^\topβ+\varepsilon_i$, where $\varepsilon_i\sim N(0,σ_i^2)$ and the variances are unknown to the estimator. One natural measurement of the difficulty of this problem is the number of samples $m$ for which $σ_i^2\le1$ (larger $m$ is easier). We obtain a polynomial-time estimator with rate $\tilde{O}((nd^3/m^4)^{1/6})$ when $m\gg d^{3/4}n^{1/4}$, as well as nearly-matching lower bounds. For $d=O(1)$, our estimator achieves error $o(1)$ when $m\gg n^{1/4}$, whereas $L_1$ regression and other traditional approaches require $m\gg n^{1/2}$. In adaptive linear regression, the errors are drawn i.i.d. from an unknown distribution $p$, and our goal is to design a generic estimator that performs nearly as well as the best custom estimator that knows $p$. We introduce a (computationally inefficient) adaptive estimator that, so long as $p$ is a mixture of $k$ symmetric log-concave densities, achieves error comparable with the optimal estimator that knows $p$ and has $\tildeΘ(n/k)$ samples. For $k=1$, we show that $L_q$ regression (with data-dependent $q$) gives a polynomial-time estimator. Finally, to study the computational limits of both problems, we introduce the planted linear regression problem, where $X_i\sim N(0,I_d)$, $m$ unknown samples are noiseless, and the rest have error $\varepsilon_i\sim N(0,1)$. We conjecture that recovering $β$ up to error $\ll\sqrt{d/n}$ (or exactly) may have an information-computation gap between $m=d+1$ and $m\sim d^{3/4}n^{1/4}$, as is suggested by our near-matching polynomial-time estimator and statistical query (SQ) lower bound.

math.ST