arXiv ScienceSearch

arXiv subjects

Yuta Koike

Publications and source records attributed to Yuta Koike.

At least 19 recordsLinked to original sources

Martingale central limit theorems in $p$-Wasserstein distance

We obtain multivariate martingale central limit theorems in $p$-Wasserstein distance with respect to the $\ell_r$ norm in $\mathbb{R}^d$ for $p\geq 1$ and $r\in [1,\infty]$, which generalize the results for $p=1$ and $r=2$ in the literature. As corollaries, we obtain the Yurinskii coupling and Cramér-type moderate deviation results. We also provide an illustrative application to the stochastic gradient descent algorithm. To prove our main results, we combine Lindeberg's swapping argument with a new Gaussian convolution inequality controlling the $p$-Wasserstein distance between a Gaussian convolved with a perturbation and the Gaussian with matching mean and covariance matrix. The latter is obtained by developing the recent line of research on $p$-Wasserstein bounds.

math.PR

Connections between the Föllmer process and the denoising diffusion probabilistic model

The Föllmer process is a Brownian motion conditioned to have a pre-specified distribution at time 1. This process can be interpreted as an ``augmented'' time-compressed version of the reverse stochastic differential equation (SDE) corresponding to the denoising diffusion probabilistic model (DDPM). While this fact has been indirectly used to analyze DDPM sampling errors via discretization of the reverse SDE, the connection between direct discretization of the Föllmer process and the DDPM sampler has not yet been fully explored. This paper clarifies this point while surveying relevant results from the literature. We show that discretized Föllmer processes give natural hyper-parameter settings of the DDPM sampler while accommodating a broader class of variance schedules than discretized reverse SDEs. Moreover, this allows us to systematically recover state-of-the-art results on DDPM sampling error bounds, along with slight improvements.

stat.ML

Wasserstein bounds for denoising diffusion probabilistic models via the Föllmer process

This paper studies sampling error bounds for denoising diffusion probabilistic models (DDPMs) in the 2-Wasserstein distance. Our contributions are threefold. (i) Under general Lipschitz-type conditions on the score function and for a broad class of variance schedules, including the cosine schedule, we establish sharp upper bounds that are optimal in both the dimension and the number of steps, and recover several sharp error bounds previously obtained in the literature. (ii) We prove that the same Lipschitz-type conditions, which encompass those commonly imposed on the (learned) score, imply a logarithmic Sobolev inequality and hence a quadratic transportation cost inequality for the DDPM. As a consequence, in settings covered by existing work, an optimal Wasserstein bound, up to a logarithmic factor, follows from the recently obtained sharp error bound in the Kullback-Leibler divergence under geometric-type variance schedules. (iii) We show that for general log-concave target distributions, the optimal Wasserstein error bound remains attainable even without a quadratic transportation cost inequality for the target. Our analysis is based on viewing the DDPM sampler as a discretization of the Föllmer process rather than the conventional reverse Ornstein-Uhlenbeck process.

stat.ML

High-dimensional bootstrap and asymptotic expansion

The recent seminal work of Chernozhukov, Chetverikov and Kato has shown that bootstrap approximation for the maximum of a sum of independent random vectors is justified even when the dimension is much larger than the sample size. In this context, numerical experiments suggest that third-moment matching bootstrap approximations would outperform normal approximation even without studentization, but the existing theoretical results cannot explain this phenomenon. In this paper, we develop an asymptotic expansion formula for the bootstrap coverage probability and show that it can give an explanation for the above phenomenon. In particular, we find the following interesting blessing of dimensionality phenomenon: The third-moment matching wild bootstrap is second-order accurate in high dimensions even without studentization if the covariance matrix has identical diagonal entries and bounded eigenvalues. We also show that a double wild bootstrap method is second-order accurate regardless of the covariance structure. The validity of these results is established under the assumption that the underlying distributions admit Stein kernels.

math.ST

On lead-lag estimation of non-synchronously observed point processes

This paper introduces a new theoretical framework for analyzing lead-lag relationships between point processes, with a special focus on applications to high-frequency financial data. In particular, we are interested in lead-lag relationships between two sequences of order arrival timestamps. The seminal work of Dobrev and Schaumburg proposed model-free measures of cross-market trading activity based on cross-counts of timestamps. While their method is known to yield reliable results, it faces limitations because its original formulation inherently relies on discrete-time observations, an issue we address in this study. Specifically, we formulate the problem of estimating lead-lag relationships in two point processes as that of estimating the shape of the cross-pair correlation function (CPCF) of a bivariate stationary point process, a quantity well-studied in the neuroscience and spatial statistics literature. Within this framework, the prevailing lead-lag time is defined as the location of the CPCF's sharpest peak. Under this interpretation, the peak location in Dobrev and Schaumburg's cross-market activity measure can be viewed as an estimator of the lead-lag time in the aforementioned sense. We further propose an alternative lead-lag time estimator based on kernel density estimation and show that it possesses desirable theoretical properties and delivers superior numerical performance. Empirical evidence from high-frequency financial data demonstrates the effectiveness of our proposed method.

math.ST

Gaussian Approximation for High-Dimensional $U$-statistics with Size-Dependent Kernels

Motivated by small bandwidth asymptotics for kernel-based semiparametric estimators in econometrics, this paper establishes Gaussian approximation results for high-dimensional fixed-order $U$-statistics whose kernels depend on the sample size. Our results allow for a situation where the dominant component of the Hoeffding decomposition is absent or unknown, including cases with known degrees of degeneracy as special forms. The obtained error bounds for Gaussian approximations are sharp enough to almost recover the weakest bandwidth condition of small bandwidth asymptotics in the fixed-dimensional setting when applied to a canonical semiparametric estimation problem. We also present an application to an adaptive goodness-of-fit testing and the simultaneous inference on high-dimensional density weighted averaged derivatives, along with discussions about several potential applications.

math.ST

Adaptive deep learning for nonlinear time series models

In this paper, we develop a general theory for adaptive nonparametric estimation of the mean function of a non-stationary and nonlinear time series model using deep neural networks (DNNs). We first consider two types of DNN estimators, non-penalized and sparse-penalized DNN estimators, and establish their generalization error bounds for general non-stationary time series. We then derive minimax lower bounds for estimating mean functions belonging to a wide class of nonlinear autoregressive (AR) models that include nonlinear generalized additive AR, single index, and threshold AR models. Building upon the results, we show that the sparse-penalized DNN estimator is adaptive and attains the minimax optimal rates up to a poly-logarithmic factor for many nonlinear AR models. Through numerical simulations, we demonstrate the usefulness of the DNN methods for estimating nonlinear AR models with intrinsic low-dimensional structures and discontinuous or rough mean functions, which is consistent with our theory.

math.ST

Spectral norm bounds for high-dimensional realized covariance matrices and application to weak factor models

Motivated by statistical analysis of latent factor models for high-frequency financial data, we develop sharp upper bounds for the spectral norm of the realized covariance matrix of a high-dimensional Itô semimartingale with possibly infinite activity jumps. For this purpose, we develop Burkholder-Gundy type inequalities for matrix martingales with the help of the theory of non-commutative $L^p$ spaces. The obtained bounds are applied to estimating the number of (relevant) common factors in a continuous-time latent factor model from high-frequency data in the presence of weak factors.

math.ST

Drift estimation for a multi-dimensional diffusion process using deep neural networks

Recently, many studies have shed light on the high adaptivity of deep neural network methods in nonparametric regression models, and their superior performance has been established for various function classes. Motivated by this development, we study a deep neural network method to estimate the drift coefficient of a multi-dimensional diffusion process from discrete observations. We derive generalization error bounds for least squares estimates based on deep neural networks and show that they achieve the minimax rate of convergence up to a logarithmic factor when the drift function has a compositional structure.

math.ST

Sharp High-dimensional Central Limit Theorems for Log-concave Distributions

Let $X_1,\dots,X_n$ be i.i.d. log-concave random vectors in $\mathbb R^d$ with mean 0 and covariance matrix $Σ$. We study the problem of quantifying the normal approximation error for $W=n^{-1/2}\sum_{i=1}^nX_i$ with explicit dependence on the dimension $d$. Specifically, without any restriction on $Σ$, we show that the approximation error over rectangles in $\mathbb R^d$ is bounded by $C(\log^{13}(dn)/n)^{1/2}$ for some universal constant $C$. Moreover, if the Kannan-Lovász-Simonovits (KLS) spectral gap conjecture is true, this bound can be improved to $C(\log^{3}(dn)/n)^{1/2}$. This improved bound is optimal in terms of both $n$ and $d$ in the regime $\log n=O(\log d)$. We also give $p$-Wasserstein bounds with all $p\geq2$ and a Cramér type moderate deviation result for this normal approximation error, and they are all optimal under the KLS conjecture. To prove these bounds, we develop a new Gaussian coupling inequality that gives almost dimension-free bounds for projected versions of $p$-Wasserstein distance for every $p\geq2$. We prove this coupling inequality by combining Stein's method and Eldan's stochastic localization procedure.

math.PR

High-dimensional Central Limit Theorems by Stein's Method in the Degenerate Case

In the literature of high-dimensional central limit theorems, there is a gap between results for general limiting correlation matrix $Σ$ and the strongly non-degenerate case. For the general case where $Σ$ may be degenerate, under certain light-tail conditions, when approximating a normalized sum of $n$ independent random vectors by the Gaussian distribution $N(0,Σ)$ in multivariate Kolmogorov distance, the best-known error rate has been $O(n^{-1/4})$, subject to logarithmic factors of the dimension. For the strongly non-degenerate case, that is, when the minimum eigenvalue of $Σ$ is bounded away from 0, the error rate can be improved to $O(n^{-1/2})$ up to a $\log n$ factor. In this paper, we show that the $O(n^{-1/2})$ rate up to a $\log n$ factor can still be achieved in the degenerate case, provided that the minimum eigenvalue of the limiting correlation matrix of any three components is bounded away from 0. We prove our main results using Stein's method in conjunction with previously unexplored inequalities for the integral of the first three derivatives of the standard Gaussian density over convex polytopes. These inequalities were previously known only for hyperrectangles. Our proof demonstrates the connection between the three-components condition and the third moment Berry--Esseen bound.

math.PR

Improved Central Limit Theorem and bootstrap approximations in high dimensions

This paper deals with the Gaussian and bootstrap approximations to the distribution of the max statistic in high dimensions. This statistic takes the form of the maximum over components of the sum of independent random vectors and its distribution plays a key role in many high-dimensional econometric problems. Using a novel iterative randomized Lindeberg method, the paper derives new bounds for the distributional approximation errors. These new bounds substantially improve upon existing ones and simultaneously allow for a larger class of bootstrap methods.

math.ST

From $p$-Wasserstein Bounds to Moderate Deviations

We use a new method via $p$-Wasserstein bounds to prove Cramér-type moderate deviations in (multivariate) normal approximations. In the classical setting that $W$ is a standardized sum of $n$ independent and identically distributed (i.i.d.) random variables with sub-exponential tails, our method recovers the optimal range of $0\leq x=o(n^{1/6})$ and the near optimal error rate $O(1)(1+x)(\log n+x^2)/\sqrt{n}$ for $P(W>x)/(1-Φ(x))\to 1$, where $Φ$ is the standard normal distribution function. Our method also works for dependent random variables (vectors) and we give applications to the combinatorial central limit theorem, Wiener chaos, homogeneous sums and local dependence. The key step of our method is to show that the $p$-Wasserstein distance between the distribution of the random variable (vector) of interest and a normal distribution grows like $O(p^αΔ)$, $1\leq p\leq p_0$, for some constants $α, Δ$ and $p_0$. In the above i.i.d. setting, $α=1, Δ=1/\sqrt{n}, p_0=n^{1/3}$. For this purpose, we obtain general $p$-Wasserstein bounds in (multivariate) normal approximations using Stein's method.

math.PR

High-dimensional Data Bootstrap

This article reviews recent progress in high-dimensional bootstrap. We first review high-dimensional central limit theorems for distributions of sample mean vectors over the rectangles, bootstrap consistency results in high dimensions, and key techniques used to establish those results. We then review selected applications of high-dimensional bootstrap: construction of simultaneous confidence sets for high-dimensional vector parameters, multiple hypothesis testing via stepdown, post-selection inference, intersection bounds for partially identified parameters, and inference on best policies in policy evaluation. Finally, we also comment on a couple of future research directions.

math.ST

Notes on the dimension dependence in high-dimensional central limit theorems for hyperrectangles

Let $X_1,\dots,X_n$ be independent centered random vectors in $\mathbb{R}^d$. This paper shows that, even when $d$ may grow with $n$, the probability $P(n^{-1/2}\sum_{i=1}^nX_i\in A)$ can be approximated by its Gaussian analog uniformly in hyperrectangles $A$ in $\mathbb{R}^d$ as $n\to\infty$ under appropriate moment assumptions, as long as $(\log d)^5/n\to0$. This improves a result of Chernozhukov, Chetverikov & Kato [Ann. Probab. 45 (2017) 2309-2353] in terms of the dimension growth condition. When $n^{-1/2}\sum_{i=1}^nX_i$ has a common factor across the components, this condition can be further improved to $(\log d)^3/n\to0$. The corresponding bootstrap approximation results are also developed. These results serve as a theoretical foundation of simultaneous inference for high-dimensional models.

math.ST

High-dimensional central limit theorems for homogeneous sums

This paper develops a quantitative version of de Jong's central limit theorem for homogeneous sums in a high-dimensional setting. More precisely, under appropriate moment assumptions, we establish an upper bound for the Kolmogorov distance between a multi-dimensional vector of homogeneous sums and a Gaussian vector so that the bound depends polynomially on the logarithm of the dimension and is governed by the fourth cumulants and the maximal influences of the components. As a corollary, we obtain high-dimensional versions of fourth moment theorems, universality results and Peccati-Tudor type theorems for homogeneous sums. We also sharpen some existing (quantitative) central limit theorems by applications of our result.

math.PR

Nearly optimal central limit theorem and bootstrap approximations in high dimensions

In this paper, we derive new, nearly optimal bounds for the Gaussian approximation to scaled averages of $n$ independent high-dimensional centered random vectors $X_1,\dots,X_n$ over the class of rectangles in the case when the covariance matrix of the scaled average is non-degenerate. In the case of bounded $X_i$'s, the implied bound for the Kolmogorov distance between the distribution of the scaled average and the Gaussian vector takes the form $$C (B^2_n \log^3 d/n)^{1/2} \log n,$$ where $d$ is the dimension of the vectors and $B_n$ is a uniform envelope constant on components of $X_i$'s. This bound is sharp in terms of $d$ and $B_n$, and is nearly (up to $\log n$) sharp in terms of the sample size $n$. In addition, we show that similar bounds hold for the multiplier and empirical bootstrap approximations. Moreover, we establish bounds that allow for unbounded $X_i$'s, formulated solely in terms of moments of $X_i$'s. Finally, we demonstrate that the bounds can be further improved in some special smooth and zero-skewness cases.

math.PR

Large-dimensional Central Limit Theorem with Fourth-moment Error Bounds on Convex Sets and Balls

We prove the large-dimensional Gaussian approximation of a sum of $n$ independent random vectors in $\mathbb{R}^d$ together with fourth-moment error bounds on convex sets and Euclidean balls. We show that compared with classical third-moment bounds, our bounds have near-optimal dependence on $n$ and can achieve improved dependence on the dimension $d$. For centered balls, we obtain an additional error bound that has a sub-optimal dependence on $n$, but recovers the known result of the validity of the Gaussian approximation if and only if $d=o(n)$. We discuss an application to the bootstrap. We prove our main results using Stein's method.

math.PR