arXiv ScienceSearch

arXiv subjects

Boris Nectoux

Publications and source records attributed to Boris Nectoux.

At least 19 recordsLinked to original sources

Quantifying Uncertainty In Wide Two-Layer Neural Networks: On The Law Of The Limiting Fluctuation Process

Uncertainty quantification in neural networks prediction is a main issue for usual applications. Our approach seeks at reducing computation costs by directly evaluating uncertainty using PDE's information on the asymptotic variance, rather than the deep ensemble method which may be seen as a Monte Carlo estimation of the prediction, requiring the training of multiple networks. We thus study the law of the limiting process describing the random fluctuations around the mean-field limit of wide two-layer neural networks trained by stochastic gradient descent in a weak-noise regime. Building on a recent trajectorial central limit theorem, in which this limit is characterized as the weak solution of a linear stochastic evolution equation, we identify its law explicitly. More precisely, we show that it is a centered Gaussian process in the dual of a weighted Sobolev space, and we derive a closed covariance representation for the finite-dimensional distributions obtained by testing it against smooth functions. This covariance is expressed through the solution of a backward transport equation with a nonlocal source term, whose coefficients are driven by the mean-field trajectory. As a consequence, by testing against the activation function at a fixed input, we obtain an expression for the limiting variance of the corresponding network-output fluctuations. We illustrate this result numerically on a one-dimensional regression example.

cs.NE

Uniform-in-time concentration in two-layer neural networks via transportation inequalities

We quantify, uniformly over time and with high probability, the discrepancy between the predictions of a two-layer neural network trained by stochastic gradient descent (SGD) and their mean-field limit, for quadratic loss and ridge regularization. As a key ingredient, we establish T p transportation inequalities (p $\in$ {1, 2}) for the law of the SGD parameters, with explicit constants independent of the iteration index. We then prove uniform-in-time concentration of the empirical parameter measure around its mean-field limit in the Wasserstein distance W 1 , and we translate these bounds into prediction-error estimates against a fixed test function $\Phi$. We also derive analogous concentration bounds in the sliced-Wasserstein distance SW 1 , leading to dimension-free rates.

cs.NE

Eyring-Kramers formula for the mean exit time of non-Gibbsian elliptic processes: the non characteristic boundary case

In this work, we derive a new sharp asymptotic equivalent in the small temperature regime $h\to 0$ for the mean exit time from a bounded domain for the non-reversible process $dX\_t=b(X\_t)dt + \sqrt h \, dB\_t$ under a generic orthogonal decomposition of $b$ and when the boundary of $\Omega$ is assumed to be \textit{non characteristic}. The main contribution of this work lies in the fact that we do not assume that the process $(X\_t,t\ge 0)$ is \textit{Gibbsian}. In this case, a new correction term characterizing the \textit{non-Gibbsianness} of the process appears in the equivalent of the mean exit time. The proof is mainly based on tools from spectral and semi-classical analysis.

math.AP

Quasi-stationarity of the Dyson Brownian motion with collisions

In this work, we investigate the ergodic behavior of a system of particules, subject to collisions, before it exits a fixed subdomain of its state space. This system is composed of several one-dimensional ordered Brownian particules in interaction with electrostatic repulsions, which is usually referred as the (generalized) Dyson Brownian motion. The starting points of our analysis are the work [E. C{\'e}pa and D. L{\'e}pingle, 1997 Probab. Theory Relat. Fields] which provides existence and uniqueness of such a system subject to collisions via the theory of multivalued SDEs and a Krein-Rutman type theorem derived in [A. Guillin, B. Nectoux, L. Wu, 2020 J. Eur. Math. Soc.].

math.PR

Large deviations of the empirical measures of a strong-Feller Markov process inside a subset and quasi-ergodic distribution

In this work, we establish, for a strong Feller process, the large deviation principle for the occupation measure conditioned not to exit a given subregion. The rate function vanishes only at a unique measure, which is the so-called quasi-ergodic distribution of the process in this subregion. In addition, we show that the rate function is the Dirichlet form in the particular case when the process is reversible. We apply our results to several stochastic processes such as the solutions of elliptic stochastic differential equations driven by a rotationally invariant $\alpha$-stable process, the kinetic Langevin process, and the overdamped Langevin process driven by a Brownian motion.

math.PR

Long time behavior of killed Feynman-Kac semigroups with singular Schr{\"o}dinger potentials

In this work, we investigate the compactness and the long time behavior of killed Feynman-Kac semigroups of various processes arising from statistical physics with very general singular Schr{\"o}dinger potentials. The processes we consider cover a large class of processes used in statistical physics, with strong links with quantum mechanics and (local or not) Schr{\"o}dinger operators (including e.g. fractional Laplacians). For instance we consider solutions to elliptic differential equations, L{\'e}vy processes, the kinetic Langevin process with locally Lipschitz gradient fields, and systems of interacting L{\'e}vy particles. Our analysis relies on a Perron-Frobenius type theorem derived in a previous work [A. Guillin, B. Nectoux, L. Wu, 2020 J. Eur. Math. Soc.] for Feller kernels and on the tools introduced in [L. Wu, 2004, Probab. Theory Relat. Fields] to compute bounds on the essential spectral radius of a bounded nonnegative kernel.

math.PR

Central Limit Theorem for Bayesian Neural Network trained with Variational Inference

In this paper, we rigorously derive Central Limit Theorems (CLT) for Bayesian two-layerneural networks in the infinite-width limit and trained by variational inference on a regression task. The different networks are trained via different maximization schemes of the regularized evidence lower bound: (i) the idealized case with exact estimation of a multiple Gaussian integral from the reparametrization trick, (ii) a minibatch scheme using Monte Carlo sampling, commonly known as Bayes-by-Backprop, and (iii) a computationally cheaper algorithm named Minimal VI. The latter was recently introduced by leveraging the information obtained at the level of the mean-field limit. Laws of large numbers are already rigorously proven for the three schemes that admits the same asymptotic limit. By deriving CLT, this work shows that the idealized and Bayes-by-Backprop schemes have similar fluctuation behavior, that is different from the Minimal VI one. Numerical experiments then illustrate that the Minimal VI scheme is still more efficient, in spite of bigger variances, thanks to its important gain in computational complexity.

stat.ML

Generalized Langevin And Nos{\'e}-hoover Processes Absorbed At The Boundary Of A Metastable Domain

In this paper, we prove in a very weak regularity setting existence and uniqueness of quasi-stationary distributions as well as exponential conver- gence towards the quasi-stationary distribution for the generalized Langevin and the Nos{\'e}-Hoover processes, two processes which are widely used in molecular dynamics. The case of singular potentials is considered. With the techniques used in this work, we are also able to greatly improve existing results on quasi-stationary distributions for the kinetic Langevin process to a weak regularity setting.

math.PR

Law of Large Numbers for Bayesian two-layer Neural Network trained with Variational Inference

We provide a rigorous analysis of training by variational inference (VI) of Bayesian neural networks in the two-layer and infinite-width case. We consider a regression problem with a regularized evidence lower bound (ELBO) which is decomposed into the expected log-likelihood of the data and the Kullback-Leibler (KL) divergence between the a priori distribution and the variational posterior. With an appropriate weighting of the KL, we prove a law of large numbers for three different training schemes: (i) the idealized case with exact estimation of a multiple Gaussian integral from the reparametrization trick, (ii) a minibatch scheme using Monte Carlo sampling, commonly known as Bayes by Backprop, and (iii) a new and computationally cheaper algorithm which we introduce as Minimal VI. An important result is that all methods converge to the same mean-field limit. Finally, we illustrate our results numerically and discuss the need for the derivation of a central limit theorem.

stat.ML

Recent advances in the long-time analysis of killed degenerate processes and their particle approximation

We review some recent results of quantitative long-time convergence for the law of a killed Markov process conditioned to survival toward a quasi-stationary distribution, and on the analogous question for the particle systems used in practice to sample these distributions. With respect to the existing literature, one of the novelties of these works is the degeneracy of the underlying process with respect to classical elliptic diffusion, namely it can be a non-elliptic hypoelliptic diffusion, a piecewise deterministic Markov process or an Euler numerical scheme.

math.PR

Exit time and principal eigenvalue of non-reversible elliptic diffusions

In this work, we analyse the metastability of non-reversible diffusion processes $$dX_t=\boldsymbol{b}(X_t)dt+\sqrt h\,dB_t$$ on a bounded domain $\Omega$ when $\mathbf{b}$ admits the decomposition $\mathbf{b}=-(\nabla f+\mathbf{\ell})$ and $\nabla f \cdot \mathbf{\ell}=0$. In this setting, we first show that, when $h\to 0$, the principal eigenvalue of the generator of $(X_t)_{t\ge 0}$ with Dirichlet boundary conditions on the boundary $\partial\Omega$ of $\Omega$ is exponentially close to the inverse of the mean exit time from $\Omega$, uniformly in the initial conditions $X_0=x$ within the compacts of $\Omega$. The asymptotic behavior of the law of the exit time in this limit is also obtained. The main novelty of these first results follows from the consideration of non-reversible elliptic diffusions whose associated dynamical systems $\dot X=\mathbf{b}(X)$ admit equilibrium points on $\partial\Omega$. In a second time, when in addition $\div \mathbf{\ell} =0$, we derive a new sharp asymptotic equivalent in the limit $h\to 0$ of the principal eigenvalue of the generator of the process and of its mean exit time from $\Omega$. Our proofs combine tools from large deviations theory and from semiclassical analysis, and truly relies on the notion of quasi-stationary distribution.

math.PR

Law of large numbers and central limit theorem for wide two-layer neural networks: the mini-batch and noisy case

In this work, we consider a wide two-layer neural network and study the behavior of its empirical weights under a dynamics set by a stochastic gradient descent along the quadratic loss with mini-batches and noise. Our goal is to prove a trajectorial law of large number as well as a central limit theorem for their evolution. When the noise is scaling as 1/N $\beta$ and 1/2 < $\beta$ $\le$ $\infty$, we rigorously derive and generalize the LLN obtained for example in [CRBVE20, MMM19, SS20b]. When 3/4 < $\beta$ $\le$ $\infty$, we also generalize the CLT (see also [SS20a]) and further exhibit the effect of mini-batching on the asymptotic variance which leads the fluctuations. The case $\beta$ = 3/4 is trickier and we give an example showing the divergence with time of the variance thus establishing the instability of the predictions of the neural network in this case. It is illustrated by simple numerical examples.

math.PR

Eyring-Kramers exit rates for the overdamped Langevin dynamics: the case with saddle points on the boundary

Let $(X_t)_{t\ge 0}$ be the stochastic process solution to the overdamped Langevin dynamics $$dX_t=-\nabla f(X_t) \, dt +\sqrt h \, dB_t$$ and let $\Omega \subset \mathbb R^d $ be the basin of attraction of a local minimum of $f: \mathbb R^d \to \mathbb R$. Up to a small perturbation of $\Omega$ to make it smooth, we prove that the exit rates of $(X_t)_{t\ge 0}$ from $\Omega$ through each of the saddle points of $f$ on $\partial \Omega$ can be parametrized by the celebrated Eyring-Kramers laws, in the limit $h \to 0$. This result provides firm mathematical grounds to jump Markov models which are used to model the evolution of molecular systems, as well as to some numerical methods which use these underlying jump Markov models to efficiently sample metastable trajectories of the overdamped Langevin dynamics.

math.PR

Eyring-Kramers type formulas for some piecewise deterministic Markov processes

In this work, we give sharp asymptotic equivalents in the small temperature regime of the smallest eigenvalues of the generator of some piecewise deterministic Markov processes (including the ZigZag process and the Bouncy Particle Sampler process) with refreshment rate $\alpha$ on the one-dimensional torus T. These asymptotic equivalents are usually called Eyring-Kramers type formulas in the literature. The case when the refreshment rate $\alpha$ vanishes on T is also considered.

math-ph

The exit from a metastable state: concentration of the exit point distribution on the low energy saddle points, part 2

We consider the first exit point distribution from a bounded domain $\Omega$ of the stochastic process $(X_t)_{t\ge 0}$ solution to the overdamped Langevin dynamics $$d X_t = -\nabla f(X_t) d t + \sqrt{h} \ d B_t$$ starting from deterministic initial conditions in $\Omega$, under rather general assumptions on $f$ (for instance, $f$ may have several critical points in $\Omega$). This work is a continuation of the previous paper \cite{DLLN-saddle1} where the exit point distribution from $\Omega$ is studied when $X_0$ is initially distributed according to the quasi-stationary distribution of $(X_t)_{t\ge 0}$ in $\Omega$. The proofs are based on analytical results on the dependency of the exit point distribution on the initial condition, large deviation techniques and results on the genericity of Morse functions.

math.AP

Small eigenvalues of the Witten Laplacian with Dirichlet boundary conditions: the case with critical points on the boundary

In this work, we give sharp asymptotic equivalents in the limit $h\to 0$ of the small eigenvalues of the Witten Laplacian, that is the operator associated with the quadratic form $$ \psi\in H^1_0(\Omega)\mapsto h^2 \int_\Omega \big \vert \nabla \big (e^{\frac 1hf} \psi\big )\big \vert^2\, e^{-\frac 2hf},$$where $\overline\Omega=\Omega\cup \partial \Omega$ is an oriented $C^\infty$ compact and connected Riemannian manifold with non empty boundary $\partial \Omega$ and $f: \overline \Omega\to \mathbb R$ is a $C^\infty$ Morse function. The function $f$ is allowed to admit critical points on $ \partial \Omega$, which is the main novelty of this work in comparison with the existing literature.

math.SP

Repartition of the quasi-stationary distribution and first exit point density for a double-well potential

Let f : R d $\rightarrow$ R be a smooth function and (Xt) t$\ge$0 be the stochastic process solution to the overdamped Langevin dynamics dXt = ----f (Xt)dt + $\sqrt$ h dBt. Let $\Omega$ $\subset$ R d be a smooth bounded domain and assume that f | $\Omega$ is a double-well potential with degenerate barriers. In this work, we study in the small temperature regime, i.e. when h $\rightarrow$ 0 + , the asymptotic repartition of the quasi-stationary distribution of (Xt) t$\ge$0 in $\Omega$ within the two wells of f | $\Omega$. We show that this distribution generically concentrates in precisely one well of f | $\Omega$ when h $\rightarrow$ 0 + but can nevertheless concentrate in both wells when f | $\Omega$ admits sufficient symmetries. This phenomenon corresponds to the so-called tunneling effect in semiclassical analysis. We also investigate in this setting the asymptotic behaviour when h $\rightarrow$ 0 + of the first exit point distribution from $\Omega$ of (Xt) t$\ge$0 when X0 is distributed according to the quasi-stationary distribution. 1 Setting and results 1.1 Quasi-stationary distribution and purpose of this work Let (X t) t$\ge$0 be the stochastic process solution to the overdamped Langevin dynamics in R d : dX t = ----f (X t)dt + $\sqrt$ h dB t , (1) where f : R d $\rightarrow$ R is the potential (chosen C $\infty$ in all this work), h > 0 is the temperature and (B t) t$\ge$0 is a standard d-dimensional Brownian motion. Let $\Omega$ be a C $\infty$ bounded open and connected subset of R d and introduce $\tau$ $\Omega$ = inf{t $\ge$ 0 | X t / $\in$ $\Omega$} the first exit time from $\Omega$. A quasi-stationary distribution for the process (1) on $\Omega$ is a probability measure $\mu$ h on $\Omega$ such that, when X 0 $\sim$ $\mu$ h , it holds for any time t > 0 and any Borel set A $\subset$ $\Omega$, P(X t $\in$ A | t < $\tau$ $\Omega$) = $\mu$ h (A).

math.AP

The exit from a metastable state: concentration of the exit point distribution on the low energy saddle points

We consider the first exit point distribution from a bounded domain $\Omega$ of the stochastic process $(X_t)_{t\ge 0}$ solution to the overdamped Langevin dynamics $$d X_t = -\nabla f(X_t) d t + \sqrt{h} \ d B_t$$ starting from the quasi-stationary distribution in $\Omega$. In the small temperature regime ($h\to 0$) and under rather general assumptions on $f$ (in particular, $f$ may have several critical points in $\Omega$), it is proven that the support of the distribution of the first exit point concentrates on some points realizing the minimum of $f$ on $\partial \Omega$. The proof relies on tools to study tunnelling effects in semi-classical analysis. Extensions of the results to more general initial distributions than the quasi-stationary distribution are also presented.

math.AP