arXiv ScienceSearch

arXiv subjects

David Belius

Publications and source records attributed to David Belius.

At least 19 recordsLinked to original sources

Variational formula for the logarithmic potential of free additive convolutions

We establish a general variational formula for the logarithmic potential of the free additive convolution of two compactly supported probability measure on $\R$. The formula is given in terms of the $R$-transform of the first measure, and the logarithmic potential of second measure. The result applies in particular to the additive convolution with the semicircle or Marchenko-Pastur laws, for which the formula simplifies. The logarithmic potential of additive convolutions appears for instance in estimates of the determinant of sums of independent random matrices.

math.PR

A Comprehensive Analysis on the Learning Curve in Kernel Ridge Regression

This paper conducts a comprehensive study of the learning curves of kernel ridge regression (KRR) under minimal assumptions. Our contributions are three-fold: 1) we analyze the role of key properties of the kernel, such as its spectral eigen-decay, the characteristics of the eigenfunctions, and the smoothness of the kernel; 2) we demonstrate the validity of the Gaussian Equivalent Property (GEP), which states that the generalization performance of KRR remains the same when the whitened features are replaced by standard Gaussian vectors, thereby shedding light on the success of previous analyzes under the Gaussian Design Assumption; 3) we derive novel bounds that improve over existing bounds across a broad range of setting such as (in)dependent feature vectors and various combinations of eigen-decay rates in the over/underparameterized regimes.

cs.LG

Characterizing Overfitting in Kernel Ridgeless Regression Through the Eigenspectrum

We derive new bounds for the condition number of kernel matrices, which we then use to enhance existing non-asymptotic test error bounds for kernel ridgeless regression (KRR) in the over-parameterized regime for a fixed input dimension. For kernels with polynomial spectral decay, we recover the bound from previous work; for exponential decay, our bound is non-trivial and novel. Our contribution is two-fold: (i) we rigorously prove the phenomena of tempered overfitting and catastrophic overfitting under the sub-Gaussian design assumption, closing an existing gap in the literature; (ii) we identify that the independence of the features plays an important role in guaranteeing tempered overfitting, raising concerns about approximating KRR generalization using the Gaussian design assumption in previous literature.

cs.LG

On the determinant in Bray-Moore's TAP complexity formula

In the computation of the TAP complexity, originally carried out by Bray and Moore, a fundamental step is to calculate the determinant of a random Hessian. As the replica method does not give a clear prescription, physicists debated how to perform this computation and its consequences on the TAP complexity for a long time. In this paper we prove the original Bray and Moore formula for the behaviour of the determinant at exponential scale to be correct, and compute an important prefactor coming from a small outlier in the spectrum.

math.PR

A Theoretical Analysis of the Test Error of Finite-Rank Kernel Ridge Regression

Existing statistical learning guarantees for general kernel regressors often yield loose bounds when used with finite-rank kernels. Yet, finite-rank kernels naturally appear in several machine learning problems, e.g.\ when fine-tuning a pre-trained deep neural network's last layer to adapt it to a novel task when performing transfer learning. We address this gap for finite-rank kernel ridge regression (KRR) by deriving sharp non-asymptotic upper and lower bounds for the KRR test error of any finite-rank KRR. Our bounds are tighter than previously derived bounds on finite-rank KRR, and unlike comparable results, they also remain valid for any regularization parameters.

cs.LG

Fluctuations of the ground state of the spiked spherical Sherrington-Kirkpatrick model

The Sherrington-Kirkpatrick Hamiltonian is a random quadratic function on the high-dimensional sphere. This article studies the ground state (i.e. maximum) of this Hamiltonian with external field, or more generally with a non-linear "spike" term. We compute the level of the maximum to leading order, and under appropriate condition its first- and second-order fluctuations. The equivalent results are also derived for the maximum of the model's TAP free energy on the ball.

math.PR

TAP variational principle for the constrained overlap multiple spherical Sherrington-Kirkpatrick model

Spin glass models involving multiple replicas with constrained overlaps have been studied in [FPV92; PT07; Pan18a]. For the spherical versions of these models [Ko19; Ko20] showed that the limiting free energy is given by a Parisi type minimization. In this work we show that for Sherrington-Kirkpatrick (i.e. 2-spin) interactions, it can also be expressed in terms of a Thouless-Andersson-Palmer (TAP) variational principle. This is only the second spin glass model where a mathematically rigorous TAP computation of the free energy at all temperatures and external fields has been achieved. The variational formula we derive here also confirms that the model is replica symmetric, a fact which is natural but not obviously deducible from its Parisi formula.

math.PR

Injectivity of ReLU networks: perspectives from statistical physics

When can the input of a ReLU neural network be inferred from its output? In other words, when is the network injective? We consider a single layer, $x \mapsto \mathrm{ReLU}(Wx)$, with a random Gaussian $m \times n$ matrix $W$, in a high-dimensional setting where $n, m \to \infty$. Recent work connects this problem to spherical integral geometry giving rise to a conjectured sharp injectivity threshold for $\alpha = \frac{m}{n}$ by studying the expected Euler characteristic of a certain random set. We adopt a different perspective and show that injectivity is equivalent to a property of the ground state of the spherical perceptron, an important spin glass model in statistical physics. By leveraging the (non-rigorous) replica symmetry-breaking theory, we derive analytical equations for the threshold whose solution is at odds with that from the Euler characteristic. Furthermore, we use Gordon's min--max theorem to prove that a replica-symmetric upper bound refutes the Euler characteristic prediction. Along the way we aim to give a tutorial-style introduction to key ideas from statistical physics in an effort to make the exposition accessible to a broad audience. Our analysis establishes a connection between spin glasses and integral geometry but leaves open the problem of explaining the discrepancies.

cond-mat.dis-nn

Complexity of local maxima of given radial derivative for mixed $p$-spin Hamiltonians

We study the number of local maxima with given radial derivative of spherical mixed $p$-spin models and prove that the second moment matches the square of the first moment on exponential scale for arbitrary mixtures and any radial derivative. This is surprising, since for the number of local maxima with given radial derivative and given energy the corresponding result is only true for specific mixtures [Sub17; BSZ20]. We use standard Kac-Rice computations to derive formulas for the first and second moment at exponential scale, and then find a remarkable analytic argument that shows that the second moment formula is bounded by twice the first moment formula in this general setting. This also leads to a new proof of a central inequality used to prove concentration of the number critical points of pure $p$-spin models of given energy in [Sub17] and removes the need for the computer assisted argument used in that paper for $3 \leq p \leq 10$.

math.PR

Phase diagram for the tap energy of the $p$-spin spherical mean field spin glass model

We solve the Thouless-Anderson-Palmer (TAP) variational principle associated to the spherical pure $p$-spin mean field spin glass Hamiltonian and present a detailed phase diagram. In the high temperature phase the maximum of variational principle is the annealed free energy of the model. In the low temperature phase the maximum, for which we give a formula, is strictly smaller. The high temperature phase consists of three subphases. (1) In the first phase $m=0$ is the unique relevant TAP maximizer. (2) In the second phase there are exponentially many TAP maximizers, but $m=0$ remains dominant. (3) In the third phase, after the so called dynamic phase transition, $m=0$ is no longer a relevant TAP maximizer, and exponentially many non-zero relevant TAP solutions add up to give the annealed free energy. Finally in the low temperature phase a subexponential number of TAP maximizers of near-maximal TAP energy dominate.

cond-mat.dis-nn

Learning Multiscale Convolutional Dictionaries for Image Reconstruction

Convolutional neural networks (CNNs) have been tremendously successful in solving imaging inverse problems. To understand their success, an effective strategy is to construct simpler and mathematically more tractable convolutional sparse coding (CSC) models that share essential ingredients with CNNs. Existing CSC methods, however, underperform leading CNNs in challenging inverse problems. We hypothesize that the performance gap may be attributed in part to how they process images at different spatial scales: While many CNNs use multiscale feature representations, existing CSC models mostly rely on single-scale dictionaries. To close the performance gap, we thus propose a multiscale convolutional dictionary structure. The proposed dictionary structure is derived from the U-Net, arguably the most versatile and widely used CNN for image-to-image learning problems. We show that incorporating the proposed multiscale dictionary in an otherwise standard CSC framework yields performance competitive with state-of-the-art CNNs across a range of challenging inverse problems including CT and MRI reconstruction. Our work thus demonstrates the effectiveness and scalability of the multiscale CSC approach in solving challenging inverse problems.

cs.CV

High temperature TAP upper bound for the free energy of mean field spin glasses

This work proves an upper bound for the free energy of the Sherrington-Kirkpatrick model and its generalizations in terms of the Thouless-Anderson-Palmer (TAP) energy. The result applies to models with spherical or Ising spins and any mixed $p$-spin Hamiltonian with external field or with a non-linear spike term. The bound is expected to be tight to leading order at high temperature, and is non-trivial in the presence of an external field. For the proof a geometric microcanonical method is employed, in which one covers the spin space with sets, each of which is centered at a magnetization vector $m$ and whose contribution to the partition function is bounded in terms of the TAP energy at $m$.

math.PR

Fluctuations of the free energy of the mixed $p$-spin mean field spin glass model

We prove the convergence in distribution of the fluctuations of the free energy of the mixed $p$-spin Sherrington-Kirkpatrick model with non-vanishing $2$-spin component at high enough temperature. The limit is Gaussian, and the fluctuations are seen to arise from weighted cycle counts in the complete graph on the spin indices weighted by the $2$-spin interaction matrix.

math.PR

Triviality of the geometry of mixed $p$-spin spherical Hamiltonians with external field

We study isotropic Gaussian random fields on the high-dimensional sphere with an added deterministic linear term, also known as mixed p-spin Hamiltonians with external field. We prove that if the external field is sufficiently strong, then the resulting function has trivial geometry, that is only two critical points. This contrasts with the situation of no or weak external field where these functions typically have an exponential number of critical points. We give an explicit threshold $h_c$ for the magnitude of the external fieldnecessary for trivialization and conjecture $h_c$ to be sharp. The Kac-Rice formula is our main tool. Our work extends [Fyo15], which identified the trivial regime for the special case of pure p-spin Hamiltonians with random external field.

math.PR

On the Empirical Neural Tangent Kernel of Standard Finite-Width Convolutional Neural Network Architectures

The Neural Tangent Kernel (NTK) is an important milestone in the ongoing effort to build a theory for deep learning. Its prediction that sufficiently wide neural networks behave as kernel methods, or equivalently as random feature models, has been confirmed empirically for certain wide architectures. It remains an open question how well NTK theory models standard neural network architectures of widths common in practice, trained on complex datasets such as ImageNet. We study this question empirically for two well-known convolutional neural network architectures, namely AlexNet and LeNet, and find that their behavior deviates significantly from their finite-width NTK counterparts. For wider versions of these networks, where the number of channels and widths of fully-connected layers are increased, the deviation decreases.

cs.LG

Tightness for the Cover Time of the two dimensional sphere

Let $C^*_{ε,S^2}$ denote the cover time of the two dimensional sphere by a Wiener sausage of radius $ε$. We prove that $$\sqrt{C^{*}_{ε,S^2} } -\sqrt{\frac{2A_{S^2}}π}(\log ε^{-1}-\frac14\log\log ε^{-1})$$ is tight, where $A_{S^2}=4π$ denotes the Riemannian area of $S^2$.

math.PR

Maximum of the Ginzburg-Landau fields

We study two dimensional massless field in a box with potential $V\left( \nabla ϕ\left( \cdot \right) \right) $ and zero boundary condition, where $V$ is any symmetric and uniformly convex function. Naddaf-Spencer and Miller proved the macroscopic averages of this field converge to a continuum Gaussian free field. In this paper we prove the distribution of local marginal $ϕ\left( x\right) $, for any $x$ in the bulk, has a Gaussian tail. We further characterize the leading order of the maximum and dimension of high points of this field, thus generalize the results of Bolthausen-Deuschel-Giacomin and Daviaud for the discrete Gaussian free field.

math.PR

The TAP-Plefka variational principle for the spherical SK model

We reinterpret the Thouless-Anderson-Palmer approach to mean field spin glass models as a variational principle in the spirit of the Gibbs variational principle and the Bragg-Williams approximation. We prove this TAP-Plefka variational principle rigorously in the case of the spherical Sherrington-Kirkpatrick model.

math.PR