arXiv ScienceSearch

arXiv subjects

Matthieu Wyart

Publications and source records attributed to Matthieu Wyart.

At least 37 records · Page 2Linked to original sources

The Role of Excitations in Supercooled Liquids: Density, Geometry, and Relaxation Dynamics

Low-energy excitations play a key role in all condensed-matter systems, yet there is limited understanding of their nature in glasses, where they correspond to local rearrangements of groups of particles. Here we introduce an algorithm to systematically uncover these excitations up to the activation energy scale relevant to structural relaxation. We use it in a model system to measure the density of states on a scale never achieved before, confirming that this quantity shifts to higher energy under cooling, precisely as the activation energy does. Secondly, we show that the excitations' energetic and spatial features allow one to predict with great accuracy the dynamic propensity, i.e. the location of future relaxation dynamics. Finally, we find that excitations have a core whose properties, including the displacement of the most mobile particle, scale as a power-law of their activation energy and are independent of temperature. Additionally, they exhibit an outer deformation field that depends on the material's stability and, therefore, on temperature. We build a scaling description of these findings. Overall, our analysis supports that excitations play a crucial role in regulating relaxation dynamics near the glass transition, effectively suppressing the transition to dynamical arrest predicted by mean-field theories while also being strongly influenced by it.

cond-mat.soft

How Deep Networks Learn Sparse and Hierarchical Data: the Sparse Random Hierarchy Model

Understanding what makes high-dimensional data learnable is a fundamental question in machine learning. On the one hand, it is believed that the success of deep learning lies in its ability to build a hierarchy of representations that become increasingly more abstract with depth, going from simple features like edges to more complex concepts. On the other hand, learning to be insensitive to invariances of the task, such as smooth transformations for image datasets, has been argued to be important for deep networks and it strongly correlates with their performance. In this work, we aim to explain this correlation and unify these two viewpoints. We show that by introducing sparsity to generative hierarchical models of data, the task acquires insensitivity to spatial transformations that are discrete versions of smooth transformations. In particular, we introduce the Sparse Random Hierarchy Model (SRHM), where we observe and rationalize that a hierarchical representation mirroring the hierarchical model is learnt precisely when such insensitivity is learnt, thereby explaining the strong correlation between the latter and performance. Moreover, we quantify how the sample complexity of CNNs learning the SRHM depends on both the sparsity and hierarchical structure of the task.

stat.ML

A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Data

Understanding the structure of real data is paramount in advancing modern deep-learning methodologies. Natural data such as images are believed to be composed of features organized in a hierarchical and combinatorial manner, which neural networks capture during learning. Recent advancements show that diffusion models can generate high-quality images, hinting at their ability to capture this underlying compositional structure. We study this phenomenon in a hierarchical generative model of data. We find that the backward diffusion process acting after a time $t$ is governed by a phase transition at some threshold time, where the probability of reconstructing high-level features, like the class of an image, suddenly drops. Instead, the reconstruction of low-level features, such as specific details of an image, evolves smoothly across the whole diffusion process. This result implies that at times beyond the transition, the class has changed, but the generated sample may still be composed of low-level elements of the initial image. We validate these theoretical insights through numerical experiments on class-unconditional ImageNet diffusion models. Our analysis characterizes the relationship between time and scale in diffusion models and puts forward generative models as powerful tools to model combinatorial data properties.

stat.ML

Dynamical heterogeneities of thermal creep in pinned interfaces

Disordered systems under applied loading display slow creep flows at finite temperature, which can lead to the material rupture. Renormalization group arguments predicted that creep proceeds via thermal avalanches of activated events. Recently, thermal avalanches were argued to control the dynamics of liquids near their glass transition. Both theoretical approaches are markedly different. Here we provide a scaling description that seeks to unify dynamical heterogeneities in both phenomena, confirm it in simple models of pinned elastic interfaces, and discuss its experimental implications.

cond-mat.dis-nn

Ductile-to-brittle transition and yielding in soft amorphous materials: perspectives and open questions

Soft amorphous materials are viscoelastic solids ubiquitously found around us, from clays and cementitious pastes to emulsions and physical gels encountered in food or biomedical engineering. Under an external deformation, these materials undergo a noteworthy transition from a solid to a liquid state that reshapes the material microstructure. This yielding transition was the main theme of a workshop held from January 9 to 13, 2023 at the Lorentz Center in Leiden. The manuscript presented here offers a critical perspective on the subject, synthesizing insights from the various brainstorming sessions and informal discussions that unfolded during this week of vibrant exchange of ideas. The result of these exchanges takes the form of a series of open questions that represent outstanding experimental, numerical, and theoretical challenges to be tackled in the near future.

cond-mat.soft

On the different regimes of Stochastic Gradient Descent

Modern deep networks are trained with stochastic gradient descent (SGD) whose key hyperparameters are the number of data considered at each step or batch size $B$, and the step size or learning rate $\eta$. For small $B$ and large $\eta$, SGD corresponds to a stochastic evolution of the parameters, whose noise amplitude is governed by the ''temperature'' $T\equiv \eta/B$. Yet this description is observed to break down for sufficiently large batches $B\geq B^*$, or simplifies to gradient descent (GD) when the temperature is sufficiently small. Understanding where these cross-overs take place remains a central challenge. Here, we resolve these questions for a teacher-student perceptron classification model and show empirically that our key predictions still apply to deep networks. Specifically, we obtain a phase diagram in the $B$-$\eta$ plane that separates three dynamical phases: (i) a noise-dominated SGD governed by temperature, (ii) a large-first-step-dominated SGD and (iii) GD. These different phases also correspond to different regimes of generalization error. Remarkably, our analysis reveals that the batch size $B^*$ separating regimes (i) and (ii) scale with the size $P$ of the training set, with an exponent that characterizes the hardness of the classification problem.

cs.LG

Testing theories of the glass transition with the same liquid, but many kinetic rules

We study the glass transition by exploring a broad class of kinetic rules that can significantly modify the normal dynamics of super-cooled liquids, while maintaining thermal equilibrium. Beyond the usual dynamics of liquids, this class includes dynamics in which a fraction $(1-f_R)$ of the particles can perform pairwise exchange or 'swap moves', while a fraction $f_P$ of the particles can only move along restricted directions. We find that (i) the location of the glass transition varies greatly but smoothly as $f_P$ and $f_R$ change and (ii) it is governed by a linear combination of $f_P$ and $f_R$. (iii) Dynamical heterogeneities are not governed by the static structure of the material. Instead, they are similar at the glass transition across the ($f_R$, $f_P$) diagram. These observations are negative items for some existing theories of the glass transition, particularly those reliant on growing thermodynamic order or locally favored structure, and open new avenues to test other approaches.

cond-mat.soft

How Deep Neural Networks Learn Compositional Data: The Random Hierarchy Model

Deep learning algorithms demonstrate a surprising ability to learn high-dimensional tasks from limited examples. This is commonly attributed to the depth of neural networks, enabling them to build a hierarchy of abstract, low-dimensional data representations. However, how many training examples are required to learn such representations remains unknown. To quantitatively study this question, we introduce the Random Hierarchy Model: a family of synthetic tasks inspired by the hierarchical structure of language and images. The model is a classification task where each class corresponds to a group of high-level features, chosen among several equivalent groups associated with the same class. In turn, each feature corresponds to a group of sub-features chosen among several equivalent ones and so on, following a hierarchy of composition rules. We find that deep networks learn the task by developing internal representations invariant to exchanging equivalent groups. Moreover, the number of data required corresponds to the point where correlations between low-level features and classes become detectable. Overall, our results indicate how deep networks overcome the curse of dimensionality by building invariant representations, and provide an estimate of the number of data required to learn a hierarchical task.

cs.LG

Scaling Description of Dynamical Heterogeneity and Avalanches of Relaxation in Glass-Forming Liquids

We provide a theoretical description of dynamical heterogeneities in glass-forming liquids, based on the premise that relaxation occurs via local rearrangements coupled by elasticity. In our framework, the growth of the dynamical correlation length $\xi$ and of the correlation volume $\chi_4$ are controlled by a zero-temperature fixed point. We connect this critical behavior to the properties of the distribution of local energy barriers at zero temperature. Our description makes a direct connection between dynamical heterogeneities and avalanche-type relaxation associated to dynamic facilitation, allowing us to relate the size distribution of heterogeneities to their time evolution. Within an avalanche, a local region relaxes multiple times, the more the larger is the avalanche. This property, related to the nature of the zero-temperature fixed point, directly leads to decoupling of particle diffusion and relaxation time (the so-called Stokes-Einstein violation). Our most salient predictions are tested and confirmed by numerical simulations of scalar and tensorial thermal elasto-plastic models.

cond-mat.soft

Local vs. Cooperative: Unraveling Glass Transition Mechanisms with SEER

Which phenomenon slows down the dynamics in super-cooled liquids and turns them into glasses is a long-standing question of condensed-matter. Most popular theories posit that as the temperature decreases, many events must occur in a coordinated fashion on a growing length scale for relaxation to occur. Instead, other approaches consider that local barriers associated with the elementary rearrangement of a few particles or `excitations' govern the dynamics. To resolve this conundrum, our central result is to introduce an algorithm, SEER, which can systematically extract hundreds of excitations and their energy from any given configuration. We also provide a novel measurement of the activation energy, characterizing the liquid dynamics, based on fast quenching and reheating. We use these two methods in a popular liquid model of polydisperse particles. Such polydisperse models are known to capture the hallmarks of the glass transition and can be equilibrated efficiently up to millisecond time scales. The analysis reveals that cooperative effects do not control the fragility of such liquids: the change of energy of local barriers determines the change of activation energy. More generally, these methods can now be used to measure the degree of cooperativity of any liquid model.

cond-mat.soft

Dissecting the Effects of SGD Noise in Distinct Regimes of Deep Learning

Understanding when the noise in stochastic gradient descent (SGD) affects generalization of deep neural networks remains a challenge, complicated by the fact that networks can operate in distinct training regimes. Here we study how the magnitude of this noise $T$ affects performance as the size of the training set $P$ and the scale of initialization $\alpha$ are varied. For gradient descent, $\alpha$ is a key parameter that controls if the network is `lazy'($\alpha\gg1$) or instead learns features ($\alpha\ll1$). For classification of MNIST and CIFAR10 images, our central results are: (i) obtaining phase diagrams for performance in the $(\alpha,T)$ plane. They show that SGD noise can be detrimental or instead useful depending on the training regime. Moreover, although increasing $T$ or decreasing $\alpha$ both allow the net to escape the lazy regime, these changes can have opposite effects on performance. (ii) Most importantly, we find that the characteristic temperature $T_c$ where the noise of SGD starts affecting the trained model (and eventually performance) is a power law of $P$. We relate this finding with the observation that key dynamical quantities, such as the total variation of weights during training, depend on both $T$ and $P$ as power laws. These results indicate that a key effect of SGD noise occurs late in training by affecting the stopping process whereby all data are fitted. Indeed, we argue that due to SGD noise, nets must develop a stronger `signal', i.e. larger informative weights, to fit the data, leading to a longer training time. A stronger signal and a longer training time are also required when the size of the training set $P$ increases. We confirm these views in the perceptron model, where signal and noise can be precisely measured. Interestingly, exponents characterizing the effect of SGD depend on the density of data near the decision boundary, as we explain.

cs.LG

Armouring of a frictional interface by mechanical noise

A dry frictional interface loaded in shear often displays stick-slip. The amplitude of this cycle depends on the probability that a microscopic event nucleates a rupture and on the rate at which microscopic events are triggered. The latter is determined by the distribution of soft spots, $P(x)$, which is the density of microscopic regions that yield if the shear load is increased by some amount $x$. In minimal models of a frictional interface - that include disorder, inertia and long-range elasticity - we discovered an 'armouring' mechanism by which the interface is greatly stabilised after a large slip event: $P(x)$ then vanishes at small argument as $P(x)\sim x^\theta$ [1]. The exponent $\theta$ is non-zero only in the presence of inertia (otherwise $\theta=0$). It was found to depend on the statistics of the disorder in the model, a phenomenon that was not explained. Here, we show that a single-particle toy model with inertia and disorder captures the existence of a non-trivial exponent $\theta>0$, which we can analytically relate to the statistics of the disorder.

cond-mat.soft

How deep convolutional neural networks lose spatial information with training

A central question of machine learning is how deep nets manage to learn tasks in high dimensions. An appealing hypothesis is that they achieve this feat by building a representation of the data where information irrelevant to the task is lost. For image datasets, this view is supported by the observation that after (and not before) training, the neural representation becomes less and less sensitive to diffeomorphisms acting on images as the signal propagates through the net. This loss of sensitivity correlates with performance, and surprisingly correlates with a gain of sensitivity to white noise acquired during training. These facts are unexplained, and as we demonstrate still hold when white noise is added to the images of the training set. Here, we (i) show empirically for various architectures that stability to image diffeomorphisms is achieved by both spatial and channel pooling, (ii) introduce a model scale-detection task which reproduces our empirical observations on spatial pooling and (iii) compute analitically how the sensitivity to diffeomorphisms and noise scales with depth due to spatial pooling. The scalings are found to depend on the presence of strides in the net architecture. We find that the increased sensitivity to noise is due to the perturbing noise piling up during pooling, after being rectified by ReLU units.

cs.LG

Avalanches and deformation in glasses and disordered systems

In this chapter, we discuss avalanches in glasses and disordered systems, and the macroscopic dynamical behavior that they mediate. We briefly review three classes of systems where avalanches are observed: depinning transition of disordered interfaces, yielding of amorphous materials, and the jamming transition. Without extensive formalism, we discuss results gleaned from theoretical approaches -- mean-field theory, scaling and exponent relations, the renormalization group, and a few results from replica theory. We focus both on the remarkably sophisticated physics of avalanches and on relatively new approaches to the macroscopic flow behavior exhibited past the depinning/yielding transition.

cond-mat.dis-nn

What Can Be Learnt With Wide Convolutional Neural Networks?

Understanding how convolutional neural networks (CNNs) can efficiently learn high-dimensional functions remains a fundamental challenge. A popular belief is that these models harness the local and hierarchical structure of natural data such as images. Yet, we lack a quantitative understanding of how such structure affects performance, e.g., the rate of decay of the generalisation error with the number of training samples. In this paper, we study infinitely-wide deep CNNs in the kernel regime. First, we show that the spectrum of the corresponding kernel inherits the hierarchical structure of the network, and we characterise its asymptotics. Then, we use this result together with generalisation bounds to prove that deep CNNs adapt to the spatial scale of the target function. In particular, we find that if the target function depends on low-dimensional subsets of adjacent input variables, then the decay of the error is controlled by the effective dimensionality of these subsets. Conversely, if the target function depends on the full set of input variables, then the error decay is controlled by the input dimension. We conclude by computing the generalisation error of a deep CNN trained on the output of another deep CNN with randomly-initialised parameters. Interestingly, we find that, despite their hierarchical structure, the functions generated by infinitely-wide deep CNNs are too rich to be efficiently learnable in high dimension.

stat.ML

Learning sparse features can lead to overfitting in neural networks

It is widely believed that the success of deep networks lies in their ability to learn a meaningful representation of the features of the data. Yet, understanding when and how this feature learning improves performance remains a challenge: for example, it is beneficial for modern architectures trained to classify images, whereas it is detrimental for fully-connected networks trained for the same task on the same data. Here we propose an explanation for this puzzle, by showing that feature learning can perform worse than lazy training (via random feature kernel or the NTK) as the former can lead to a sparser neural representation. Although sparsity is known to be essential for learning anisotropic data, it is detrimental when the target function is constant or smooth along certain directions of input space. We illustrate this phenomenon in two settings: (i) regression of Gaussian random functions on the d-dimensional unit sphere and (ii) classification of benchmark datasets of images. For (i), we compute the scaling of the generalization error with number of training points, and show that methods that do not learn features generalize better, even when the dimension of the input space is large. For (ii), we show empirically that learning features can indeed lead to sparse and thereby less smooth representations of the image predictors. This fact is plausibly responsible for deteriorating the performance, which is known to be correlated with smoothness along diffeomorphisms.

stat.ML

Scaling theory for the statistics of slip at frictional interfaces

Slip at a frictional interface occurs via intermittent events. Understanding how these events are nucleated, can propagate, or stop spontaneously remains a challenge, central to earthquake science and tribology. In the absence of disorder, rate-and-state approaches predict a diverging nucleation length at some stress $\sigma^*$, beyond which cracks can propagate. Here we argue for a flat interface that disorder is a relevant perturbation to this description. We justify why the distribution of slip contains two parts: a powerlaw corresponding to `avalanches', and a `narrow' distribution of system-spanning `fracture' events. We derive novel scaling relations for avalanches, including a relation between the stress drop and the spatial extension of a slip event. We compute the cut-off length beyond which avalanches cannot be stopped by disorder, leading to a system-spanning fracture, and successfully test these predictions in a minimal model of frictional interfaces.

cond-mat.dis-nn

Failure and success of the spectral bias prediction for Kernel Ridge Regression: the case of low-dimensional data

Recently, several theories including the replica method made predictions for the generalization error of Kernel Ridge Regression. In some regimes, they predict that the method has a `spectral bias': decomposing the true function $f^*$ on the eigenbasis of the kernel, it fits well the coefficients associated with the O(P) largest eigenvalues, where $P$ is the size of the training set. This prediction works very well on benchmark data sets such as images, yet the assumptions these approaches make on the data are never satisfied in practice. To clarify when the spectral bias prediction holds, we first focus on a one-dimensional model where rigorous results are obtained and then use scaling arguments to generalize and test our findings in higher dimensions. Our predictions include the classification case $f(x)=$sign$(x_1)$ with a data distribution that vanishes at the decision boundary $p(x)\sim x_1^{\chi}$. For $\chi>0$ and a Laplace kernel, we find that (i) there exists a cross-over ridge $\lambda^*_{d,\chi}(P)\sim P^{-\frac{1}{d+\chi}}$ such that for $\lambda\gg \lambda^*_{d,\chi}(P)$, the replica method applies, but not for $\lambda\ll\lambda^*_{d,\chi}(P)$, (ii) in the ridge-less case, spectral bias predicts the correct training curve exponent only in the limit $d\rightarrow\infty$.

cs.LG