arXiv ScienceSearch

arXiv subjects

Sho Yaida

Publications and source records attributed to Sho Yaida.

At least 19 recordsLinked to original sources

Effective Theory of Transformers at Initialization

We perform an effective-theory analysis of forward-backward signal propagation in wide and deep Transformers, i.e., residual neural networks with multi-head self-attention blocks and multilayer perceptron blocks. This analysis suggests particular width scalings of initialization and training hyperparameters for these models. We then take up such suggestions, training Vision and Language Transformers in practical setups.

cs.LG

Meta-Principled Family of Hyperparameter Scaling Strategies

In this note, we first derive a one-parameter family of hyperparameter scaling strategies that interpolates between the neural-tangent scaling and mean-field/maximal-update scaling. We then calculate the scalings of dynamical observables -- network outputs, neural tangent kernels, and differentials of neural tangent kernels -- for wide and deep neural networks. These calculations in turn reveal a proper way to scale depth with width such that resultant large-scale models maintain their representation-learning ability. Finally, we observe that various infinite-width limits examined in the literature correspond to the distinct corners of the interconnected web spanned by effective theories for finite-width neural networks, with their training dynamics ranging from being weakly-coupled to being strongly-coupled.

cs.LG

The Principles of Deep Learning Theory

This book develops an effective theory approach to understanding deep neural networks of practical relevance. Beginning from a first-principles component-level picture of networks, we explain how to determine an accurate description of the output of trained networks by solving layer-to-layer iteration equations and nonlinear learning dynamics. A main result is that the predictions of networks are described by nearly-Gaussian distributions, with the depth-to-width aspect ratio of the network controlling the deviations from the infinite-width Gaussian description. We explain how these effectively-deep networks learn nontrivial representations from training and more broadly analyze the mechanism of representation learning for nonlinear models. From a nearly-kernel-methods perspective, we find that the dependence of such models' predictions on the underlying learning algorithm can be expressed in a simple and universal way. To obtain these results, we develop the notion of representation group flow (RG flow) to characterize the propagation of signals through the network. By tuning networks to criticality, we give a practical solution to the exploding and vanishing gradient problem. We further explain how RG flow leads to near-universal behavior and lets us categorize networks built from different activation functions into universality classes. Altogether, we show that the depth-to-width ratio governs the effective model complexity of the ensemble of trained networks. By using information-theoretic techniques, we estimate the optimal aspect ratio at which we expect the network to be practically most useful and show how residual connections can be used to push this scale to arbitrary depths. With these tools, we can learn in detail about the inductive bias of architectures, hyperparameters, and optimizers.

cs.LG

Non-Gaussian processes and neural networks at finite widths

Gaussian processes are ubiquitous in nature and engineering. A case in point is a class of neural networks in the infinite-width limit, whose priors correspond to Gaussian processes. Here we perturbatively extend this correspondence to finite-width neural networks, yielding non-Gaussian processes as priors. The methodology developed herein allows us to track the flow of preactivation distributions by progressively integrating out random variables from lower to higher layers, reminiscent of renormalization-group flow. We further develop a perturbative procedure to perform Bayesian inference with weakly non-Gaussian priors.

stat.ML

Robust Learning with Jacobian Regularization

Design of reliable systems must guarantee stability against input perturbations. In machine learning, such guarantee entails preventing overfitting and ensuring robustness of models against corruption of input data. In order to maximize stability, we analyze and develop a computationally efficient implementation of Jacobian regularization that increases classification margins of neural networks. The stabilizing effect of the Jacobian regularizer leads to significant improvements in robustness, as measured against both random and adversarial input perturbations, without severely degrading generalization properties on clean data.

stat.ML

Fluctuation-dissipation relations for stochastic gradient descent

The notion of the stationary equilibrium ensemble has played a central role in statistical mechanics. In machine learning as well, training serves as generalized equilibration that drives the probability distribution of model parameters toward stationarity. Here, we derive stationary fluctuation-dissipation relations that link measurable quantities and hyperparameters in the stochastic gradient descent algorithm. These relations hold exactly for any stationary state and can in particular be used to adaptively set training schedule. We can further use the relations to efficiently extract information pertaining to a loss-function landscape such as the magnitudes of its Hessian and anharmonicity. Our claims are empirically verified.

stat.ML

Morphology of renormalization-group flow for the de Almeida-Thouless-Gardner universality class

A replica-symmetry-breaking phase transition is predicted in a host of disordered media. The criticality of the transition has, however, long been questioned below its upper critical dimension, six, due to the absence of a critical fixed point in the renormalization-group flows at one-loop order. A recent two-loop analysis revealed a possible strong-coupling fixed point but, given the uncontrolled nature of perturbative analysis in the strong-coupling regime, debate persists. Here we examine the nature of the transition as a function of spatial dimension and show that the strong-coupling fixed point can go through a Hopf bifurcation, resulting in a critical limit cycle and a concomitant discrete scale invariance. We further investigate a different renormalization scheme and argue that the basin of attraction of the strong-coupling fixed point/limit cycle may thus stay finite for all dimensions.

cond-mat.stat-mech

Zero-temperature glass transition in two dimensions

The nature of the glass transition is theoretically understood in the mean-field limit of infinite spatial dimensions, but the problem remains totally open in physical dimensions. Nontrivial finite-dimensional fluctuations are hard to control analytically, and experiments fail to provide conclusive evidence regarding the nature of the glass transition. Here, we use Monte Carlo simulations that fully bypass the glassy slowdown, and access equilibrium states in two-dimensional glass-forming liquids at low enough temperatures to directly probe the transition. We find that the liquid state terminates at a thermodynamic glass transition at zero temperature, which is associated with an entropy crisis and a diverging static correlation length.

cond-mat.stat-mech

Configurational entropy measurements in extremely supercooled liquids that break the glass ceiling

Liquids relax extremely slowly upon approaching the glass state. One explanation is that an entropy crisis, due to the rarefaction of available states, makes it increasingly arduous to reach equilibrium in that regime. Validating this scenario is challenging, because experiments offer limited resolution, while numerical studies lag more than eight orders of magnitude behind experimentally-relevant timescales. In this work we not only close the colossal gap between experiments and simulations but manage to create in-silico configurations that have no experimental analog yet. Deploying a range of computational tools, we obtain four estimates of their configurational entropy. These measurements consistently confirm that the steep entropy decrease observed in experiments is also found in simulations, even beyond the experimental glass transition. Our numerical results thus extend the new observational window into the physics of glasses and reinforce the relevance of an entropy crisis for understanding their formation.

cond-mat.stat-mech

Cycle-expansion method for the Lyapunov exponent, susceptibility, and higher moments

Lyapunov exponents characterize the chaotic nature of dynamical systems by quantifying the growth rate of uncertainty associated with the imperfect measurement of initial conditions. Finite-time estimates of the exponent, however, experience fluctuations due to both the initial condition and the stochastic nature of the dynamical path. The scale of these fluctuations is governed by the Lyapunov susceptibility, the finiteness of which typically provides a sufficient condition for the law of large numbers to apply. Here, we obtain a formally exact expression for this susceptibility in terms of the Ruelle dynamical zeta function for one-dimensional systems. We further show that, for systems governed by sequences of random matrices, the cycle expansion of the zeta function enables systematic computations of the Lyapunov susceptibility and of its higher-moment generalizations. The method is here applied to a class of dynamical models that maps to static disordered spin chains with interactions stretching over a varying distance, and is tested against Monte Carlo simulations.

cond-mat.stat-mech

Optimizing collective fieldtaxis of swarming agents through reinforcement learning

Swarming of animal groups enthralls scientists in fields ranging from biology to physics to engineering. Complex swarming patterns often arise from simple interactions between individuals to the benefit of the collective whole. The existence and success of swarming, however, nontrivially depend on microscopic parameters governing the interactions. Here we show that a machine-learning technique can be employed to tune these underlying parameters and optimize the resulting performance. As a concrete example, we take an active matter model inspired by schools of golden shiners, which collectively conduct phototaxis. The problem of optimizing the phototaxis capability is then mapped to that of maximizing benefits in a continuum-armed bandit game. The latter problem accepts a simple reinforcement-learning algorithm, which can tune the continuous parameters of the model. This result suggests the utility of machine-learning methodology in swarm-robotics applications.

physics.bio-ph

Nontrivial critical fixed point for replica-symmetry-breaking transitions

The transformation of the free-energy landscape from smooth to hierarchical is one of the richest features of mean-field disordered systems. A well-studied example is the de Almeida-Thouless transition for spin glasses in a magnetic field, and a similar phenomenon--the Gardner transition--has recently been predicted for structural glasses. The existence of these replica-symmetry-breaking phase transitions has, however, long been questioned below their upper critical dimension, d_u=6. Here, we obtain evidence for the existence of these transitions in d<d_u using a two-loop calculation. Because the critical fixed point is found in the strong-coupling regime, we corroborate the result by resumming the perturbative series with inputs from a three-loop calculation and an analysis of its large-order behavior. Our study offers a resolution of the long-lasting controversy surrounding phase transitions in finite-dimensional disordered systems.

cond-mat.stat-mech

Point-to-set lengths, local structure, and glassiness

The growing sluggishness of glass-forming liquids is thought to be accompanied by growing structural order. The nature of such order, however, remains hotly debated. A decade ago, point-to-set (PTS) correlation lengths were proposed as measures of amorphous order in glass formers, but recent results raise doubts as to their generality. Here, we extend the definition of PTS correlations to agnostically capture any type of growing order in liquids, be it local or amorphous. This advance enables the formulation of a clear distinction between slowing down due to conventional critical ordering and that due to glassiness, and provides a unified framework to assess the relative importance of specific local order and generic amorphous order in glass formation.

cond-mat.stat-mech

Linking dynamical heterogeneity to static amorphous order

Glass-forming liquids grow dramatically sluggish upon cooling. This slowdown has long been thought to be accompanied by a growing correlation length. Characteristic dynamical and static length scales, however, have been observed to grow at different rates, which perplexes the relationship between the two and with the slowdown. Here, we show the existence of a direct link between dynamical sluggishness and static point-to-set correlations, holding at the local level as we probe different environments within a liquid. This link, which is stronger and more general than that observed with locally preferred structures, suggests the existence of an intimate relationship between structure and dynamics in a broader range of glass-forming liquids than previously thought.

cond-mat.dis-nn

Efficient measurement of point-to-set correlations and overlap fluctuations in glass-forming liquids

Cavity point-to-set correlations are real-space tools to detect the roughening of the free-energy landscape that accompanies the dynamical slowdown of glass-forming liquids. Measuring these correlations in model glass formers remains, however, a major computational challenge. Here, we develop a general parallel-tempering method that provides orders-of-magnitude improvement for sampling and equilibrating configurations within cavities. We apply this improved scheme to the canonical Kob-Andersen binary Lennard-Jones model for temperatures down to the mode-coupling theory crossover. Most significant improvements are noted for small cavities, which have thus far been the most difficult to study. This methodological advance also enables us to study a broader range of physical observables associated with thermodynamic fluctuations. We measure the probability distribution of overlap fluctuations in cavities, which displays a non-trivial temperature evolution. The corresponding overlap susceptibility is found to provide a robust quantitative estimate of the point-to-set length scale requiring no fitting. By resolving spatial fluctuations of the overlap in the cavity, we also obtain quantitative information about the geometry of overlap fluctuations. We can thus examine in detail how the penetration length as well as its fluctuations evolve with temperature and cavity size.

cond-mat.stat-mech

Glassy slowdown and replica-symmetry-breaking instantons

Glass-forming liquids exhibit a dramatic dynamical slowdown as the temperature is lowered. This can be attributed to relaxation proceeding via large structural rearrangements whose characteristic size increases as the system cools. These cooperative rearrangements are well modeled by instantons in a replica effective field theory, with the size of the dominant instanton encoding the liquid's cavity point-to-set correlation length. Varying the parameters of the effective theory corresponds to varying the statistics of the underlying free-energy landscape. We demonstrate that, for a wide range of parameters, replica-symmetry-breaking instantons dominate. The detailed structure of the dominant instanton provides a rich window into point-to-set correlations and glassy dynamics.

cond-mat.stat-mech

Effective Field Theory for Supercooled Liquids

Starting from a microscopic model of liquids, we construct an effective theory of an overlap field through duplication of the system and coarse-graining. We then propose a recipe to extract a relaxation time and two characteristic length scales of a supercooled liquid from this effective field theory. Appealing to the Ginzburg-Landau-Wilson paradigm near the putative critical point, we further conclude that this effective field theory resides within the Ising universality class.

cond-mat.dis-nn

Critical Exponents for Supercooled Liquids

We compute critical exponents governing universal features of supercooled liquids through the effective theory of an overlap field. The correlation length diverges with the Ising exponent; the size of dynamically heterogeneous patches grows more rapidly; and the relaxation time obeys a generalized Vogel-Fulcher-Tammann relation.

cond-mat.dis-nn