arXiv ScienceSearch

arXiv subjects

Dmitry Vaintrob

Publications and source records attributed to Dmitry Vaintrob.

15 recordsLinked to original sources

Width-Robust Learnability in Mean-Field Bayesian Neural Networks

Infinite-width limits are a standard way to reason about neural networks, but it is not automatic that the limiting learner has the same complexity-theoretic inductive bias as large finite networks. We study this question for Bayesian neural networks at the mean-field, or critical feature-learning, scaling. The central quantity is the \emph{reduced entropy} \[ s_\infty(y,\varepsilon)=\limsup_N -\frac{1}{N}\log \pi_N^0(L\le \varepsilon), \] the intensive prior cost of representing a target function $y$ to population mean-squared error $\varepsilon$. Our main result is a width-robust learnability theorem. At fixed depth, a family of Boolean-cube targets is learnable from polynomially many samples at infinite width if and only if it is learnable at polynomial width, if and only if its reduced entropy is polynomially bounded. Equivalently, up to polynomial slack in accuracy, the Bayesian mean-field learner generalizes exactly on the targets that can be represented by polynomial-size networks. The forward direction is proved by a form of subsampling: from the infinitely many hidden neurons in the mean-field solution, one can select polynomially many representatives and still preserve the learned function on every input simultaneously. At the critical scaling this subsampling has both an ``active'' component, which keeps the data-dependent low-dimensional statistics, and a ``lazy'' component, which resamples the entropy-dominated directions from the prior. Thus the infinite-width mean-field limit gives a clean analytic description of learning without introducing spurious width-dependent generalization power.

stat.ML

Towards Worst-Case Guarantees with Scale-Aware Interpretability

Neural networks organize information according to the hierarchical, multi-scale structure of natural data. Methods to interpret model internals should be similarly scale-aware, explicitly tracking how features compose across resolutions and guaranteeing bounds on the influence of fine-grained structure that is discarded as irrelevant noise. We posit that the renormalisation framework from physics can meet this need by offering technical tools that can overcome limitations of current methods. Moreover, relevant work from adjacent fields has now matured to a point where scattered research threads can be synthesized into practical, theory-informed tools. To combine these threads in an AI safety context, we propose a unifying research agenda -- \emph{scale-aware interpretability} -- to develop formal machinery and interpretability tools that have robustness and faithfulness properties supported by statistical physics.

hep-th

Mathematical Models of Computation in Superposition

Superposition -- when a neural network represents more ``features'' than it has dimensions -- seems to pose a serious challenge to mechanistically interpreting current AI systems. Existing theory work studies \emph{representational} superposition, where superposition is only used when passing information through bottlenecks. In this work, we present mathematical models of \emph{computation} in superposition, where superposition is actively helpful for efficiently accomplishing the task. We first construct a task of efficiently emulating a circuit that takes the AND of the $\binom{m}{2}$ pairs of each of $m$ features. We construct a 1-layer MLP that uses superposition to perform this task up to $\varepsilon$-error, where the network only requires $\tilde{O}(m^{\frac{2}{3}})$ neurons, even when the input features are \emph{themselves in superposition}. We generalize this construction to arbitrary sparse boolean circuits of low depth, and then construct ``error correction'' layers that allow deep fully-connected networks of width $d$ to emulate circuits of width $\tilde{O}(d^{1.5})$ and \emph{any} polynomial depth. We conclude by providing some potential applications of our work for interpreting neural networks that implement computation in superposition.

cs.LG

Sylow theorems for supergroups

We introduce Sylow subgroups and $0$-groups to the theory of complex algebraic supergroups, which mimic Sylow subgroups and $p$-groups in the theory of finite groups. We prove that Sylow subgroups are always $0$-groups, and show that they are unique up to conjugacy. Further, we give an explicit classification of $0$-groups which will be very useful for future applications. Finally, we prove an analogue of Sylow's third theorem on the number of Sylow subgroups of a supergroup.

math.RT

Formality of little disks and algebraic geometry

We construct a canonical chain of formality quasiisomorphisms for the operad of chains on framed little disks and the operad of chains on little disks. The construction is done in terms of logarithmic algebraic geometry and is remarkable for being rational (and indeed definable integrally) in de Rham cohomology.

math.AT

The Deligne-Mumford operad as a trivialization of the circle action

We prove that the tree-like Deligne-Mumford operad is a homotopical model for the trivialization of the circle in the higher-genus framed little discs operad. Our proof is based on a geometric argument involving nodal annuli. We use as a model for the higher-genus framed little discs an operad of Riemann surfaces with analytically parametrized boundary. We develop the formalism of topological moduli problems as a framework to accommodate the orbifold nature of the Deligne-Mumford operad.

math.AT

Moduli of framed formal curves

We introduce framed formal curves, which are formal algebraic curves with boundary components parametrized by the punctured formal disk. We study the moduli space of nodal framed formal curves, which we endow with a logarithmic structure. We show that this moduli space is a smooth formal logarithmic stack. The remarkable property of our construction is that framed formal curves admit a natural operation of "gluing along the boundary" which works well in families and preserves smoothness (both in a formal and in a logarithmic sense), and this induces gluing maps on the level of moduli. Using moduli spaces of framed formal curves we enhance the operad E2 of little disks (as well as its cousin, the framed little disks operad) to a fully log motivic operad. We use this structure to obtain a purely algebro-geometric proof of the formality of chains on these classical operads (initially proven for little disks by Tamarkin using analytic methods). We also recover and extend the known Galois action on the l-adic cohomology of framed and unframed little disks, and on Drinfeld associators, and extend it to the action on integral chains of a larger group scheme: the logarithmic motivic Galois group. Our methods generalize to a higher genus context, giving new "motivic" enrichments (for example, action by the absolute Galois group and by the log motivic Galois group) on the operad of chains in the oriented geometric bordism operad of Ayala and Lurie, which encodes the algebraic structure on the Hochschild cochains of any fully dualizable DG category.

math.AG

Categorical Logarithmic Hodge Theory, I

We write down a new "logarithmic" quasicoherent category $\operatorname{Qcoh}_{log}(U, X, D)$ attached to a smooth open algebraic variety $U$ with toroidal compactification $X$ and boundary divisor $D$. This is a (large) symmetric monoidal Abelian category, which we argue can be thought of as the categorical substrate for logarithmic Hodge theory of $U$. We show that its Hochschild homology theory coincides with the theory of log-forms on $X$ with logarithmic structure induced by $D$, and in particular, that the noncommutative Hodge-to de Rham sequence on $\operatorname{Qcoh}_{log}(U, X, D)$ recovers known log Hodge structure on the de Rham cohomology of the open variety $U$. As an application, we compute the Hochschild homology of the category of coherent sheaves on the infinite root stack of Talpo and Vistoli in the toroidal setting. We prove a derived invariance result for this theory: namely, that strictly toroidal changes of compactification do not change the derived category of $\operatorname{Qcoh}_{log}(U, X, D)$. The definition is motivated by the coherent object appearing in the author's microlocal mirror symmetry result [20]. In this paper, the first in a series, we work over an algebraically closed field of characteristic zero. The next installment will develop the characteristic p and mixed-characteristic theories.

math.AG

The Gauss-Manin connection on the periodic cyclic homology

It is expected that the periodic cyclic homology of a DG algebra over the field of complex numbers (and, more generally, the periodic cyclic homology of a DG category) carries a lot of additional structure similar to the mixed Hodge structure on the de Rham cohomology of algebraic varieties. Whereas a construction of such a structure seems to be out of reach at the moment its counterpart in finite characteristic is much better understood thanks to recent groundbreaking works of Kaledin. In particular, it is proven by Kaledin that under some assumptions on a DG algebra $A$ over a perfect field $k$ of characteristic $p$, a lifting of $A$ over the ring of second Witt vectors $W_2(k)$ specifies the structure of a Fontaine-Laffaille module on the periodic cyclic homology of $A$. The purpose of this paper is to develop a relative version of Kaledin's theory for DG algebras over a base $k$-algebra $R$ incorporating in the picture the Gauss-Manin connection on the relative periodic cyclic homology constructed by Getzler. Our main result asserts that, under some assumptions on $A$, the Gauss-Manin connection on its periodic cyclic homology can be recovered from the Hochschild homology of $A$ equipped with the action of the Kodaira-Spencer operator as the inverse Cartier transform (in the sense of Ogus-Vologodsky). As an application, we prove, using the reduction modulo $p$ technique, that, for a smooth and proper DG algebra over a complex punctured disk, the monodromy of the Gauss-Manin connection on its periodic cyclic homology is quasi-unipotent.

math.AG

Determinants of Subquotients of Galois Representations Associated to Abelian Varieties

Given an abelian variety $A$ of dimension $g$ over a number field $K$, and a prime $\ell$, the $\ell^n$-torsion points of $A$ give rise to a representation $ρ_{A, \ell^n} : \gal(\bar{K} / K) \to \gl_{2g}(\zz/\ell^n\zz)$. In particular, we get a mod-$\ell$ representation $ρ_{A, \ell} : \gal(\bar{K} / K) \to \gl_{2g}(\ff_\ell)$and an $\ell$-adic representation $ρ_{A, \ell} : \gal(\bar{K} / K) \to \gl_{2g}(\zz_\ell)$. In this paper, we describe the possible determinants of subrepresentations (or more generally, subquotients) of these two representation for $\ell$ a prime number, as $A$ varies over all $g$-dimensional abelian varieties. Note that it is certainly not the case that any mod-$\ell$ subquotient lifts to an $\ell$-adic one. Nevertheless, the list of possible mod-$\ell$ characters turns out to be remarkably similar to the list of possible $\ell$-adic characters.

math.NT

On the Surjectivity of Galois Representations Associated to Elliptic Curves over Number Fields

Given an elliptic curve $E$ over a number field $K$, the $\ell$-torsion points $E[\ell]$ of $E$ define a Galois representation $\gal(\bar{K}/K) \to \gl_2(\ff_\ell)$. A famous theorem of Serre states that as long as $E$ has no Complex Multiplication (CM), the map $\gal(\bar{K}/K) \to \gl_2(\ff_\ell)$ is surjective for all but finitely many $\ell$. We say that a prime number $\ell$ is exceptional (relative to the pair $(E,K)$) if this map is not surjective. Here we give a new bound on the largest exceptional prime, as well as on the product of all exceptional primes of $E$. We show in particular that conditionally on the Generalized Riemann Hypothesis (GRH), the largest exceptional prime of an elliptic curve $E$ without CM is no larger than a constant (depending on $K$) times $\log N_E$, where $N_E$ is the absolute value of the norm of the conductor. This answers affirmatively a question of Serre.

math.NT

Introduction to representation theory

These are lecture notes that arose from a representation theory course given by the first author to the remaining six authors in March 2004 within the framework of the Clay Mathematics Institute Research Academy for high school students, and its extended version given by the first author to MIT undergraduate math students in the Fall of 2008. The notes cover a number of standard topics in representation theory of groups, Lie algebras, and quivers, and contain many problems and exercises. They should be accessible to students with a strong background in linear algebra and a basic knowledge of abstract algebra, and may be used for an undergraduate or introductory graduate course in representation theory.

math.RT

The string topology BV algebra, Hochschild cohomology and the Goldman bracket on surfaces

In 1999 Chas and Sullivan discovered that the homology H_*(LX) of the space of free loops on a closed oriented smooth manifold X has a rich algebraic structure called string topology. They proved that H_*(LX) is naturally a Batalin-Vilkovisky (BV) algebra. There are several conjectures connecting the string topology BV algebra with algebraic structures on the Hochschild cohomology of algebras related to the manifold X, but none of them has been verified for manifolds of dimension n>1. In this work we study string topology in the case when X is aspherical (i.e. its homotopy groups π_i(X) vanish for i > 1). In this case the Hochschild cohomology Gerstenhaber algebra HH^*(A) of the group algebra A of the fundamental group of X has a BV structure. Our main result is a theorem establishing a natural isomorphism between the Hochschild cohomology BV algebra HH^*(A) and the string topology BV algebra H_*(LX). In particular, for a closed oriented surface X of hyperbolic type we obtain a complete description of the BV algebra operations on H_*(LX) and HH^*(A) in terms of the Goldman bracket of loops on X. The only manifolds for which the BV algebra structure on H_*(LX) was known before were spheres and complex Stiefel manifolds. Our proof is based on a combination of topological and algebraic constructions allowing us to compute and compare multiplications and BV operators on both H_*(LX) and HH^*(A).

math.AT