arXiv ScienceSearch

arXiv subjects

David S. Berman

Publications and source records attributed to David S. Berman.

At least 19 recordsLinked to original sources

Do Quantum Models Scale Like LLMs?

In this work, we study the neural scaling laws of RydbergGPT, an autoregressive transformer model trained on qubit projective measurement data gathered from interacting Rydberg atom arrays. The quantum system is known to exhibit a finite-size remnant of a critical point as the laser detuning parameter is varied. We find that near the critical point the transformer loss as a function of training dataset size is well described by a power-law with a loss floor correction. However, away from criticality the quality of the power-law description is substantially reduced. We then compare the statistical structure of both Rydberg measurements and natural-language corpora using an entropy-normalised, finite sample corrected mutual information "two-point" function. We find that near-critical statistics of the two point functions are closest to those observed in natural-language, whilst other qubit configurations far from the critical point have two-point functions that decay more rapidly. This supports the hypothesis that multi-scale dependence contributes to stable neural scaling, and that scaling behaviour should be viewed as a property of the model-data pair.

cs.LG

Optimal Pruning for Neural Architectures using Fisher Information Distances

A new scheme for parameter pruning is introduced, derived from the differential-geometric distance in model space. Pruning a parameter sets its value to zero, representing a displacement of the model to the hypersurface on which that parameter vanishes. The minimal distance from the unpruned model to this hypersurface is naturally computed via the geodesic distance in the model space as determined by the Fisher information metric. This distance determines the true change in the model, and its performance, under pruning. By analysing progressively more faithful approximations of this geodesic distance a natural hierarchy of optimality for pruning methods is determined. This starts with the traditional magnitude pruning, then develops into new more sophisticated and effective pruning schemes. The method is demonstrated for both fully-connected networks and vision transformers, on MNIST and CIFAR-10, over the complete $0$-$100\%$ pruning range and across five random seeds. It outperforms pruning by parameter magnitude and by the local Fisher information alone in every architecture and dataset combination considered, on both accuracy and the Matthews correlation coefficient. Additionally, analysis of different levels of geodesic approximation produces intermediate pruning schemes that are computationally efficient and maintain near-optimal performance. This geometric picture supplies not only a state-of-the-art pruning methodology for AI models, but also a verified and mathematically-motivated justification for pruning schemes.

cs.AI

Modeling financial time series with $ϕ^{4}$ quantum field theory

We use a $ϕ^{4}$ quantum field theory with inhomogeneous couplings and explicit symmetry-breaking to model an ensemble of financial time series from the S$\&$P 500 index. The continuum nature of the $ϕ^4$ theory avoids the inaccuracies that occur in Ising-based models which require a discretization of the time series. We demonstrate this using the example of the 2008 global financial crisis. The $ϕ^{4}$ quantum field theory is expressive enough to reproduce the higher-order statistics such as the market kurtosis, which can serve as an indicator of possible market shocks. Accurate reproduction of high kurtosis is absent in binarized models. Therefore Ising models, despite being widely employed in econophysics, are incapable of fully representing empirical financial data, a limitation not present in the generalization of the $ϕ^{4}$ scalar field theory. We then investigate the scaling properties of the $ϕ^{4}$ machine learning algorithm and extract exponents which govern the behavior of the learned couplings (or weights and biases in ML language) in relation to the number of stocks in the model. Finally, we use our model to forecast the price changes of the AAPL, MSFT, and NVDA stocks. We conclude by discussing how the $ϕ^{4}$ scalar field theory could be used to build investment strategies and the possible intuitions that the QFT operations of dimensional compactification and renormalization can provide for financial modelling.

q-fin.ST

AI and the Research-Education Environment of Physics

In the current era of AI transforming the research-education environment of physics, variety of issues and concerns arise. The KITP program "Generative AI for High and Low Energy Physics'' offered a discussion session on this, and here presented is a summary of the opinions provided in the discussion. The material is formulated such that it can serve as a starting point for further discussions in readers' research community/institution/group.

physics.ed-ph

A path to natural language through tokenisation and transformers

Natural languages exhibit striking regularities in their statistical structure, including notably the emergence of Zipf's and Heaps' laws. Despite this, it remains broadly unclear how these properties relate to the modern tokenisation schemes used in contemporary transformer models. In this note, we analyse the information content (as measured by the Shannon entropy) of various corpora under the assumption of a Zipfian frequency distribution, and derive a closed-form expression for the slot entropy expectation value. We then empirically investigate how byte--pair encoding (BPE) transforms corpus statistics, showing that recursive applications of BPE drive token frequencies toward a Zipfian power law while inducing a characteristic growth pattern in empirical entropy. Utilizing the ability of transformers to learn context dependent token probability distributions, we train language models on corpora tokenised at varying BPE depths, revealing that the model predictive entropies increasingly agree with Zipf-derived predictions as the BPE depth increases. Attention-based diagnostics further indicate that deeper tokenisation reduces local token dependencies, bringing the empirical distribution closer to the weakly dependent (near IID) regime. Together, these results clarify how BPE acts not only as a compression mechanism but also as a statistical transform that reconstructs key informational properties of natural language.

cs.CL

NCoder -- A Quantum Field Theory approach to encoding data

In this paper we present a novel approach to interpretable AI inspired by Quantum Field Theory (QFT) which we call the NCoder. The NCoder is a modified autoencoder neural network whose latent layer is prescribed to be a subset of $n$-point correlation functions. Regarding images as draws from a lattice field theory, this architecture mimics the task of perturbatively constructing the effective action of the theory order by order in an expansion using Feynman diagrams. Alternatively, the NCoder may be regarded as simulating the procedure of statistical inference whereby high dimensional data is first summarized in terms of several lower dimensional summary statistics (here the $n$-point correlation functions), and subsequent out-of-sample data is generated by inferring the data generating distribution from these statistics. In this way the NCoder suggests a fascinating correspondence between perturbative renormalizability and the sufficiency of models. We demonstrate the efficacy of the NCoder by applying it to the generation of MNIST images, and find that generated images can be correctly classified using only information from the first three $n$-point functions of the image distribution.

hep-th

Grokking vs. Learning: Same Features, Different Encodings

Grokking typically achieves similar loss to ordinary, "steady", learning. We ask whether these different learning paths - grokking versus ordinary training - lead to fundamental differences in the learned models. To do so we compare the features, compressibility, and learning dynamics of models trained via each path in two tasks. We find that grokked and steadily trained models learn the same features, but there can be large differences in the efficiency with which these features are encoded. In particular, we find a novel "compressive regime" of steady training in which there emerges a linear trade-off between model loss and compressibility, and which is absent in grokking. In this regime, we can achieve compression factors 25x times the base model, and 5x times the compression achieved in grokking. We then track how model features and compressibility develop through training. We show that model development in grokking is task-dependent, and that peak compressibility is achieved immediately after the grokking plateau. Finally, novel information-geometric measures are introduced which demonstrate that models undergoing grokking follow a straight path in information space.

cs.LG

Curvature of an exotic 7-sphere

We study the geometry of the Gromoll-Meyer sphere, one of Milnor's exotic $7$-spheres. We focus on a Kaluza-Klein Ansatz, with a round $S^4$ as base space, unit $S^3$ as fibre, and $k=1,2$ $SU(2)$ instantons as gauge fields, where all quantities admit an elegant description in quaternionic language. The metric's moduli space coincides with the $k=1,2$ instantons' moduli space quotiented by the isometry of the base, plus an additional $\mathbb{R}^+$ factor corresponding to the radius of the base, $r$. We identify a "center" of the $k=2$ instanton moduli space with enhanced symmetry. This $k=2$ solution is used together with the maximally symmetric $k=1$ solution to obtain a metric of maximal isometry, $SO(3)\times O(2)$, and to explicitly compute its Ricci tensor. This allows us to put a bound on $r$ to ensure positive Ricci curvature, which implies various energy conditions for an $8$-dimensional static space-time. This construction then enables a concrete examination of the properties of the sectional curvature.

hep-th

The Geometry, Branes and Applications of Exceptional Field Theory

This is a review of exceptional field theory: a generalisation of Kaluza-Klein theory that unifies the metric and $p$-form gauge field degrees of freedom of supergravity into a generalised or extended geometry, whose additional coordinates may be viewed as conjugate to brane winding modes. This unifies the maximal supergravities, treating their previously-hidden exceptional Lie symmetries as a fundamental geometric symmetry. Duality orbits of solutions simplify into single objects, that in many cases have simple geometric interpretations, for instance as wave or monopole-type solutions. It also provides a route to explore exotic or non-geometric aspects of M-theory, such as exotic branes, U-folds, and more novel sorts of non-Riemannian spaces.

hep-th

The Inverse of Exact Renormalization Group Flows as Statistical Inference

We build on the view of the Exact Renormalization Group (ERG) as an instantiation of Optimal Transport described by a functional convection-diffusion equation. We provide a new information theoretic perspective for understanding the ERG through the intermediary of Bayesian Statistical Inference. This connection is facilitated by the Dynamical Bayesian Inference scheme, which encodes Bayesian inference in the form of a one parameter family of probability distributions solving an integro-differential equation derived from Bayes' law. In this note, we demonstrate how the Dynamical Bayesian Inference equation is, itself, equivalent to a diffusion equation which we dub Bayesian Diffusion. Identifying the features that define Bayesian Diffusion, and mapping them onto the features that define the ERG, we obtain a dictionary outlining how renormalization can be understood as the inverse of statistical inference.

hep-th

Bayesian Renormalization

In this note we present a fully information theoretic approach to renormalization inspired by Bayesian statistical inference, which we refer to as Bayesian Renormalization. The main insight of Bayesian Renormalization is that the Fisher metric defines a correlation length that plays the role of an emergent RG scale quantifying the distinguishability between nearby points in the space of probability distributions. This RG scale can be interpreted as a proxy for the maximum number of unique observations that can be made about a given system during a statistical inference experiment. The role of the Bayesian Renormalization scheme is subsequently to prepare an effective model for a given system up to a precision which is bounded by the aforementioned scale. In applications of Bayesian Renormalization to physical systems, the emergent information theoretic scale is naturally identified with the maximum energy that can be probed by current experimental apparatus, and thus Bayesian Renormalization coincides with ordinary renormalization. However, Bayesian Renormalization is sufficiently general to apply even in circumstances in which an immediate physical scale is absent, and thus provides an ideal approach to renormalization in data science contexts. To this end, we provide insight into how the Bayesian Renormalization scheme relates to existing methods for data compression and data generation such as the information bottleneck and the diffusion learning paradigm. We conclude by designing an explicit form of Bayesian Renormalization inspired by Wilson's momentum shell renormalization scheme in Quantum Field Theory. We apply this Bayesian Renormalization scheme to a simple Neural Network and verify the sense in which it organizes the parameters of the model according to a hierarchy of information theoretic importance.

hep-th

Twisted Self-duality

We examine a generalisation of the usual self-duality equations for Yang-Mills theory when the colour space admits a non-trivial involution. This involution allows us to construct a non-trivial twist which may be combined with the Hodge star to form a twisted self-dual curvature. We will construct a simple example of twisted self-duality for $su(2) \oplus su(2)$ gauge theory along with its explicit solutions and then dimensionally reduce from four dimensions to obtain families of non-trivial non-linear equations in lower dimensions. This twisted self-duality constraint will be shown to arise in E_7 exceptional field theory through a Scherk-Schwarz reduction and we will show how an Eguchi-Hanson gravitational instanton also obeys the twisted self-duality condition.

hep-th

On the Dynamics of Inference and Learning

Statistical Inference is the process of determining a probability distribution over the space of parameters of a model given a data set. As more data becomes available this probability distribution becomes updated via the application of Bayes' theorem. We present a treatment of this Bayesian updating process as a continuous dynamical system. Statistical inference is then governed by a first order differential equation describing a trajectory or flow in the information geometry determined by a parametric family of models. We solve this equation for some simple models and show that when the Cramér-Rao bound is saturated the learning rate is governed by a simple $1/T$ power-law, with $T$ a time-like variable denoting the quantity of data. The presence of hidden variables can be incorporated in this setting, leading to an additional driving term in the resulting flow equation. We illustrate this with both analytic and numerical examples based on Gaussians and Gaussian Random Processes and inference of the coupling constant in the 1D Ising model. Finally we compare the qualitative behaviour exhibited by Bayesian flows to the training of various neural networks on benchmarked data sets such as MNIST and CIFAR10 and show how that for networks exhibiting small final losses the simple power-law is also satisfied.

cond-mat.dis-nn

Double copying Exceptional Field theories

We examine exceptional field theory through the lens of the generalised double copy formalism. This allows us to construct classical solutions in M-theory using a generalised Kerr-Schild ansatz and along the way indicates hints towards a single copy of M-theory. Based on a talk at the Nankai Symposium given by DSB.

hep-th

Machine Learning Calabi-Yau Hypersurfaces

We revisit the classic database of weighted-P4s which admit Calabi-Yau 3-fold hypersurfaces equipped with a diverse set of tools from the machine-learning toolbox. Unsupervised techniques identify an unanticipated almost linear dependence of the topological data on the weights. This then allows us to identify a previously unnoticed clustering in the Calabi-Yau data. Supervised techniques are successful in predicting the topological parameters of the hypersurface from its weights with an accuracy of R^2 > 95%. Supervised learning also allows us to identify weighted-P4s which admit Calabi-Yau hypersurfaces to 100% accuracy by making use of partitioning supported by the clustering behaviour.

hep-th

The Classical Double Copy for M-theory from a Kerr-Schild Ansatz for Exceptional Field Theory

We construct the classical double copy formalism for M-theory. This extends the current state of the art by including the three form potential of eleven dimensional supergravity along with the metric. The key for this extension is to construct a Kerr-Schild type Ansatz for exceptional field theory. This Kerr-Schild Ansatz then allows us to find the solutions of charged objects such as the membrane from a set of single copy fields. The exceptional field theory formalism then automatically produces the IIB Kerr-Schild ansatz allowing the construction of the single copy for the fields of IIB supergravity (with manifest $SL(2)$ symmetry).

hep-th

The single copy of the gravitational holonomy

The double copy is a well-established relationship between gravity and gauge theories. It relates perturbative scattering amplitudes as well as classical solutions, and recently there has been mounting evidence that it also applies to non-perturbative information. In this paper, we consider the holonomy properties of manifolds in gravity and prescribe a single copy of gravitational holonomy that differs from the holonomy in gauge theory. We discuss specific cases and give examples where the single copy holonomy group is reduced. Our results may prove useful in extending the classical double copy. We also clarify previous misconceptions in the literature regarding gravitational Wilson lines and holonomy.

hep-th

Double Field Theory and Geometric Quantisation

We examine various properties of double field theory and the doubled string sigma model in the context of geometric quantisation. In particular we look at T-duality as the symplectic transformation related to an alternative choice of polarisation in the construction of the quantum bundle for the string. Following this perspective we adopt a variety of techniques from geometric quantisation to study the doubled space. One application is the construction of the double coherent state that provides the shortest distance in any duality frame and a stringy deformed Fourier transform.

hep-th