arXiv ScienceSearch

arXiv subjects

Mason Kamb

Publications and source records attributed to Mason Kamb.

6 recordsLinked to original sources

An exact information theory of generalization phase transitions in Bayesian diffusion models

How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead of memorizing it, remains a fundamental mystery. To address this, we introduce analytically tractable Bayesian information restricted diffusion (BIRD) models, in which each pixel observes restricted information about noisy data. A BIRD model time-reverses diffusion by inferring which past training sample produced its current restricted observation using the Bayesian posterior. This model class generalizes existing analytical diffusion models that use spatially local information restriction. We show that spatially local BIRD models closely approximate trained diffusion models \textit{early in training}, across different architectures such as UNets and DiTs. Under minimal assumptions on the data distribution, we identify an information-theoretic phase boundary between memorization and generalization in the joint space of amount of training data, time in the reverse generative process, and amount of information restriction: a BIRD model memorizes when the mutual information between its restricted noisy observations and the training data exceeds the log number of training points, and it generalizes otherwise. Experiments across a range of datasets confirm our theoretically predicted location for the transition. We find that generation proceeds near the edge of memorization: both spatially local BIRD models and early-training diffusion models track the memorization-generalization phase boundary by increasingly restricting information over time. Overall, our results reveal a fundamental role for information restriction in generative AI to circumvent the curse of dimensionality.

cs.LG

There Will Be a Scientific Theory of Deep Learning

In this paper, we make the case that a scientific theory of deep learning is emerging. By this we mean a theory which characterizes important properties and statistics of the training process, hidden representations, final weights, and performance of neural networks. We pull together major strands of ongoing research in deep learning theory and identify five growing bodies of work that point toward such a theory: (a) solvable idealized settings that provide intuition for learning dynamics in realistic systems; (b) tractable limits that reveal insights into fundamental learning phenomena; (c) simple mathematical laws that capture important macroscopic observables; (d) theories of hyperparameters that disentangle them from the rest of the training process, leaving simpler systems behind; and (e) universal behaviors shared across systems and settings which clarify which phenomena call for explanation. Taken together, these bodies of work share certain broad traits: they are concerned with the dynamics of the training process; they primarily seek to describe coarse aggregate statistics; and they emphasize falsifiable quantitative predictions. We argue that the emerging theory is best thought of as a mechanics of the learning process, and suggest the name learning mechanics. We discuss the relationship between this mechanics perspective and other approaches for building a theory of deep learning, including the statistical and information-theoretic perspectives. In particular, we anticipate a symbiotic relationship between learning mechanics and mechanistic interpretability. We also review and address common arguments that fundamental theory will not be possible or is not important. We conclude with a portrait of important open directions in learning mechanics and advice for beginners. We host further introductory materials, perspectives, and open questions at learningmechanics.pub.

stat.ML

CMT-Benchmark: A Benchmark for Condensed Matter Theory Built by Expert Researchers

Large language models (LLMs) have shown remarkable progress in coding and math problem-solving, but evaluation on advanced research-level problems in hard sciences remains scarce. To fill this gap, we present CMT-Benchmark, a dataset of 50 problems covering condensed matter theory (CMT) at the level of an expert researcher. Topics span analytical and computational approaches in quantum many-body, and classical statistical mechanics. The dataset was designed and verified by a panel of expert researchers from around the world. We built the dataset through a collaborative environment that challenges the panel to write and refine problems they would want a research assistant to solve, including Hartree-Fock, exact diagonalization, quantum/variational Monte Carlo, density matrix renormalization group (DMRG), quantum/classical statistical mechanics, and model building. We evaluate LLMs by programmatically checking solutions against expert-supplied ground truth. We developed machine-grading, including symbolic handling of non-commuting operators via normal ordering. They generalize across tasks too. Our evaluations show that frontier models struggle with all of the problems in the dataset, highlighting a gap in the physical reasoning skills of current LLMs. Notably, experts identified strategies for creating increasingly difficult problems by interacting with the LLMs and exploiting common failure modes. The best model, GPT5, solves 30\% of the problems; average across 17 models (GPT, Gemini, Claude, DeepSeek, Llama) is 11.4\pm2.1\%. Moreover, 18 problems are solved by none of the 17 models, and 26 by at most one. These unsolved problems span Quantum Monte Carlo, Variational Monte Carlo, and DMRG. Answers sometimes violate fundamental symmetries or have unphysical scaling dimensions. We believe this benchmark will guide development toward capable AI research assistants and tutors.

cs.LG

An analytic theory of creativity in convolutional diffusion models

We obtain an analytic, interpretable and predictive theory of creativity in convolutional diffusion models. Indeed, score-matching diffusion models can generate highly original images that lie far from their training data. However, optimal score-matching theory suggests that these models should only be able to produce memorized training examples. To reconcile this theory-experiment gap, we identify two simple inductive biases, locality and equivariance, that: (1) induce a form of combinatorial creativity by preventing optimal score-matching; (2) result in fully analytic, completely mechanistically interpretable, local score (LS) and equivariant local score (ELS) machines that, (3) after calibrating a single time-dependent hyperparameter can quantitatively predict the outputs of trained convolution only diffusion models (like ResNets and UNets) with high accuracy (median $r^2$ of $0.95, 0.94, 0.94, 0.96$ for our top model on CIFAR10, FashionMNIST, MNIST, and CelebA). Our model reveals a locally consistent patch mosaic mechanism of creativity, in which diffusion models create exponentially many novel images by mixing and matching different local training set patches at different scales and image locations. Our theory also partially predicts the outputs of pre-trained self-attention enabled UNets (median $r^2 \sim 0.77$ on CIFAR10), revealing an intriguing role for attention in carving out semantic coherence from local patch mosaics.

cs.LG

The Anthropocene by the Numbers: A Quantitative Snapshot of Humanity's Influence on the Planet

The presence and action of humans on Earth has exerted a strong influence on the evolution of the planet over the past $\approx$ 10,000 years, the consequences of which are now becoming broadly evident. Despite a deluge of tightly-focused and necessarily technical studies exploring each facet of "human impacts" on the planet, their integration into a complete picture of the human-Earth system lags far behind. Here, we quantify twelve dimensionless ratios which put the magnitude of human impacts in context, comparing the magnitude of anthropogenic processes to their natural analogues. These ratios capture the extent to which humans alter the terrestrial surface, hydrosphere, biosphere, atmosphere, and biogeochemistry of Earth. In almost all twelve cases, the impact of human processes rivals or exceeds their natural counterparts. The values and corresponding uncertainties for these impacts at global and regional resolution are drawn from the primary scientific literature, governmental and international databases, and industry reports. We present this synthesis of the current "state of affairs" as a graphical snapshot designed to be used as a reference. Furthermore, we establish a searchable database termed the Human Impacts Database (www.anthroponumbers.org) which houses all quantities reported here and many others with extensive curation and annotation. While necessarily incomplete, this work collates and contextualizes a set of essential numbers summarizing the broad impacts of human activities on Earth's atmosphere, land, water, and biota.

physics.soc-ph

Time-Delay Observables for Koopman: Theory and Applications

Nonlinear dynamical systems are ubiquitous in science and engineering, yet analysis and prediction of these systems remains a challenge. Koopman operator theory circumvents some of these issues by considering the dynamics in the space of observable functions on the state, in which the dynamics are intrinsically linear and thus amenable to standard techniques from numerical analysis and linear algebra. However, practical issues remain with this approach, as the space of observables is infinite-dimensional and selecting a subspace of functions in which to accurately represent the system is a nontrivial task. In this work we consider time-delay observables to represent nonlinear dynamics in the Koopman operator framework. We prove the surprising result that Koopman operators for different systems admit universal (system-independent) representations in these coordinates, and give analytic expressions for these representations. In addition, we show that for certain systems a restricted class of these observables form an optimal finite-dimensional basis for representing the Koopman operator, and that the analytic representation of the Koopman operator in these coordinates coincides with results computed by the dynamic mode decomposition. We provide numerical examples to complement our results. In addition to being theoretically interesting, these results have implications for a number of linearization algorithms for dynamical systems.

math.NA