arXiv Science⌕ Search

arXiv · 2609.31101

A Flatness-Generalization Relation in the Teacher-Student Tree-Committee Machine

Abstract

The flatness of the loss landscape at a minimizer is a widely used heuristic for reasoning about neural-network generalization, yet evidence for this relation is mostly empirical and controversial. We study this relation in a teacher-student tree committee machine, where both the ERM estimator and the Hessian spectrum are analytically tractable in the proportional high-dimensional limit. First, we use a zero-temperature Gibbs formulation to obtain predictions for the observables of the typical minimizers of the empirical loss. Secondly, we use Edwards-Jones formalism to derive the limiting Hessian resolvent around these typical minimizers. All predictions agree with finite-size gradient-descent simulations. Finally, we study three measures of flatness, namely the left and right edges and the spectral mean, and check if a decrease in generalization error as the dataset size is increased corresponds to an increase in flatness. We find that the answer strongly depends on the learning task and on the ratio of the number of parameters to the number of data points. In regression, the spectral mean and right edge correlate with the generalization error, while the left edge does so only in the overparametrized regime. In classification this correlation reliably holds only in the highly overparametrized phase, while for underparametrized networks it can even reverse.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Brandon Livio Annesi, Davide Straziota, Enrico Maria Malatesta. 2026-09-25. A Flatness-Generalization Relation in the Teacher-Student Tree-Committee Machine. https://arxiv.org/abs/2609.31101

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Locking transition in coupled disordered systems

When a system of many interacting constituents is extended across space, every region carries the same disordered landscape, and it is unclear whether distant regions settle into the same state. We study copies of the Random Energy Model (REM), sharing one disorder realization, coupled along a chain. A first-order transition separates a phase with short-range correlations, from a locked phase in which the entire chain occupies the lowest-energy state with probability one, even at positive temperature. No intermediate ordering lengths occur. In the Generalized REM the chain locks either to the lowest-energy state or to the lowest free-energy valley.

cond-mat.dis-nn↗

Pre-registered tests of solid-state-physics-inspired LLM compression: a cluster-level negative result at small-language-model scale

We report a three-month autonomous research-agent program testing five solid-state-physics-inspired compression mappings on pretrained language models, with predictions committed to git before any pilot data and a 3-sigma gate deciding PASS or SHELVE. The common anchor -- area-law / Kohn-nearsighted decay of the one-particle density matrix -- has a distance face (P001 Wannier, P002 tight-binding) and a rank face (P003 DMRG-truncated MLPs, P005 Wilson-RG, P011 tensor-train embeddings). P005 was pre-empted at Phase 1; three of four Phase-3 pilots were falsified. On the attention face, GPT-2-medium attention-versus-distance is best fit by a stretched exponential in 12 of 16 median-layer heads once probe padding is excluded, and a tight-binding cutoff costs +96% perplexity (P002); on Pythia-160M the Wannier sparsity 0.054 +/- 0.004 is indistinguishable from PCA, random-Haar and identity baselines (P001). On the rank face, per-token tensor-train bond dimension does not track surprisal (r = 0.016 vs a pre-registered 0.65) and the format inflates rather than compresses (P011). P003 is mixed: its scaling claim shelved (r = -0.434), its MPO premise died at stage-0, and its cross-paper check, r = 0.523 as first written, collapses to 0.047 under the same correction, leaving both cross-paper checks null. The results invert the pre-registered prediction that most attention heads behave like Kohn-nearsighted insulators, pointing instead to critical, glassy or heavy-tailed regimes; the inversion is specific to the <= 350M scale tested, while the rank-face no-gain result held to 7-8B. We contribute the pre-registration + 3-sigma + cluster-framing + append-only-catalogue discipline -- including why our own enforcement gate was designed but not deployed -- four pre-registered negative results with full data release, and the inversion. The catalogue holds eighteen concluded studies, seventeen negative.

cond-mat.dis-nn↗

Localization in tight-binding models with power-law distributed couplings

We study the localization properties of 1D and 2D tight-binding models with power-law distributed couplings by comparing the spectrum and the localization properties of the eigenmodes of the Laplacian and the adjacency matrix, using numerical diagonalization of these matrices for different system sizes and connectivities. These two matrices are relevant for different types of dynamical processes. While all eigenmodes of the adjacency matrix are localized for sufficiently large system sizes, the Laplacian matrix always leads to a small proportion of system-spanning modes due to a conservation law, and therefore to power-law tails in the probability distribution of the participation ratio and its relation to the eigenvalues. In one dimension, the exponent of these power laws change continuously with the exponent that characterizes the distribution of couplings. In two dimensions, the modes with the largest relaxation times change from system-spanning to localized when the exponent of the distribution of couplings becomes larger than 0.75. We provide phenomenological explanations for all these findings.

cond-mat.dis-nn↗