arXiv Science⌕ Search

arXiv · 2610.09044

Quadratic Weak-to-Strong Generalization in Random Feature Networks via Random Matrix Theory

Abstract

Weak-to-strong generalization is the phenomenon where a strong student model trained with labels produced by a weak teacher model is able to generalize better than the teacher. In this paper, we study this phenomenon in two-layer random feature networks where the model strength is determined by its width. Using tools from random matrix theory, we derive deterministic equivalents for the population errors of an optimally trained teacher and a student trained with gradient flow. For ReLU activation and a pure spherical harmonic target, we obtain sharp asymptotics under a Gaussian universality assumption, showing a quadratic improvement: the student error scales as the square of the teacher error. These results attain the general lower bound of Medvedev at al (2025). We also analyze how the student behaves under more general stopping times and targets supported on multiple harmonic degrees, characterizing the regimes in which weak-to-strong generalization occurs and identifying the transition between quadratic, non-quadratic, and no improvement.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Deborah Oliveira, Elliot Paquette. 2026-10-06. Quadratic Weak-to-Strong Generalization in Random Feature Networks via Random Matrix Theory. https://arxiv.org/abs/2610.09044

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Are Good Generators Good Decision-Makers? Policy Learning for General Interventions via Retargeted Counterfactual Generation

Generative models are increasingly used to support decision-making in complex systems, where interventions may be joint and high-dimensional, and outcomes are high-dimensional. However, using generators for these decision-making settings are challenged by three problems. First, they are often trained on noisy logs with limited intervention data. Second, generator learning can be noisy across different environments. Third, generators are not optimized for the decisions they support. We introduce policy learning via retargeted counterfactual generation, which trains a generator for the decisions it supports in three steps. We (1) learn a doubly-robust, invariant counterfactual generator for high-dimensional interventions and outcomes, (2) conduct policy-learning based on its rollouts, and (3) retarget the generator toward the learned policy's interventions and relearn the policy, so the generator is accurate where decisions are made. Theoretically, our generator's excess counterfactual risk has a doubly robust product remainder, and retargeting removes the worst-case density ratio between the logged and learned policies from the regret. We demonstrate the efficacy of our approach across synthetic data, Cell Painting images of SARS-CoV-2-infected cells, and physical video simulations.

stat.ML↗

Non-asymptotic Convergence of Stochastic Gradient Descent in Score-based Generative Models

Score-based Generative Models (SGMs) have achieved impressive performance in data generation across a wide range of applications. While the statistical properties of their sampling procedures are increasingly well understood, the optimization dynamics underlying their training remain less explored. SGMs are typically trained by minimizing a weighted denoising score-matching objective, yet optimization guarantees with stochastic gradients remain limited. In this work, we study Stochastic Gradient Descent (SGD) for SGMs, contributing results in two complementary regimes. For general score parameterizations, we derive a non-convex analysis of SGD for the weighted denoising score-matching objective, making explicit how the resulting optimization bound depends on the loss weighting and time-sampling distribution. We then consider overparameterized two-layer ReLU networks and develop a Neural Tangent Kernel analysis tailored to diffusion training with stochastic gradients, yielding score-approximation error bounds along the SGD trajectory. Our analysis quantifies the role of the reweighting factor in these bounds, providing a theoretical characterization of weighting choices used in practice.

stat.ML↗

Ordinary Nonconvex SGD under Distance-Dependent Moments: Finite-Horizon Stationarity and Nagaev Bounds

Uniform noise-moment bounds exclude stochastic gradients whose variability increases with the iterate. We study ordinary, single-sample stochastic gradient descent for smooth, lower-bounded, possibly nonconvex objectives under distance-dependent conditional moments. Under second moments alone, a direct descent--displacement argument yields $T^{-1/3}$ expected average squared-gradient stationarity with a horizon-dependent stepsize. An explicit oracle-complexity corollary matches the known smooth Blum--Gladyshev (BG-0) lower bound, including the $Lb_2Δ^3\varepsilon^{-6}$ and $LΔσ^2\varepsilon^{-4}$ stochastic terms, where $Δ$ is the initial objective gap and $σ^2+b_2\|x-x_1\|^2$ bounds the variance. Thus unchanged SGD attains the minimax stochastic complexity in this second-moment class. For $p>2$, predictable localization and a Hilbert-space Fuk--Nagaev inequality yield a high-probability bound separating logarithmic variance and polynomial rare-shock contributions. The localization radius is derived from the recursion: no bounded-iterate assumption, clipping, normalization, momentum, or increasing batch size is needed. We also give increasing-confidence rates, an objective-gap-growth refinement recovering root-$T$ stationarity, and stochastic $L^p$-Lipschitz examples. The broad BG-0 optimality statement is distinguished from the smaller mean-square-smooth class, in which additional oracle structure permits faster algorithms.

stat.ML↗