arXiv · 2608.06766
Hidden Gauge Controls Feature Specialization in ReLU Networks
Abstract
The success of deep learning depends on learning useful representations, yet predicting how training organizes these representations across neurons remains difficult. In this work, we show that changing the scale of initial weights can determine which neurons learn a feature without altering any neuron's initial contribution. We construct ReLU networks with identical initial features and predictions that reach the same final predictions with different roles for their neurons. In one, all neurons share the learned feature. In the other, one neuron acquires it while every other neuron's contribution vanishes. The only change is the relative scale of each neuron's input and output weights. Our analysis explains how an initial learning advantage persists through convergence: as one neuron learns the target, it reduces the error driving the others and limits their subsequent adaptation. We prove this outcome in a nonlinear model under gradient flow and small-step gradient descent, and quantify how scale changes the speed and path of feature learning. Experiments verify the predicted dynamics and show that scale also changes feature assignment when two features compete. These results reveal how initialization can control the organization of a learned representation without changing what the network initially represents.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Tongxi Wang. 2026-09-13. Hidden Gauge Controls Feature Specialization in ReLU Networks. https://arxiv.org/abs/2608.06766
Cite the original work for its findings. Save a collection to share your selection of sources.