arXiv · 2610.00771
Localizing Transfer Between Memorization Tasks
Abstract
A central puzzle in transfer learning is why pre-training on one task can accelerate training or improve performance on another task, and what mechanisms underlie this transfer. In this work, we examine the transfer between memorization tasks of random input-output mappings. We find two surprising transfer patterns: equivalent transfer, where each additional pre-training epoch saves approximately one downstream fine-tuning epoch; and non-equivalent transfer, where pre-training on a mismatched task can be even more efficient than directly training on the downstream task itself. Through ablation experiments, we decompose and localize the transfer into two separate effects: a "trivial" magnitude-driven transfer in the last layer, and a "non-trivial" structure-driven transfer, partially attributable to the covariance of the other layers. These results advance our understanding of the underlying mechanisms of transfer learning and have the potential to lead to principled pre-training strategies.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yimiao Yu, Florentin Guth. 2026-09-30. Localizing Transfer Between Memorization Tasks. https://arxiv.org/abs/2610.00771
Cite the original work for its findings. Save a collection to share your selection of sources.