arXiv Science⌕ Search

arXiv · 2610.08503

Replica Fragmentation and Glassy Dynamics in Parity Learning

Abstract

We study how independently trained Transformer neural networks reconstruct a binary string from its local domain walls. Runs sharing the data and training protocol can realize different functions. We treat them as replicas and measure truth alignment $m$, prediction confidence $q_{\mathrm{self}}$, and cross-replica agreement $q_{\mathrm{cross}}$. Confident disagreement defines the finite-size replica fragmentation that we call glass-like. With small training sets, replicas predict all training examples correctly but remain confident in incorrect predictions for unseen inputs, a regime we call memorization. With larger sets, runs can generalize and then retreat. Retreat occurs when outputs start to deviate from truth while confidence remains high. The frontier between learned and unlearned outputs recedes toward shorter strings. Many later recover as the frontier advances again. The self--cross gap $q_{\mathrm{self}}-q_{\mathrm{cross}}$ clearly distinguishes the three learning regimes of memorization, retreat, and recovery. With overall and position-dependent truth alignment subtracted, their residual correlations also differ: memorizing replicas have weakly and uniformly correlated residuals, while retreat has the largest fraction of replica pairs whose residuals are anti-correlated. That is, on the inputs where one replica of such a pair does better than its average, the other tends to do worse. These finite-size observations distinguish persistent memorization from ongoing retreat--recovery dynamics. The thermodynamic and long-time limits, and extensions to other learning tasks, remain open.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Han Ma. 2026-10-06. Replica Fragmentation and Glassy Dynamics in Parity Learning. https://arxiv.org/abs/2610.08503

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Optimizing optimal transport: Role of final distributions in finite-time thermodynamics

Performing thermodynamic tasks within finite time while minimizing thermodynamic costs is a central challenge in stochastic thermodynamics. Here, we develop a unified framework for optimizing the thermodynamic cost of performing various tasks in finite time for overdamped Langevin systems. Conventional optimization of thermodynamic cost based on optimal transport theory leaves room for varying the final distributions according to the intended task, enabling further optimization. Taking advantage of this freedom, we use Lagrange multipliers to derive the optimal final distribution that minimizes the thermodynamic cost. Our framework applies to a wide range of thermodynamic tasks, including particle transport, thermal squeezing, and information processing such as information erasure, measurement, and feedback. Our results are expected to provide design principles for information-processing devices and thermodynamic machines that operate at high speed with low energetic costs.

cond-mat.stat-mech↗

Non-Markovian escape under stochastic resetting

Stochastic resetting is a powerful strategy known to optimize target-search processes at microscopic scales. While its effects on Markovian systems are well understood, its influence on memory-driven systems, such as in viscoelastic baths, has not been adequately investigated. In this work, we study the first-passage properties of escape for a harmonically trapped particle in a non-Markovian environment under stochastic resetting. We employ a complete renewal approach and find that the characteristic non-exponential heavy tail of the first passage time (FPT) distribution becomes exponential when resetting is introduced. We further find that optimal resetting is achievable at a lower reset rate when the dynamics are weakly correlated; however, for stronger correlations, the process needs to be reset more frequently. Therefore, resetting in memory-driven dynamics can be used as an effective control strategy to initiate faster escape, thereby regulating efficient transport mechanisms in complex chemical and biomolecular environments that follow non-Markovian dynamics.

cond-mat.stat-mech↗

Violation of the method of images in non-Markovian processes and its connection to stochastic heat

This article discusses a failure of the widely used method of images to describe the time evolution of probability distributions in diffusive processes with memory. A walker in one-dimensional space draws a dead or a surviving path depending on whether it has touched a target during stochastic evolution. For the dead walker, we define its conjugate twin paired by paths spatially reflected at the first passage time. The probability distribution of the reflected dead path coincides with that of the free image walker in a physical domain, but not generally with that of the original dead walker for non-Markovian processes, which violates the method of images. For systems reducible to the generalized Langevin equation with the fluctuation-dissipation relation, we propose an energetic interpretation in terms of a path-memory force, where the path-probability ratio of the dead to the reflected path obeys an analogous relation to the fluctuation theorem associated with heat from the reflected to the original dead walker. This framework provides a quantitative basis as well as an intuitive picture of how and why the method of images breaks down for non-Markovian processes.

cond-mat.stat-mech↗