arXiv ScienceSearch

arXiv subjects

Entao Yang

Publications and source records attributed to Entao Yang.

9 recordsLinked to original sources

One Policy Is Enough: Single-Agent Reinforcement Learning Outperforms Tree Search for Chemistry Tool Learning

Chemistry questions often demand exact computation and database lookups that a language model cannot supply from its parameters, so it must reach for external tools. Tool use here is a three-part problem: select the right tool from a large pool, fill it with correctly typed arguments, and chain calls so that each consumes the outputs of the last. CheMatAgent, a previously published system, addresses this with hierarchical evolutionary MCTS: separate policy and execution models searching tool-call trees under two learned critics, one regressed partly onto GPT-assigned scores. We show that a single policy suffices. Our model interleaves reasoning, tool calls, and returns in one left-to-right generation, trained by a supervised warm-up and then outcome-level reinforcement learning against a programmatic reward read directly off the gold call chain, which leaves no learned critic and no judge in the training loop. On ChemToolBench multiple-tool comprehensive chemistry, on both backbones CheMatAgent use, we improve Tool F1 by 5.5% and Return F1 by 9.6% on Qwen-2.5-7B, and by 3.7% and 3.9% on Llama-3.1-8B, compared with their strongest search configuration, at one model invocation per question, against a search whose cost grows with the tree; we also lead answer Pass Rate on Qwen-2.5-7B.

cs.LG

Tunneling the Loss Landscape: Bypassing Memorization with Monte Carlo Parameter Swapping

Grokking is a striking phenomenon in neural network training, where a model can undergo a prolonged period of pure memorization before abrupt generalization. While previous works have attempted to interpret it through classical machine learning mechanisms like weight norm, recent research draws an analogy from statistical physics, framing grokking as a form of computational glass relaxation. This theory defines the initial memorization as a result of `fast cooling' where the training loss is reduced so quickly that a glass state is formed, followed by a `slow relaxation' towards final generalization. Although providing a unifying framework for representative grokking theories, this perspective has remained largely at the theoretical on macroscopic level without direct empirical validation on training dynamics. Here we introduce a three-component framework to directly characterize the training dynamics via parameter mobility (PM), and two representative measurements from glassy dynamics: replica correlation (RC) and fractal dimension (FD). We demonstrate that standard optimization presents clear signatures of glass dynamics and inherently traps the grokking network in a kinetic arrested memorization state with a collapsed mobility, strong history dependence, and channel-like motions. This quantitative agreement motivates us to introduce State-Aware Monte Carlo Parameter Swapping (SAM-Swap), an optimization plug-in that can accelerate generalization, inspired by swap Monte Carlo algorithm widely used in glass dynamics. Comparing SAM-Swap, weight decay, and Gaussian gradient noise, we find that accelerated generalization is consistently associated with random exploration in the parameter space, similar to diffusion in physics.

cond-mat.dis-nn

The Grokked Illusion: True Equilibrium Mitigates Catastrophic Forgetting

While neural networks are typically evaluated by their training and test performance, these metrics do not reveal how robust a learned representation is. Recent studies have shown that solutions occupying larger volumes in parameter space, as quantified by Boltzmann entropy, often exhibit superior generalizability compared to those reached by conventional optimization, a phenomenon known as the high entropy advantage. Here we ask whether this advantage persists beyond generalization. Specifically, we investigate models' robustness, the ability to retain the learned knowledge when the model is subsequently trained to acquire new information. Using grokking in modular arithmetic as a controlled setting, we design a noise injection experiment to evaluate the robustness difference between AdamW-trained transformers and high-entropy model sampled from Wang-Landau Molecular Dynamics with identical saturated performance. By forcing both models to fully remember new data with random labels, we find that AdamW-trained models suffer from catastrophic forgetting, with original task test accuracy dropping from 100% to below 75%, whereas the high-entropy models maintain approximately 95% test accuracy. We term this hidden fragility behind apparent generalization the "grokked illusion." Through singular value decomposition of the neural network weights, we discover that high-entropy neural networks possess significantly higher effective rank in attention and MLP layers both before and after noise injection, indicating richer feature representations can serve as a buffer against catastrophic forgetting. Our findings demonstrate that perfect generalization does not imply equal robustness, offering a new perspective on what makes a trained model robust to interference.

cs.LG

Is Grokking a Computational Glass Relaxation?

Understanding neural network's (NN) generalizability remains a central question in deep learning research. The special phenomenon of grokking, where NNs abruptly generalize long after the training performance reaches a near-perfect level, offers a unique window to investigate the underlying mechanisms of NNs' generalizability. Here we propose an interpretation for grokking by framing it as a computational glass relaxation: viewing NNs as a physical system where parameters are the degrees of freedom and train loss is the system energy, we find memorization process resembles a rapid cooling of liquid into non-equilibrium glassy state at low temperature and the later generalization is like a slow relaxation towards a more stable configuration. This mapping enables us to sample NNs' Boltzmann entropy (states of density) landscape as a function of training loss and test accuracy. Our experiments in transformers on arithmetic tasks suggests that there is NO entropy barrier in the memorization-to-generalization transition of grokking, challenging previous theory that defines grokking as a first-order phase transition. We identify a high-entropy advantage under grokking, an extension of prior work linking entropy to generalizability but much more significant. Inspired by grokking's far-from-equilibrium nature, we develop a toy optimizer WanD based on Wang-landau molecular dynamics, which can eliminate grokking without any constraints and find high-norm generalizing solutions. This provides strictly-defined counterexamples to theory attributing grokking solely to weight norm evolution towards the Goldilocks zone and also suggests new potential ways for optimizer design.

cs.LG

High-entropy Advantage in Neural Networks' Generalizability

One of the central challenges in modern machine learning is understanding how neural networks generalize knowledge learned from training data to unseen test data. While numerous empirical techniques have been proposed to improve generalization, a theoretical understanding of the mechanism of generalization remains elusive. Here we introduce the concept of Boltzmann entropy into neural networks by re-conceptualizing such networks as hypothetical molecular systems where weights and biases are atomic coordinates, and the loss function is the potential energy. By employing molecular simulation algorithms, we compute entropy landscapes as functions of both training loss and test accuracy (or test loss), on networks with up to 1 million parameters, across four distinct machine learning tasks: arithmetic question, real-world tabular data, image recognition, and language modeling. Our results reveal the existence of high-entropy advantage, wherein high-entropy network states generally outperform those reached via conventional training techniques like stochastic gradient descent. This entropy advantage provides a thermodynamic explanation for neural network generalizability: the generalizable states occupy a larger part of the parameter space than its non-generalizable analog at low train loss. Furthermore, we find this advantage more pronounced in narrower neural networks, indicating a need for different training optimizers tailored to different sizes of networks.

cs.LG

Machine learning-informed structuro-elastoplasticity predicts ductility of disordered solids

All solids yield under sufficiently high mechanical loads. Below yield, the mechanical responses of all disordered solids are nearly alike, but above yield every different disordered solid responds in its own way. Brittle systems can shatter without warning, like ordinary window glass, or exhibit strain localization prior to fracture, like metallic or polymeric glasses. Ductile systems, e.g. foams like shaving cream or emulsions like mayonnaise, can flow indefinitely with no strain localization. While there are empirical strategies for tuning the degree of strain localization, there is no framework that explains their effectiveness or limitations. We show that Structuro-Elastoplastic (StEP) models provide microscopic understanding of how strain localization depends on the interplay of structure, plasticity and elasticity.

cond-mat.soft

Structuro-elasto-plasticity (StEP) model for plasticity in disordered solids

Elastoplastic lattice models for the response of solids to deformation typically incorporate structure only implicitly via a local yield strain that is assigned to each site. However, the local yield strain can change in response to a nearby or even distant plastic event in the system. This interplay is key to understanding phenomena such as avalanches in which one plastic event can trigger another, leading to a cascade of events, but typically is neglected in elastoplastic models. To include the interplay one could calculate the local yield strain for a given particulate system and follow its evolution, but this is expensive and requires knowledge of particle interactions, which is often hard to extract from experiments. Instead, we introduce a structural quantity, "softness," obtained using machine learning to correlate with imminent plastic rearrangements. We show that softness also correlates with local yield strain. We incorporate softness to construct a "structuro-elasto-plasticity" model that reproduces particle simulation results quantitatively for several observable quantities, confirming that we capture the influence of the interplay of local structure, plasticity, and elasticity on material response.

cond-mat.soft

Understanding Creep Suppression Mechanism in Polymer Nanocomposites through Machine Learning

While recent efforts have shown how local structure plays an essential role in the dynamic heterogeneity of homogeneous glass-forming materials, systems containing interfaces such as thin films or composite materials remain poorly understood. It is known that interfaces perturb the molecular packing nearby, however, numerous studies show the dynamics are modified over a much larger range. Here, we examine the dynamics in polymer nanocomposites (PNCs) using a combination of simulations and experiments and quantitatively separate the role of polymer packing from other effects on the dynamics, as a function of distance from the nanoparticle surfaces. After showing good qualitative agreement between the simulations and experiments in glassy structure and creep compliance, we use a recently developed machine learning technique to decompose polymer dynamics in our simulated PNCs into structure-dependent and structure-independent processes. With this decomposition, the free energy barrier for polymer rearrangement can be described as a combination of packing-dependent and packing-independent barriers. We find both barriers are higher near nanoparticles and decrease with applied stress, quantitatively demonstrating that the slow interfacial dynamics is not solely due to polymer packing differences, but also the change of structure-dynamics relationships. Finally, we present how this decomposition can be used to accurately predict strain-time creep curves for PNCs from their static configuration, providing additional insights into the effects of polymer-nanoparticle interfaces on creep suppression in PNCs.

cond-mat.soft

The Role of Local Structure in the Enhanced Dynamics of Deformed Glasses

External stress can accelerate molecular mobility of amorphous solids by several orders of magnitude. The changes in mobility are commonly interpreted through the Eyring model, which invokes an empirical activation volume whose origin remains poorly understood. Here, we analyze constant-stress molecular dynamics simulations and propose an extension of the Eyring model with a machine-learned field, softness. Our model connects the activation volume, an empirical parameter, to a structural property (softness). We show that stress has an inhomogeneous effect on the mobility that depends on local structure, which explains the narrower distribution of relaxation time observed under stress.

cond-mat.soft