arXiv · 2609.33594
SymbolicLM: Training Language Models as Symbolic Regressors
Abstract
Large Language Models (LLMs) have shown promising capabilities in scientific reasoning, yet scientific discovery ultimately requires deriving precise laws directly from observational data, known as Symbolic Regression (SR). This poses a challenge for LLMs due to the gap between probabilistic text generation and the exact structural requirements of SR. Existing approaches rely on complex external scaffolds, which are computationally expensive and separate symbolic reasoning from the model itself. To address this limitation, we propose to directly equip LLMs with symbolic regression capabilities through dedicated numerical-symbolic and physical supervision. We introduce PhysSymbArena, a large-scale benchmark containing over 160,000 equations and 1.8B tokens of numerical-symbolic data with physical descriptions, enabling systematic training and evaluation. Based on PhysSymbArena, we develop SymbolicLM, which enhances the symbolic regression ability of LLMs through mathematical and physical supervision. During inference, we further introduce SymbolicSGA, a refinement framework that leverages quantitative feedback to iteratively improve generated equations. Experiments on multiple symbolic regression benchmarks show that SymbolicLM substantially improves structural recovery while maintaining competitive numerical fitting performance. These results demonstrate that symbolic regression can be explicitly learned as an intrinsic capability of LLMs.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jun Yao, Yingfan Hua, Ruikun Li, Shixiang Tang, Bin Liu, Wanli Ouyang, Yan Lu. 2026-09-27. SymbolicLM: Training Language Models as Symbolic Regressors. https://arxiv.org/abs/2609.33594
Cite the original work for its findings. Save a collection to share your selection of sources.