arXiv ScienceSearch

arXiv · 2609.02929

Scaling Neural Network Quantum States for Ab Initio Quantum Chemistry

Abstract

Neural-network quantum states (NNQSs) can represent many-electron wave functions without explicitly enumerating the determinant space, but their accuracy depends jointly on model size and variational-optimization effort. Here we characterize this dependence for a physics-conditioned autoregressive NNQS trained separately on two six-molecule source benchmarks. Across eight model sizes and five optimization milestones, we find that model size and optimization steps jointly shape the energy error. The capacity advantage of larger models becomes more apparent with sufficient optimization, while the returns from additional optimization vary with model size. We capture this coupling using an interaction scaling law and quantify the cumulative compute of each evaluated configuration. The resulting error-compute Pareto frontiers provide a practical decision rule for jointly selecting model size and optimization steps under a given compute budget within the evaluated range. Furthermore, we find that this beneficial scaling trend persists during fine-tuning on held-out N$_2$. Pretrained models show decreasing error with increasing model size, with a steeper reduction following pretraining on the Hard benchmark. Together, these results place autoregressive neural quantum states within the broader landscape of empirical neural scaling and open a quantitative route toward the systematic scaling of neural quantum solvers for ab initio quantum chemistry.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Chenxi Yu, Hanlin Kong, Jianan Wei, Lizhong Fu, Honghui Shang, Wenguan Wang, Jinlong Yang. 2026-08-27. Scaling Neural Network Quantum States for Ab Initio Quantum Chemistry. https://arxiv.org/abs/2609.02929

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Dynamic Response Functions for Cavity Quantum Electrodynamics Hartree-Fock Theory

Frequency-dependent linear and quadratic response functions are implemented for a cavity quantum electrodynamics (QED) generalization of Hartree-Fock (HF) theory. Dynamic electric dipole polarizability, optical rectification hyperpolarizability, and second harmonic generation hyperpolarizability tensors are evaluated for molecules strongly coupled to a single-mode optical cavity, and the results are benchmarked against the same tensors obtained from real-time time-dependent QED-HF simulations. Substantial cavity-induced changes to the frequency-dependent properties are observed for certain perturbing frequencies at large electron-photon coupling strengths. We also find special perturbing frequencies at which there is no cavity effect, regardless of the coupling strength.

physics.chem-ph

Perturbatively Corrected Linear Response Selected Configuration Interaction

Selected configuration interaction (SCI) methods have emerged as powerful, lower-cost alternatives to full configuration interaction (FCI) for ground- and excited-state energies. Still, calculating molecular response properties with SCI remains a significant challenge. In this work, we introduce perturbative corrections to the linear response selected configuration interaction (LR-SCI) framework, using an order-by-order Epstein-Nesbet perturbation expansion through second order. We demonstrate that in this theoretical framework, the finite-order perturbative treatment preserves the pole structure of the parent variational LR-SCI theory, which means that although the method can be useful for static properties, it is not suitable for frequency-dependent molecular response properties. Numerical benchmarks targeting the static polarizabilities of water, ethene, boron hydride, and hydrogen chloride demonstrate systematic convergence toward the FCI limit for both ground and excited electronic states. While first-order corrections yield marginal improvements, the inclusion of second-order corrections substantially enhances accuracy over underlying variational treatments and diminishes oscillatory convergence behavior present in the parent variational LR-SCI method. Combined with extrapolation techniques, LR-SCI-PT achieves excellent agreement with high-level coupled-cluster references, establishing a powerful route toward near-FCI quality molecular properties for systems otherwise inaccessible to exact FCI treatments.

physics.chem-ph

Nonclassical condensation pathways revealed by the multivariable theory of nucleation

We extend classical nucleation theory (CNT) by explicitly incorporating the multidimensional nature of nucleation and the coupled roles of kinetics and thermodynamics. Specifically, we treat the cluster density as an independent variable, within both sharp-interface and diffuse-interface descriptions. The kinetics are governed by dynamical density functional theory. Applied to liquid condensation in the Lennard-Jones system, our two-variable (size--density) and three-variable (size--interface width--density) models reveal nonclassical nucleation mechanism. At low supersaturation, both models recover the classical picture, in which clusters nucleate and grow at the equilibrium liquid density. As supersaturation increases, a nonclassical behavior emerges: the critical cluster density decreases, and the nucleation pathway involves concomitant evolution in cluster size, density, and, within the diffuse-interface description, interfacial width. Our model with diffuse interface reveals a rapid increase in interfacial diffuseness at high supersaturation. Near the spinodal limit, both models predict the critical cluster with diverging sizes, densities approaching that of the metastable initial phase, and vanishing work of formation, which provides a smooth connection between nucleation and spinodal decomposition. Comparison with molecular dynamics simulations demonstrates that both models substantially outperform CNT. However, the weak non-monotonic dependence of the critical cluster density observed at very low supersaturation is captured only by diffuse-interface models. Overall, our findings indicate that CNT should be applied only in the low-supersaturation regime, and our work provides a robust foundation for its refinement beyond this limit.

physics.chem-ph