arXiv ScienceSearch

arXiv · 2609.19419

Agentic AI for Density-Functional Development: Revisiting r2SCAN

Abstract

We demonstrate physics-constrained agentic development of a meta-generalized gradient approximation (meta-GGA) functional using a large language model (LLM) to assist the search and optimization of a band-gap-oriented revision of r2SCAN. No real bonded systems were fitted, preserving r2SCAN's nonempirical philosophy. Across the finalist set, band gaps and several molecular subsets improve relative to r2SCAN; the top finalist, r2SCAN+, reduces the band-gap MAE on a benchmark comprising 24 solids from 1.26 to 0.96 eV and the aggregate MAE on 329 molecular properties from 4.79 to 4.42 kcal/mol. We first curated 78 exchange and 90 correlation candidate correction terms from r2SCAN's dimensionless ingredients, spanning polynomial terms through third degree, exponentials, exponentially damped products, and ratios. Allowing each candidate to combine one to three correction terms from the exchange catalog, the correlation catalog, or both yields about 8 x 10^5 distinct forms, making exhaustive high-throughput screening impractical. We defined the search criteria for the LLM agent using r2SCAN's exact constraints, physical norms, and the targeted iso-orbital derivative response. The agent then combined these criteria with its pretrained knowledge and accumulated search feedback to propose and refine sparse forms, prioritizing terms tied to the iso-orbital response; a second LLM critic screened proposals before deterministic verification. Compared with uniform random search, the workflow learned from prior evaluations, incurred far fewer downstream rejections (0.6% versus 24.6%), and located stronger high-response candidates: 51 agentic candidates exceeded the best random-search response of 1.263, with the overall best reaching 1.331. These results show that agentic search can support density-functional development when flexible hypothesis generation is coupled to automated physical verification.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Santosh Adhikari, Kelsey A. Parker, Etinosa Osaro, Swagata Roy, Dario Rocca. 2026-09-16. Agentic AI for Density-Functional Development: Revisiting r2SCAN. https://arxiv.org/abs/2609.19419

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Global approximations to the error function of real argument for vectorized computation

The error function of real argument can be approximated to a given uniform relative accuracy by a single closed-form expression for the whole variable range either in terms of addition, multiplication, division, and square root operations only, or also using the exponential function. The coefficients have been tabulated for up to 128-bit precision. Tests of a computer code implementation using the standard single- and double-precision floating-point arithmetic show good performance and vectorizability. Approximations to the complementary error function have also been found, including those where only additions, multiplications, and one division are needed to reach a uniform absolute accuracy.

physics.chem-ph

From Heuristics to Machine Learning: The Performance Ceiling for Single-Ion Magnets and Its Electronic Origin

Machine learning (ML) is expected to speed up the discovery of single-ion magnets (SIMs), but does the structural information available before synthesis allow such predictions? For 1215 lanthanide complexes from the SIMDAVIS 1.2.1 database we compared three increasing levels of structural description: tabular features of the coordination site, continuous symmetry measures of the coordination polyhedron, and the complete 3D arrangement of atoms. All three converge to an accuracy near 76%, only slightly above the 71% of the single rule "predict SIM for Dy3+". To explain the failures, we combined multireference ab initio calculations with an inspection of the structures behind the high-confidence errors. The SIMs missed by the geometric models are field-induced relaxers whose ground Kramers doublets are prone to tunnelling, a property invisible to geometric descriptors. Many false positives contain several lanthanide centers or radicals, so their relaxation is collective and outside the single-ion picture. The electronic-structure and connectivity information needed to identify SIMs is therefore not accessible to geometric methods alone. Geometric models remain useful: restricting the screening to compounds with high prediction confidence raises the accuracy to 88% while retaining 48% of the dataset. Building on the analysis of the failures, we propose a strategy that combines simple filters for nuclearity and for radicals with ligand-field descriptors from ab initio calculations.

physics.chem-ph

Composition-Dependent Self-Diffusion Coefficients in Liquid Mixtures from Hybrid Machine Learning

Self-diffusion coefficients are key descriptors of molecular mobility, yet experimental data remain scarce, highlighting the need for reliable prediction methods. In previous work, we introduced the hybrid Enhanced Stokes-Einstein (ESE) model, which advanced the state of the art in the physically consistent prediction of self-diffusion coefficients of solutes at infinite dilution in pure solvents by integrating the Stokes-Einstein equation with machine learning (ML). Here, we extend this approach to concentration-dependent self-diffusion coefficients and multicomponent solvents with HADES. This hybrid architecture leverages a deep-set neural network to connect pure-component and mixture prediction within a single framework. HADES predicts self-diffusion coefficients in liquid mixtures with any number of components at any composition and temperature. The only required inputs are SMILES-encoded molecular structures of the components and the pure-component viscosities, making the method broadly applicable. Trained and evaluated on a comprehensive dataset of 2526 data points for 600 systems, HADES significantly outperforms benchmark prediction methods. The trained model and its source code are fully disclosed, and the application is available via an interactive website https://ml-prop.mv.rptu.de/.

physics.chem-ph