arXiv ScienceSearch

arXiv subjects

Keith T. Butler

Publications and source records attributed to Keith T. Butler.

At least 19 recordsLinked to original sources

Revealing low-energy surfaces of multinary compounds by controlling surface coordination environments

When modeling surfaces of multinary compounds, conventional cleavage planes often cut through strongly bonded polyhedra, resulting in unphysical surface energies. Here, we introduce SALAMI (Symmetric Atomic Layers for Arbitrary Multinary Interfaces), a Python package that generates symmetric, charge-neutral, dipole-free, and low-energy slab models for multinary compounds. SALAMI performs combinatorial searches to selectively remove surface atoms and generate corrugated terminations that preserve optimal coordination environments. We applied this workflow to all symmetrically inequivalent crystallographic orientations with Miller indices up to 2 for two prototypical structures: the solid-state electrolyte Li3PS4 and the transparent conducting oxide ZnSb2O6. Density functional theory calculations reveal that Li3PS4 must preserve all PS4 units to achieve the minimum surface energy. For ZnSb2O6, low-energy surfaces are achieved by partial undercoordination of surface Sb atoms to SbO5 or SbO4 from the bulk SbO6, depending on the surface orientations. Compared to surface models generated with unconstrained coordination, applying constraints to achieve optimal local coordination environments significantly lowers surface energies, shrinking the volume of the predicted Wulff shape by approximately 20%. Our results demonstrate that meticulous control of local coordination environments is necessary for accurately predicting the surface energetics of multinary compounds.

cond-mat.mtrl-sci

Building atomistic models of heterointerfaces with optimal transport

Heterogeneous interfaces underpin technologies from microelectronics to energy conversion and storage, but their configurational complexity precludes exhaustive first-principles screening of interface registries. Although data-driven approaches can alleviate this burden, they remain limited by sparse interface datasets. Here, we introduce an energy-independent workflow that represents coherent interfaces as attributed graphs, quantifies their similarity to parent bulk environments using the fused Gromov-Wasserstein (FGW) distance, and couples this metric with Bayesian optimization over the in-plane registry space. We assess the approach for KI/NaCl, GaP/GaAs and GaN/$\mathrm{Al_{2}O_{3}}$ interfaces spanning ionic, covalent and mixed-bonding regimes, using hierarchical validation with MACE and density functional theory (DFT). Comparison with single-point energy landscapes shows that the FGW distance captures registry-dependent periodicity, while interfaces exhibit deviations between structural and energetic extrema, reflecting additional chemistry-specific contributions. Furthermore, FGW distances show an overall association with relaxed energies. Under limited screening budgets, FGW-guided registry selection consistently outperforms random search and is more robust across interface systems than selection guided by pretrained MACE energies. The workflow converts the qualitative notion of bulk-like continuity into a quantitative prescreening criterion, enabling efficient registry exploration and providing physically informed candidate structures for materials discovery workflows.

cond-mat.mtrl-sci

Six Open Questions in Machine-Learned Interatomic Potential Foundation Models

Machine-learned interatomic potentials (MLIPs) have had a profound impact on molecular modelling in recent years, promising to resolve the long-standing tension between the scale and accuracy of simulations. There has been a proliferation of new models and designs, and recently the paradigm of ``foundational'' MLIPs has become prevalent. Broadly speaking, foundation models are trained on large diverse datasets and promise to work well for new systems with minimal updates required. However, in such a new and fast moving field, there are many unanswered questions. In this article, we set out to articulate and explore what we see as the most important among these questions. We start by developing a working definition for foundational MLIPs and use this definition to frame the subsequent open questions. Despite the rapid progress in the field of MLIP models, we believe that these are fundamental questions which will continue to define cutting edge research in MLIPs in the years to come.

cond-mat.mtrl-sci

Conditional Generative Models Enable Targeted Exploration of MAX Phase Design Space

MAX phases (M$_{n+1}$AX$_n$), precursors to MXenes, span a vast compositional space, motivating efficient computational screening for synthesisable candidates. We employ CrystaLLM$-\pi$, a large language model fine-tuned on 6,179 double transition-metal MAX phases, and demonstrate its ability to generate out-of-sample structures consistent with known experimental trends. Using a conditioning vector with two dimensions (a statistically derived MXene derivative count and a surrogate for A-site binding energy), the model was able to target MXene-favourable regions of phase space for generation. Specific condition vectors double novel stable structure generation rates versus unconditioned baselines. Of ten compositionally novel candidates, five exhibit DFT-validated stability ($E_{hull} < 0.050$ eV/atom). This work showcases the potential for autoregressive generative models to explore targeted materials' spaces, offering a scalable framework for accelerated discovery in compositionally complex systems.

cond-mat.mtrl-sci

Materials Informatics Across the Length Scales

Materials informatics is increasingly used to support modelling, analysis and design across the length scales of materials science, from atomistic simulations to microstructural characterisation and continuum descriptions. Despite rapid progress, the reliability and transferability of these approaches vary strongly with scale. Here we survey data-driven methods at the nanoscale, mesoscale, and micro-to-continuum levels, highlighting established capabilities as well as unresolved challenges. Machine-learning interatomic potentials, mesoscale surrogate and operator-learning models, and learning-based analysis of experimental microstructures are discussed, with emphasis on data quality, uncertainty, interpretability, and cross-scale consistency. We further examine the role of data standards, ontologies, and emerging tools, such as autonomous laboratories, where they directly affect multiscale workflows. This perspective clarifies what can be considered reliable today and identifies key obstacles to the broader integration of materials informatics across scales.

cond-mat.mtrl-sci

A Shift-Invariant Deep Learning Framework for Automated Analysis of XPS Spectra

X-ray Photoelectron Spectroscopy (XPS) is a crucial technique for material surface analysis, yet interpreting its spectra is often challenging for both human analysts and automated methods due to the prevalence of variable spectral shifts and overlapping peaks. This project introduces a machine learning solution using a Spatial Transformer Network (STN), a type of neural network that implicitly learns to align spectra. An STN model was designed to classify the chemical environments present in an input spectrum, using functional groups as a proxy. The model was trained and tested on a large synthetic dataset of 100,000 spectra, created by linearly combining real experimental data from a library of 104 polymers. \cite{RN22} To simulate experimental variability, random uniform shifts and broadening were applied to the data. The STN was found to effectively correct for random electrostatic shifts (up to 3.0 eV) and achieved relatively high accuracy ($\sim$ 82\%) in identifying functional groups, despite utilizing a much simpler architecture than previous work. These findings demonstrate that neural networks can effectively learn the underlying relationships between spectral features and chemical composition when they are able to intrinsically account for variable shifts. This work advances the development of more reliable automated XPS analysis, offering potential as an assistive tool for researchers and as a core component in future autonomous systems like self-driving laboratories.

cond-mat.mtrl-sci

Discovering new photovoltaics using optimal transport theory

Searching by chemical and structural analogy is one of the most commonly used and successful approaches to materials discovery. However, formulating this task for algorithmic implementation raises the question of how we define similar materials. Methods have been proposed for searching materials space using vectors based on chemical composition and functional fragments in the material. Descriptors for structural similarity have also been proposed. However, the question of how to incorporate and balance structural and compositional similarity measures in a single metric remains open. Here, we adapt methods developed for calculating distances between undirected graphs and apply them to crystalline materials similarity. The Fused Gromov-Wasserstein (FGW) metric uses optimal transport theory to map between two graphs considering a balance of the graph structure and the information present at the nodes of the graph (atoms in crystals). We apply the method to exploring new photovoltaic materials. We demonstrate that FGW is competitive with embeddings from an equivariant graph neural network, trained on $> 10^6$ materials, despite minimal training. We then apply FGW to a discovery campaign to identify materials from the Materials Project database that have not previously been explored as photovoltaics, but have similarities to known high-efficiency materials. After validating predictions with hybrid density functional theory, we identify seven previously unexplored high-efficiency photovoltaic absorber candidates, including Cs$_5$Sb$_8$, which is found to have a predicted SLME of $> 30\%$ and to be thermodynamically stable. The FGW approach demonstrates the power of strong inductive biases for developing metrics for materials exploration with minimal training data.

cond-mat.mtrl-sci

Discovery and recovery of crystalline materials with property-conditioned transformers

Generative models have recently shown great promise for accelerating the design and discovery of new functional materials. Conditional generation enhances this capacity by allowing inverse design, where specific desired properties can be requested during the generation process. However, conditioning of transformer-based approaches, in particular, is constrained by discrete tokenisation schemes and the risk of catastrophic forgetting during fine-tuning. This work introduces CrystaLLM-{\pi} (property injection), a conditional autoregressive framework that integrates continuous property representations directly into the transformer's attention mechanism. Two architectures, Property-Key-Value (PKV) Prefix attention and PKV Residual attention, are presented. These methods bypass inefficient sequence-level tokenisation and preserve foundational knowledge from unsupervised pre-training on Crystallographic Information Files (CIFs) as textual input. We establish the efficacy of these mechanisms through systematic robustness studies and evaluate the framework's versatility across two distinct tasks. First, for structure recovery, the model processes high-dimensional, heterogeneous X-ray diffraction patterns, achieving structural accuracy competitive with specialised models and demonstrating applications to experimental structure recovery and polymorph differentiation. Second, for materials discovery, the model is fine-tuned on a specialised photovoltaic dataset to generate novel, stable candidates validated by Density Functional Theory (DFT). It implicitly learns to target optimal band gap regions for high photovoltaic efficiency, demonstrating a capability to map complex structure-property relationships. CrystaLLM-{\pi} provides a unified, flexible, and computationally efficient framework for inverse materials design.

cond-mat.mtrl-sci

General Learning of the Electric Response of Inorganic Materials

We introduce \texttt{MACE-Field}, a field-aware, $O(3)$-equivariant interatomic potential that learns a single electric enthalpy functional $\mathcal F(\{\mathbf R\},\mathbf E)$ and obtains $\mathbf P$, $Z^*$, and $\boldsymbol\alpha$ by exact differentiation. A uniform field couples to latent equivariant features inside the \texttt{MACE} backbone, while the scalar energy readout preserves Maxwell reciprocity, the acoustic sum rule, and crystal tensor symmetries by construction. Because this coupling is a plug-in on top of standard \texttt{MACE}, existing energy/force foundation models can be upgraded to become field-aware. Benchmarked against semilocal DFT/DFPT reference data, a directly trained cross-chemistry ferroelectric model reproduces the same-branch Berry-phase and spontaneous polarisations across diverse inorganic crystals. Starting from the multihead foundation model \texttt{mace-mp-mh-0} and its OMAT-PBE head, joint fine-tuning on dielectric, ferroelectric, and replay data yields \texttt{MACE-Field-MH-0} foundation models, which predict $Z^*$, $\boldsymbol\alpha$, derived dielectric constants, and cross-chemistry polarisation trends with fidelity that captures branch-resolved polarisation and spontaneous-polarisation, while retaining strong force-field accuracy. Further, single-material \texttt{MACE-Field} models and \texttt{MACE-Field-MH-0} reproduce \ce{BaTiO3} hysteresis loops and $\alpha$-quartz infrared, Raman, and dielectric spectra from finite-field molecular dynamics, comparable to DFPT. These results show that a simple, physics-informed field coupling can endow atomistic foundation models with transferable dielectric and ferroelectric response, while targeted single-material training remains advantageous for the most quantitative spectroscopic predictions.

cond-mat.mtrl-sci

Leveraging transfer learning for accurate estimation of ionic migration barriers in solids

Ionic mobility determines the rate performance of several applications, such as batteries, fuel cells, and electrochemical sensors and is exponentially dependent on the migration barrier ($E_m$), a difficult to measure/calculate quantity. Previous approaches to identify materials with high ionic mobility have relied on imprecise descriptors given the lack of generalizable models to predict $E_m$. Here, we present a graph neural network based architecture that leverages principles of transfer learning to efficiently and accurately predict $E_m$ across a diverse set of materials. We use a model pre-trained simultaneously on seven distinct bulk properties (labeled MPT), modify the MPT model to classify different migration pathways in a structure, and fine-tune (FT) on a manually-curated literature-derived dataset of 619 $E_m$ data points calculated with density functional theory. Importantly, our best-performing FT model (labeled MODEL-3) demonstrates substantial improvements in prediction accuracy compared to classical machine learning methods, graph models trained from scratch, and a universal machine learned interatomic potential, with a R$^2$ score of 0.703 and a mean absolute error of 0.261 eV on the test set. Notably, MODEL-3 is able to distinguish different migration pathways within a structure and also demonstrates excellent ability to generalize across intercalant compositions and chemistries. As a classifier, MODEL-3 exhibits 80\% accuracy and 82.8\% precision in identifying materials that are `good' ionic conductors (i.e., structures with $E_m <$0.65~eV). Thus, our work demonstrates the effective use of FT strategies and architectural modifications necessary for making swift and accurate $E_m$ predictions, which will be useful for materials discovery in batteries and for predicting other data-scarce material properties.

cond-mat.mtrl-sci

A literature-derived dataset of migration barriers for quantifying ionic transport in battery materials

The rate performance of any electrode or solid electrolyte material used in a battery is critically dependent on the migration barrier ($E_m$) governing the motion of the intercalant ion, which is a difficult-to-estimate quantity both experimentally and computationally. The foundation for constructing and validating accurate machine learning (ML) models that are capable of predicting $E_m$, and hence accelerating the discovery of novel electrodes and solid electrolytes, lies in the availability of high-quality dataset(s) containing $E_m$. Addressing this critical requirement, we present a comprehensive dataset comprising 619 distinct literature-reported $E_m$ values calculated using density functional theory based nudged elastic band computations, across 443 compositions and 27 structural groups consisting of various compounds that have been explored as electrodes or solid electrolytes in batteries. Our dataset includes compositions that correspond to fully charged and/or discharged states of electrode materials, with intermediate compositions incorporated in select instances. Crucially, for each compound, our dataset provides structural information, including the initial and final positions of the migrating ion, along with its corresponding $E_m$ in easy-to-use .xlsx and JSON formats. We envision our dataset to be a highly useful resource for the scientific community, facilitating the development of advanced ML models that can predict $E_m$ precisely and accelerate materials discovery.

cond-mat.mtrl-sci

Learning disentangled latent representations facilitates discovery and design of functional materials

The discovery of new materials is often constrained by the need for large labelled datasets or expensive simulations. In this study, we explore the use of Disentangling Autoencoders (DAEs) to learn compact and interpretable representations of spectral data in an entirely unsupervised manner. We demonstrate that the DAE captures physically meaningful features in optical absorption spectra, relevant to photovoltaic (PV) performance, including a latent dimension strongly correlated with the Spectroscopic Limited Maximum Efficiency (SLME)--despite being trained without access to SLME labels. This feature corresponds to a well-known spectral signature: the transition from direct to indirect optical band gaps. Compared to Principal Component Analysis (PCA) and a beta-Variational Autoencoder (beta-VAE), the DAE achieves superior reconstruction fidelity, improved correlation with efficiency metrics, and more compact encoding of relevant features. We further show that the DAE latent space enables more efficient discovery of high-performing PV materials, identifying top candidates using fewer evaluations than both VAE-guided and random search. These results highlight the potential of DAEs as a powerful tool for unsupervised structure-property learning and suggest broad applicability to other areas of materials discovery where labeled data is limited but rich structure is present in raw signals.

cond-mat.mtrl-sci

The carbon cost of materials discovery: Can machine learning really accelerate the discovery of new photovoltaics?

Computational screening has become a powerful complement to experimental efforts in the discovery of high-performance photovoltaic (PV) materials. Most workflows rely on density functional theory (DFT) to estimate electronic and optical properties relevant to solar energy conversion. Although more efficient than laboratory-based methods, DFT calculations still entail substantial computational and environmental costs. Machine learning (ML) models have recently gained attention as surrogates for DFT, offering drastic reductions in resource use with competitive predictive performance. In this study, we reproduce a canonical DFT-based workflow to estimate the maximum efficiency limit and progressively replace its components with ML surrogates. By quantifying the CO$_2$ emissions associated with each computational strategy, we evaluate the trade-offs between predictive efficacy and environmental cost. Our results reveal multiple hybrid ML/DFT strategies that optimize different points along the accuracy--emissions front. We find that direct prediction of scalar quantities, such as maximum efficiency, is significantly more tractable than using predicted absorption spectra as an intermediate step. Interestingly, ML models trained on DFT data can outperform DFT workflows using alternative exchange--correlation functionals in screening applications, highlighting the consistency and utility of data-driven approaches. We also assess strategies to improve ML-driven screening through expanded datasets and improved model architectures tailored to PV-relevant features. This work provides a quantitative framework for building low-emission, high-throughput discovery pipelines.

cond-mat.mtrl-sci

Learning Radical Excited States from Sparse Data

Emissive organic radicals are currently of great interest for their potential use in the next generation of highly efficient organic light emitting diode (OLED) devices and as molecular qubits. However, simulating their optoelectronic properties is challenging, largely due to spin-contamination and the multireference character of their excited states. Here we present a data-driven approach where, for the first time, the excited electronic states of organic radicals are learned directly from experimental excited state data, using a much smaller amount of data than typically required by Machine Learning. We adopt ExROPPP, a fast and spin-pure semiempirical method for calculation of the excited states of radicals, as a surrogate physical model for which we learn the optimal set of parameters. To achieve this we compile the largest known database of organic radical geometries and their UV-vis data, which we use to train our model. Our trained model gives Root Mean Square (RMS) and mean absolute errors for excited state energies of 0.24 and 0.16 eV respectively, improving hugely over ExROPPP with literature parameters. Four new organic radicals are synthesised and we test the model on their spectra, finding even lower errors and similar correlation as for the testing set. This model paves the way for the high throughput discovery of next generation radical-based optoelectronics.

physics.chem-ph

Optimal pre-train/fine-tune strategies for accurate material property predictions

Overcoming the challenge of limited data availability within materials science is crucial for the broad-based applicability of machine learning within materials science. One pathway to overcome this limited data availability is to use the framework of transfer learning (TL), where a pre-trained (PT) machine learning model (on a larger dataset) can be fine-tuned (FT) on a target (typically smaller) dataset. Our study systematically explores the effectiveness of various PT/FT strategies to learn and predict material properties with limited data. Specifically, we leverage graph neural networks (GNNs) to PT/FT on seven diverse curated materials datasets, encompassing sizes ranging from 941 to 132,752 datapoints. We consider datasets that cover a spectrum of material properties, ranging from band gaps (electronic) to formation energies (thermodynamic) and shear moduli (mechanical). We study the influence of PT and FT dataset sizes, strategies that can be employed for FT, and other hyperparameters on pair-wise TL among the datasets considered. We find our pair-wise PT-FT models to consistently outperform models trained from scratch on the target datasets. Importantly, we develop a GNN framework that is simultaneously PT on multiple properties (MPT), enabling the construction of generalized GNN models. Our MPT models outperform pair-wise PT-FT models on several datasets considered, and more significantly, on a 2D material band gap dataset that is completely out-of-distribution from the PT datasets. Finally, we expect our PT/FT and MPT frameworks to be generalizable to other GNNs and materials properties, which can accelerate materials design and discovery for various applications.

cond-mat.mtrl-sci

Effects of Grain Boundaries and Surfaces on Electronic and Mechanical Properties of Solid Electrolytes

Extended defects, including exposed surfaces and grain boundaries, are critical to the properties of polycrystalline solid electrolytes in all-solid-state batteries (ASSBs). These defects can significantly alter the mechanical and electronic properties of solid electrolytes, with direct manifestations on the performance of ASSBs. Here, by building a library of 590 surfaces and grain boundaries of 11 relevant solid electrolytes $-$including halides, oxides, and sulfides$-$ their electronic, mechanical, and thermodynamic characteristics are linked to the functional properties of polycrystalline solid electrolytes. It is found that the energy required to mechanically ``separate'' grain boundaries can be significantly lower than in the bulk region of materials, which can trigger preferential cracking of solid electrolyte particles in the grain boundary regions. The brittleness of ceramic solid electrolytes, inferred from the predicted low fracture toughnesses at the grain boundaries, contributes to their cracking under local pressure imparted by Lithium or Sodium penetration in the grain boundaries. Extended defects of solid electrolytes introduce new electronic ``interfacial'' states within bandgaps of solid electrolytes. These interfacial states alter and possibly increase locally the availability of free electrons and holes in solid electrolytes. Factoring effects arising from extended defects appear crucial to explain electrochemical and $-$mechanical observations in ASSBs.

cond-mat.mtrl-sci

Crystal Structure Generation with Autoregressive Large Language Modeling

The generation of plausible crystal structures is often the first step in predicting the structure and properties of a material from its chemical composition. Quickly generating and predicting inorganic crystal structures is important for the discovery of new materials, which can target applications such as energy or electronic devices. However, most current methods for crystal structure prediction are computationally expensive, slowing the pace of innovation. Seeding structure prediction algorithms with quality generated candidates can overcome a major bottleneck. Here, we introduce CrystaLLM, a methodology for the versatile generation of crystal structures, based on the autoregressive large language modeling (LLM) of the Crystallographic Information File (CIF) format. Trained on millions of CIF files, CrystaLLM focuses on modeling crystal structures through text. CrystaLLM can produce plausible crystal structures for a wide range of inorganic compounds unseen in training, as demonstrated by ab initio simulations. The integration with predictors of formation energy permits the use of a Monte Carlo Tree Search algorithm to improve the generation of meaningful structures. Our approach challenges conventional representations of crystals, and demonstrates the potential of LLMs for learning effective 'world models' of crystal chemistry, which will lead to accelerated discovery and innovation in materials science.

cond-mat.mtrl-sci

Element similarity in high-dimensional materials representations

The traditional display of elements in the periodic table is convenient for the study of chemistry and physics. However, the atomic number alone is insufficient for training statistical machine learning models to describe and extract composition-structure-property relationships. Here, we assess the similarity and correlations contained within high-dimensional local and distributed representations of the chemical elements, as implemented in an open-source Python package ElementEmbeddings. These include element vectors of up to 200 dimensions derived from known physical properties, crystal structure analysis, natural language processing, and deep learning models. A range of distance measures are compared and a clustering of elements into familiar groups is found using dimensionality reduction techniques. The cosine similarity is used to assess the utility of these metrics for crystal structure prediction, showing that they can outperform the traditional radius ratio rules for the structural classification of AB binary solids.

cond-mat.mtrl-sci