arXiv Science⌕ Search

arXiv subjects

Jason Hattrick-Simpers

Publications and source records attributed to Jason Hattrick-Simpers.

At least 37 records · Page 2Linked to original sources

Evaluating the Performance and Robustness of LLMs in Materials Science Q&A and Property Predictions

Large Language Models (LLMs) have the potential to revolutionize scientific research, yet their robustness and reliability in domain-specific applications remain insufficiently explored. In this study, we evaluate the performance and robustness of LLMs for materials science, focusing on domain-specific question answering and materials property prediction across diverse real-world and adversarial conditions. Three distinct datasets are used in this study: 1) a set of multiple-choice questions from undergraduate-level materials science courses, 2) a dataset including various steel compositions and yield strengths, and 3) a band gap dataset, containing textual descriptions of material crystal structures and band gap values. The performance of LLMs is assessed using various prompting strategies, including zero-shot chain-of-thought, expert prompting, and few-shot in-context learning. The robustness of these models is tested against various forms of 'noise', ranging from realistic disturbances to intentionally adversarial manipulations, to evaluate their resilience and reliability under real-world conditions. Additionally, the study showcases unique phenomena of LLMs during predictive tasks, such as mode collapse behavior when the proximity of prompt examples is altered and performance recovery from train/test mismatch. The findings aim to provide informed skepticism for the broad use of LLMs in materials science and to inspire advancements that enhance their robustness and reliability for practical applications.

cs.CL↗

LLM4Mat-Bench: Benchmarking Large Language Models for Materials Property Prediction

Large language models (LLMs) are increasingly being used in materials science. However, little attention has been given to benchmarking and standardized evaluation for LLM-based materials property prediction, which hinders progress. We present LLM4Mat-Bench, the largest benchmark to date for evaluating the performance of LLMs in predicting the properties of crystalline materials. LLM4Mat-Bench contains about 1.9M crystal structures in total, collected from 10 publicly available materials data sources, and 45 distinct properties. LLM4Mat-Bench features different input modalities: crystal composition, CIF, and crystal text description, with 4.7M, 615.5M, and 3.1B tokens in total for each modality, respectively. We use LLM4Mat-Bench to fine-tune models with different sizes, including LLM-Prop and MatBERT, and provide zero-shot and few-shot prompts to evaluate the property prediction capabilities of LLM-chat-like models, including Llama, Gemma, and Mistral. The results highlight the challenges of general-purpose LLMs in materials science and the need for task-specific predictive models and task-specific instruction-tuned LLMs in materials property prediction.

cond-mat.mtrl-sci↗

Open Catalyst Experiments 2024 (OCx24): Bridging Experiments and Computational Models

The search for low-cost, durable, and effective catalysts is essential for green hydrogen production and carbon dioxide upcycling to help in the mitigation of climate change. Discovery of new catalysts is currently limited by the gap between what AI-accelerated computational models predict and what experimental studies produce. To make progress, large and diverse experimental datasets are needed that are reproducible and tested at industrially-relevant conditions. We address these needs by utilizing a comprehensive high-throughput characterization and experimental pipeline to create the Open Catalyst Experiments 2024 (OCX24) dataset. The dataset contains 572 samples synthesized using both wet and dry methods with X-ray fluorescence and X-ray diffraction characterization. We prepared 441 gas diffusion electrodes, including replicates, and evaluated them using zero-gap electrolysis for carbon dioxide reduction (CO$_2$RR) and hydrogen evolution reactions (HER) at current densities up to $300$ mA/cm$^2$. To find correlations with experimental outcomes and to perform computational screens, DFT-verified adsorption energies for six adsorbates were calculated on $\sim$20,000 inorganic materials requiring 685 million AI-accelerated relaxations. Remarkably from this large set of materials, a data driven Sabatier volcano independently identified Pt as being a top candidate for HER without having any experimental measurements on Pt or Pt-alloy samples. We anticipate the availability of experimental data generated specifically for AI training, such as OCX24, will significantly improve the utility of computational models in selecting materials for experimental screening.

cond-mat.mtrl-sci↗

High-Tc superconductor candidates proposed by machine learning

We cast the relation between the chemical composition of a solid-state material and its superconducting critical temperature (Tc) as a statistical learning problem with reduced complexity. Training of query-aware similarity-based ridge regression models on experimental SuperCon data achieve average Tc prediction errors of ~5 K for unseen out-of-sample materials. Two models were trained with one excluding high pressure data in training ("ambient" model) and a second also including high pressure data ("implicit" model). Subsequent utilization of the approach to scan ~153k materials in the Materials Project enables the ranking of candidates by Tc while accounting for thermodynamic stability and small band gap. The ambient model is used to predict stable top three high-Tc candidate materials that include those with large band gaps of LiCuF4 (316 K), Ag2H12S(NO)4 (316 K), and Na2H6PtO6 (315 K). Filtering these candidates for those with small band gaps correspondingly yields LiCuF4 (316 K), Cu2P2O7 (311 K), and Cu3P2H2O9 (307 K).

cond-mat.supr-con↗

An Assessment of Commonly Used Equivalent Circuit Models for Corrosion Analysis: A Bayesian Approach to Electrochemical Impedance Spectroscopy

Electrochemical Impedance Spectroscopy (EIS) is a crucial technique for assessing corrosion of a metallic materials. The analysis of EIS hinges on the selection of an appropriate equivalent circuit model (ECM) that accurately characterizes the system under study. In this work, we systematically examined the applicability of three commonly used ECMs across several typical material degradation scenarios. By applying Bayesian Inference to simulated corrosion EIS data, we assessed the suitability of these ECMs under different corrosion conditions and identified regions where the EIS data lacks sufficient information to statistically substantiate the ECM structure. Additionally, we posit that the traditional approach to EIS analysis, which often requires measurements to very low frequencies, might not be always necessary to correctly model the appropriate ECM. Our study assesses the impact of omitting data from low to medium-frequency ranges on inference results and reveals that a significant portion of low-frequency measurements can be excluded without substantially compromising the accuracy of extracting system parameters. Further, we propose simple checks to the posterior distributions of the ECM components and posterior predictions, which can be used to quantitatively evaluate the suitability of a particular ECM and the minimum frequency required to be measured. This framework points to a pathway for expediting EIS acquisition by intelligently reducing low-frequency data collection and permitting on-the-fly EIS measurements

cond-mat.mtrl-sci↗

Probing out-of-distribution generalization in machine learning for materials

Scientific machine learning (ML) endeavors to develop generalizable models with broad applicability. However, the assessment of generalizability is often based on heuristics. Here, we demonstrate in the materials science setting that heuristics based evaluations lead to substantially biased conclusions of ML generalizability and benefits of neural scaling. We evaluate generalization performance in over 700 out-of-distribution tasks that features new chemistry or structural symmetry not present in the training data. Surprisingly, good performance is found in most tasks and across various ML models including simple boosted trees. Analysis of the materials representation space reveals that most tasks contain test data that lie in regions well covered by training data, while poorly-performing tasks contain mainly test data outside the training domain. For the latter case, increasing training set size or training time has marginal or even adverse effects on the generalization performance, contrary to what the neural scaling paradigm assumes. Our findings show that most heuristically-defined out-of-distribution tests are not genuinely difficult and evaluate only the ability to interpolate. Evaluating on such tasks rather than the truly challenging ones can lead to an overestimation of generalizability and benefits of scaling.

cond-mat.mtrl-sci↗

Roadmap on Data-Centric Materials Science

Science is and always has been based on data, but the terms "data-centric" and the "4th paradigm of" materials research indicate a radical change in how information is retrieved, handled and research is performed. It signifies a transformative shift towards managing vast data collections, digital repositories, and innovative data analytics methods. The integration of Artificial Intelligence (AI) and its subset Machine Learning (ML), has become pivotal in addressing all these challenges. This Roadmap on Data-Centric Materials Science explores fundamental concepts and methodologies, illustrating diverse applications in electronic-structure theory, soft matter theory, microstructure research, and experimental techniques like photoemission, atom probe tomography, and electron microscopy. While the roadmap delves into specific areas within the broad interdisciplinary field of materials science, the provided examples elucidate key concepts applicable to a wider range of topics. The discussed instances offer insights into addressing the multifaceted challenges encountered in contemporary materials research.

cond-mat.mtrl-sci↗

JARVIS-Leaderboard: A Large Scale Benchmark of Materials Design Methods

Lack of rigorous reproducibility and validation are major hurdles for scientific development across many fields. Materials science in particular encompasses a variety of experimental and theoretical approaches that require careful benchmarking. Leaderboard efforts have been developed previously to mitigate these issues. However, a comprehensive comparison and benchmarking on an integrated platform with multiple data modalities with both perfect and defect materials data is still lacking. This work introduces JARVIS-Leaderboard, an open-source and community-driven platform that facilitates benchmarking and enhances reproducibility. The platform allows users to set up benchmarks with custom tasks and enables contributions in the form of dataset, code, and meta-data submissions. We cover the following materials design categories: Artificial Intelligence (AI), Electronic Structure (ES), Force-fields (FF), Quantum Computation (QC) and Experiments (EXP). For AI, we cover several types of input data, including atomic structures, atomistic images, spectra, and text. For ES, we consider multiple ES approaches, software packages, pseudopotentials, materials, and properties, comparing results to experiment. For FF, we compare multiple approaches for material property predictions. For QC, we benchmark Hamiltonian simulations using various quantum algorithms and circuits. Finally, for experiments, we use the inter-laboratory approach to establish benchmarks. There are 1281 contributions to 274 benchmarks using 152 methods with more than 8 million data-points, and the leaderboard is continuously expanding. The JARVIS-Leaderboard is available at the website: https://pages.nist.gov/jarvis_leaderboard

cond-mat.mtrl-sci↗

Efficient first principles based modeling via machine learning: from simple representations to high entropy materials

High-entropy materials (HEMs) have recently emerged as a significant category of materials, offering highly tunable properties. However, the scarcity of HEM data in existing density functional theory (DFT) databases, primarily due to computational expense, hinders the development of effective modeling strategies for computational materials discovery. In this study, we introduce an open DFT dataset of alloys and employ machine learning (ML) methods to investigate the material representations needed for HEM modeling. Utilizing high-throughput DFT calculations, we generate a comprehensive dataset of 84k structures, encompassing both ordered and disordered alloys across a spectrum of up to seven components and the entire compositional range. We apply descriptor-based models and graph neural networks to assess how material information is captured across diverse chemical-structural representations. We first evaluate the in-distribution performance of ML models to confirm their predictive accuracy. Subsequently, we demonstrate the capability of ML models to generalize between ordered and disordered structures, between low-order and high-order alloys, and between equimolar and non-equimolar compositions. Our findings suggest that ML models can generalize from cost-effective calculations of simpler systems to more complex scenarios. Additionally, we discuss the influence of dataset size and reveal that the information loss associated with the use of unrelaxed structures could significantly degrade the generalization performance. Overall, this research sheds light on several critical aspects of HEM modeling and offers insights for data-driven atomistic modeling of HEMs.

cond-mat.mtrl-sci↗

An experimental high-throughput to high-fidelity study towards discovering Al-Cr containing corrosion-resistant compositionally complex alloys

Compositionally complex alloys hold the promise of simultaneously attaining superior combinations of properties, such as corrosion resistance, light-weighting, and strength. Achieving this goal is a challenge due in part to a large number of possible compositions and structures in the vast alloy design space. High-throughput methods offer a path forward, but a strong connection between the synthesis of an alloy of a given composition and structure with its properties has not been fully realized to date. Here, we present the rapid identification of corrosion-resistant alloys based on combinations of Al and Cr in a base Al-Co-Cr-Fe-Ni alloy. Previously unstudied alloy stoichiometries were identified using a combination of high-throughput experimental screening coupled with key metallurgical and electrochemical corrosion tests, identifying alloys with excellent passivation behavior. The alloy native oxide performance and its self-healing attributes were probed using rapid tests in deaerated 0.1 mol/L H2SO4. Importantly, a correlation was found between the electrochemical impedance modulus of the exposure-modified air-formed film and self-healing rate of the CCAs. Multi-element extended x-ray absorption fine structure analyses connected more ordered type chemical short-range order in the Ni-Al 1st nearest-neighbor shell to poorer corrosion resistance. This report underscores the utility of high throughput exploration of compositionally complex alloys for the identification and rapid screening of a vast stoichiometric space.

cond-mat.mtrl-sci↗

Artificial Intelligence-Enabled Optimization of Battery-Grade Lithium Carbonate Production

By 2035, the need for battery-grade lithium is expected to quadruple. About half of this lithium is currently sourced from brines and must be converted from a chloride into lithium carbonate (Li2CO3) through a process called softening. Conventional softening methods using sodium or potassium salts contribute to carbon emissions during reagent mining and battery manufacturing, exacerbating global warming. This study introduces an alternative approach using carbon dioxide (CO2(g)) as the carbonating reagent in the lithium softening process, offering a carbon capture solution. We employed an active learning-driven high-throughput method to rapidly capture CO2(g) and convert it to lithium carbonate. The model was simplified by focusing on the elemental concentrations of C, Li, and N for practical measurement and tracking, avoiding the complexities of ion speciation equilibria. This approach led to an optimized lithium carbonate process that capitalizes on CO2(g) capture and improves the battery metal supply chain's carbon efficiency.

cond-mat.mtrl-sci↗

Reproducibility in Computational Materials Science: Lessons from 'A General-Purpose Machine Learning Framework for Predicting Properties of Inorganic Materials'

The integration of machine learning techniques in materials discovery has become prominent in materials science research and has been accompanied by an increasing trend towards open-source data and tools to propel the field. Despite the increasing usefulness and capabilities of these tools, developers neglecting to follow reproducible practices creates a significant barrier for researchers looking to use or build upon their work. In this study, we investigate the challenges encountered while attempting to reproduce a section of the results presented in "A general-purpose machine learning framework for predicting properties of inorganic materials." Our analysis identifies four major categories of challenges: (1) reporting computational dependencies, (2) recording and sharing version logs, (3) sequential code organization, and (4) clarifying code references within the manuscript. The result is a proposed set of tangible action items for those aiming to make code accessible to, and useful for the community.

cond-mat.mtrl-sci↗

A High Throughput Aqueous Passivation Testing Methodology for Compositionally Complex Alloys using Scanning Droplet Cell

Compositionally complex alloy systems containing more than five principal elements allow exploring a wide range of compositions, processing, and structural variables with the hope for identifying unique properties. Such opportunities also apply to designing materials for improved corrosion resistance, regulated by a self-healing passive film. Such a rich landscape in reactivity and protectivity demands the search for high-throughput experimental testing workflows to uncover key metrics, indicative of superior properties. In this communication, one such methodology is demonstrated for evaluating passivation performance of a combinatorial library of Al0.7-x-yCoxCryFe0.15Ni0.15 thin film alloys in deaerated 0.1 mol/L H2SO4(aq), using a scanning droplet cell.

cond-mat.mtrl-sci↗

On the redundancy in large material datasets: efficient and robust learning with less data

Extensive efforts to gather materials data have largely overlooked potential data redundancy. In this study, we present evidence of a significant degree of redundancy across multiple large datasets for various material properties, by revealing that up to 95 % of data can be safely removed from machine learning training with little impact on in-distribution prediction performance. The redundant data is related to over-represented material types and does not mitigate the severe performance degradation on out-of-distribution samples. In addition, we show that uncertainty-based active learning algorithms can construct much smaller but equally informative datasets. We discuss the effectiveness of informative data in improving prediction performance and robustness and provide insights into efficient data acquisition and machine learning training. This work challenges the "bigger is better" mentality and calls for attention to the information richness of materials data rather than a narrow emphasis on data volume.

cond-mat.mtrl-sci↗

AutoEIS: automated Bayesian model selection and analysis for electrochemical impedance spectroscopy

Electrochemical Impedance Spectroscopy (EIS) is a powerful tool for electrochemical analysis; however, its data can be challenging to interpret. Here, we introduce a new open-source tool named AutoEIS that assists EIS analysis by automatically proposing statistically plausible equivalent circuit models (ECMs). AutoEIS does this without requiring an exhaustive mechanistic understanding of the electrochemical systems. We demonstrate the generalizability of AutoEIS by using it to analyze EIS datasets from three distinct electrochemical systems, including thin-film oxygen evolution reaction (OER) electrocatalysis, corrosion of self-healing multi-principal components alloys, and a carbon dioxide reduction electrolyzer device. In each case, AutoEIS identified competitive or in some cases superior ECMs to those recommended by experts and provided statistical indicators of the preferred solution. The results demonstrated AutoEIS's capability to facilitate EIS analysis without expert labels while diminishing user bias in a high-throughput manner. AutoEIS provides a generalized automated approach to facilitate EIS analysis spanning a broad suite of electrochemical applications with minimal prior knowledge of the system required. This tool holds great potential in improving the efficiency, accuracy, and ease of EIS analysis and thus creates an avenue to the widespread use of EIS in accelerating the development of new electrochemical materials and devices.

cond-mat.mtrl-sci↗

Why is EXAFS analysis for multicomponent metals so hard? Challenges and opportunities for measuring ordering in complex concentrated alloys using x-ray absorption spectroscopy

Short range order is a critical driver of properties (e.g. corrosion resistance and tensile strength) in multicomponent alloys such as complex concentrated alloys (CCAs). Extended x-ray absorption fine structure (EXAFS) is a powerful technique well suited for quantifying this short range order.Here, we described in detail the characteristics of CCAs that make the already challenging task of analyzing EXAFS data even more difficult. We then illustrate novel paths towards robust and scalable quantitative SRO analysis which will accelerate the scientific understanding and development of CCAs.

cond-mat.mtrl-sci↗

Rapid Reconstruction of 3-D Membrane Pore Structure Using a Single 2-D Micrograph

Conventional 2-D scanning electron microscopy (SEM) is commonly used to rapidly and qualitatively evaluate membrane pore structure. Quantitative 2-D analyses of pore sizes can be extracted from SEM, but without information about 3-D spatial arrangement and connectivity, which are crucial to the understanding of membrane pore structure. Meanwhile, experimental 3-D reconstruction via tomography is complex, expensive, and not easily accessible. Here, we employ data-science tools to demonstrate a proof-of-principle reconstruction of the 3-D structure of a membrane using a single 2-D image pulled from a 3-D tomographic data set. The reconstructed and experimental 3-D structures were then directly compared, with important properties such as mean pore radius, mean throat radius, coordination number and tortuosity differing by less than 15%. The developed algorithm will dramatically improve the ability of the membrane community to characterize membranes, accelerating the design and synthesis of membranes with desired structural and transport properties.

cond-mat.mtrl-sci↗

A critical examination of robustness and generalizability of machine learning prediction of materials properties

Recent advances in machine learning (ML) methods have led to substantial improvement in materials property prediction against community benchmarks, but an excellent benchmark score may not imply good generalization of performance. Here we show that ML models trained on the Materials Project 2018 (MP18) dataset can have severely degraded prediction performance on new compounds in the Materials Project 2021 (MP21) dataset. We document performance degradation in graph neural networks and traditional descriptor-based ML models for both quantitative and qualitative predictions. We find the source of the predictive degradation is due to the distribution shift between the MP18 and MP21 versions. This is revealed by the uniform manifold approximation and projection (UMAP) of the feature space. We then show that the performance degradation issue can be foreseen using a few simple tools. Firstly, the UMAP can be used to investigate the connectivity and relative proximity of the training and test data within feature space. Secondly, the disagreement between multiple ML models on the test data can illuminate out-of-distribution samples. We demonstrate that the simple yet efficient UMAP-guided and query-by-committee acquisition strategies can greatly improve prediction accuracy through adding only 1~\% of the test data. We believe this work provides valuable insights for building materials databases and ML models that enable better prediction robustness and generalizability.

cond-mat.mtrl-sci↗