arXiv ScienceSearch

arXiv subjects

Dimitrios Simatos

Publications and source records attributed to Dimitrios Simatos.

3 recordsLinked to original sources

Foundation Models for Discovery and Exploration in Chemical Space

Accurate prediction of atomistic, thermodynamic, and kinetic properties from molecular structures underpins materials innovation. Existing computational and experimental approaches lack the scalability required to navigate chemical space efficiently. Scientific foundation models trained on large unlabelled datasets offer a path towards navigating chemical space across application domains. Here, we develop MIST, a family of molecular foundation models with up to an order of magnitude more parameters and data than prior works. Trained using a novel tokenizer, Smirk, which comprehensively captures nuclear, electronic, and geometric information, MIST learns a diverse range of molecules. MIST models have been fine-tuned to predict more than 400 structure-property relationships and have been shown to match or exceed state-of-the-art performance across diverse benchmarks, from physiology to electrochemistry. We demonstrate the ability of these models to solve real-world problems across chemical space from multiobjective electrolyte solvent screening to stereochemical reasoning for organometallics and mixture property prediction. The clearest demonstration of a foundation model is its ability to solve problems that were neither explicit targets of training nor central to the intentions of its developers. We identify olfactory perception mapping as such a problem, and show that MIST accurately predicted scent profiles and learned a hierarchical representation of olfactory space consistent with hyperbolic geometry. We formulated hyperparameter aware Bayesian neural scaling laws which eliminate the need for hyperparameter sweeps at every scale, making training large compute-optimal models feasible on a limited compute budget. The methods and findings presented here represent a significant step towards accelerating materials discovery, design, and optimization using foundation models.

physics.chem-ph

Structural and dynamic disorder, not ionic trapping, controls charge transport in highly doped conducting polymers

Doped organic semiconductors are critical to emerging device applications, including thermoelectrics, bioelectronics, and neuromorphic computing devices. It is commonly assumed that low conductivities in these materials result primarily from charge trapping by the Coulomb potentials of the dopant counter-ions. Here, we present a combined experimental and theoretical study rebutting this belief. Using a newly developed doping technique, we find the conductivity of several classes of high-mobility conjugated polymers to be strongly correlated with paracrystalline disorder but poorly correlated with ionic size, suggesting that Coulomb traps do not limit transport. A general model for interacting electrons in highly doped polymers is proposed and carefully parameterized against atomistic calculations, enabling the calculation of electrical conductivity within the framework of transient localisation theory. Theoretical calculations are in excellent agreement with experimental data, providing insights into the disordered-limited nature of charge transport and suggesting new strategies to further improve conductivities.

cond-mat.mtrl-sci

Fragment Graphical Variational AutoEncoding for Screening Molecules with Small Data

In the majority of molecular optimization tasks, predictive machine learning (ML) models are limited due to the unavailability and cost of generating big experimental datasets on the specific task. To circumvent this limitation, ML models are trained on big theoretical datasets or experimental indicators of molecular suitability that are either publicly available or inexpensive to acquire. These approaches produce a set of candidate molecules which have to be ranked using limited experimental data or expert knowledge. Under the assumption that structure is related to functionality, here we use a molecular fragment-based graphical autoencoder to generate unique structural fingerprints to efficiently search through the candidate set. We demonstrate that fragment-based graphical autoencoding reduces the error in predicting physical characteristics such as the solubility and partition coefficient in the small data regime compared to other extended circular fingerprints and string based approaches. We further demonstrate that this approach is capable of providing insight into real world molecular optimization problems, such as searching for stabilization additives in organic semiconductors by accurately predicting 92% of test molecules given 69 training examples. This task is a model example of black box molecular optimization as there is minimal theoretical and experimental knowledge to accurately predict the suitability of the additives.

physics.data-an