arXiv ScienceSearch

arXiv subjects

Andrew L. Ferguson

Publications and source records attributed to Andrew L. Ferguson.

At least 19 recordsLinked to original sources

Quantum-accurate atomistic modeling of enzyme catalysis using a machine learned potential

Electronic rearrangements associated with bond forming/breaking in catalytic enzymes require quantum mechanical (QM) treatment beyond classical molecular mechanics (MM). Hybrid QM/MM methods enable tractable simulations but require system-specific setup and are sensitive to the QM region choice and treatment of the QM/MM interface. We demonstrate quantum-accurate treatment of all-atom, complete enzymes in explicit solvent comprising up to 54k atoms and 1 microsecond of total simulation time using the machine-learned interatomic potential (MLIP) eSEN-omol. We reproduce experimental barrier trends for Claisen rearrangement in chorismate mutase, resolve critical intermediate states in PETase catalyzed polymer depolymerization, and distinguish mechanistic alternatives for metal-activated phosphoryl transfer in nucleoside diphosphate kinase. We realize 1000x speedups relative to typical QM/MM calculations without system-specific tuning. These results establish MLIPs as a practical route to QM-accurate simulations of enzyme catalysis.

physics.chem-ph

Data-driven reconstruction of dynamical systems using Takens' Theorem, manifold learning, and universal function approximators

Embedding theorems can be used to provide theoretical guarantees about the relation between low-dimensional observations of a system and its full-dimensional state and dynamics. Such theorems do not, however, provide guidance on observable choice, embedding construction, or methodologies to learn the mapping between the embedding and full-dimensional state. In this work, we develop an algorithmic framework, TAkens Reconstruction (TAR), to analyze and reconstruct arbitrary dynamical systems from low-dimensional time series using an integration of Takens' Delay Embedding Theorem, manifold learning techniques, and universal function approximators. We validate TAR in applications to a variety of simulated and observed dynamical systems and use it to investigate how delay vector structure impacts reconstruction accuracy. In an ecological system, we show that simple predator-prey dynamics can be reconstructed with observations taken over a wide variety of embedding time scales. In molecular dynamics simulations of the protein Villin, we demonstrate how including multiple time delays of the same observable series can be used to improve reconstruction of systems with multiple characteristic time scales. In the trade record of Vanguard S&P 500, we show how the approach exposes underlying dynamical phenomenologies in the data and accurate return predictions over short time horizons without access to full-dimensional market observations. We develop and release an open-source software package to enable the application of TAR to arbitrary dynamical systems.

physics.comp-ph

Characterizing Defect Dynamics in Silicon Carbide Using Symmetry-Adapted Collective Variables and Machine Learning Interatomic Potentials

Silicon carbide (SiC) divacancies are attractive candidates for spin defect qubits possessing long coherence times and optical addressability. The high activation barriers associated with SiC defect formation and motion pose challenges for their study by first-principles molecular dynamics. In this work, we develop and deploy machine learning interatomic potentials (MLIPs) to accelerate defect dynamics simulations while retaining ab initio accuracy. We employ an active learning strategy comprising symmetry-adapted collective variable discovery and enhanced sampling to compile configurationally diverse training data, calculation of energies and forces using density functional theory (DFT), and training of an E(3)-equivariant MLIP based on the Allegro model. The trained MLIP reproduces DFT-level accuracy in defect transition activation free energy barriers, enables the efficient and stable simulation of multi-defect 216-atom supercells, and permits an analysis of the temperature dependence of defect thermodynamic stability and formation/annihilation kinetics to propose an optimal annealing temperature to maximally stabilize VV divacancies.

cond-mat.mtrl-sci

FlowBack-Adjoint: Physics-Aware and Energy-Guided Conditional Flow-Matching for All-Atom Protein Backmapping

Coarse-grained (CG) molecular models of proteins can substantially increase the time and length scales accessible to molecular dynamics simulations of proteins, but recovery of accurate all-atom (AA) ensembles from CG simulation trajectories can be essential for exposing molecular mechanisms of folding and docking and for calculation of physical properties requiring atomistic detail. The recently reported deep generative model FlowBack restores AA detail to protein C-alpha traces using a flow-matching architecture and demonstrates state-of-the-art performance in generation of AA structural ensembles. Training, however, is performed exclusively on structural data and the absence of any awareness of interatomic energies or forces within training results in small fractions of incorrect bond lengths, atomic clashes, and otherwise high-energy structures. In this work, we introduce FlowBack-Adjoint as a lightweight enhancement that upgrades the pre-trained FlowBack model through a one-time, physics-aware post-training pass. Auxiliary contributions to the flow introduce physical awareness of bond lengths and Lennard-Jones interactions and gradients of a molecular mechanics force field energy are incorporated via adjoint matching to steer the FlowBack-Adjoint vector field to produce lower-energy configurations. In benchmark tests against FlowBack, FlowBack-Adjoint lowers single-point energies by a median of ~78 kcal/mol.residue, reduces errors in bond lengths by >92%, eliminates >98% of molecular clashes, maintains excellent diversity of the AA configurational ensemble, and produces configurations capable of initializing stable all-atom molecular dynamics simulations without requiring energy relaxation. We propose FlowBack-Adjoint as an accurate and efficient physics-aware deep generative model for AA backmapping from C-alpha traces.

physics.chem-ph

Weighted Active Space Protocol for Multireference Machine-Learned Potentials

Multireference methods such as multiconfiguration pair-density functional theory (MC-PDFT) offer an effective means of capturing electronic correlation in systems with significant multiconfigurational character. However, their application to train machine learning-based interatomic potentials (MLPs) for catalytic dynamics has been challenging due to the sensitivity of multireference calculations to the underlying active space, which complicates achieving consistent energies and gradients across diverse nuclear configurations. To overcome this limitation, we introduce the Weighted Active-Space Protocol (WASP), a systematic approach to assign a consistent active space for a given system across uncorrelated configurations. By integrating WASP with MLPs and enhanced sampling techniques, we propose a data-efficient active learning cycle that enables the training of an MLP on multireference data. We demonstrate the method on the TiC+-catalyzed C-H activation of methane, a reaction that poses challenges for Kohn-Sham density functional theory due to its significant multireference character. This framework enables accurate and efficient modeling of catalytic dynamics, establishing a new paradigm for simulating complex reactive processes beyond the limits of conventional electronic-structure methods.

physics.chem-ph

PLUMED Tutorials: a collaborative, community-driven learning ecosystem

In computational physics, chemistry, and biology, the implementation of new techniques in a shared and open source software lowers barriers to entry and promotes rapid scientific progress. However, effectively training new software users presents several challenges. Common methods like direct knowledge transfer and in-person workshops are limited in reach and comprehensiveness. Furthermore, while the COVID-19 pandemic highlighted the benefits of online training, traditional online tutorials can quickly become outdated and may not cover all the software's functionalities. To address these issues, here we introduce ``PLUMED Tutorials'', a collaborative model for developing, sharing, and updating online tutorials. This initiative utilizes repository management and continuous integration to ensure compatibility with software updates. Moreover, the tutorials are interconnected to form a structured learning path and are enriched with automatic annotations to provide broader context. This paper illustrates the development, features, and advantages of PLUMED Tutorials, aiming to foster an open community for creating and sharing educational resources.

physics.ed-ph

A Modular and Extensible CHARMM-Compatible Model for All-Atom Simulation of Polypeptoids

Peptoids (N-substituted glycines) are a class of sequence-defined synthetic peptidomimetic polymers with applications including drug delivery, catalysis, and biomimicry. Classical molecular simulations have been used to predict and understand the conformational dynamics of single peptoid chains and their self-assembly into diverse morphologies including sheets, tubes, spheres, and fibrils. The CGenFF-NTOID model based on the CHARMM General ForceField has demonstrated success in enabling accurate all-atom molecular modeling of the structure and thermodynamic behavior of peptoids. Extension of this force field to new peptoid side chain chemistries has historically required parameterization of new side chain bonded interactions against ab initio and/or experimental data. This fitting protocol improves the accuracy of the force field but is also burdensome and time consuming, and precludes modular extensibility of the model to arbitrary peptoid sequences. In this work, we develop and demonstrate a Modular Side Chain CGenFF-NTOID (MoSiC-CGenFF-NTOID) as an extension of CGenFF-NTOID employing a modular decomposition of the peptoid backbone and side chain parameterizations wherein arbitrary side chain chemistries within the large family of substituted methyl groups (i.e., -CH3, -CH2R, -CHRR' -CRR'R'') are directly ported from CGenFF without any additional reparameterization. We validate this approach against ab initio calculations and experimental data to to develop a MoSiC-CGenFF-NTOID model for all 20 natural amino acid side chains along with 13 commonly-used synthetic side chains, and present an extensible paradigm to efficiently determine whether a novel side chain can be directly incorporated into the model or whether refitting of the CGenFF parameters is warranted. We make the model freely available to the community along with a tool to perform automated initial structure generation.

cond-mat.soft

Permutationally Invariant Networks for Enhanced Sampling (PINES): Discovery of Multi-Molecular and Solvent-Inclusive Collective Variables

The typically rugged nature of molecular free energy landscapes can frustrate efficient sampling of the thermodynamically relevant phase space due to the presence of high free energy barriers. Enhanced sampling techniques can improve phase space exploration by accelerating sampling along particular collective variables (CVs). A number of techniques exist for data-driven discovery of CVs parameterizing the important large scale motions of the system. A challenge to CV discovery is learning CVs invariant to symmetries of the molecular system, frequently rigid translation, rigid rotation, and permutational relabeling of identical particles. Of these, permutational invariance have proved a persistent challenge in frustrating the the data-driven discovery of multi-molecular CVs in systems of self-assembling particles and solvent-inclusive CVs for solvated systems. In this work, we integrate Permutation Invariant Vector (PIV) featurizations with autoencoding neural networks to learn nonlinear CVs invariant to translation, rotation, and permutation, and perform interleaved rounds of CV discovery and enhanced sampling to iteratively expand sampling of configurational phase space and obtain converged CVs and free energy landscapes. We demonstrate the Permutationally Invariant Network for Enhanced Sampling (PINES) approach in applications to the self-assembly of a 13-atom Argon cluster, association/dissociation of a NaCl ion pair in water, and hydrophobic collapse of a C45H92 n-pentatetracontane polymer chain. We make the approach freely available as a new module within the PLUMED2 enhanced sampling libraries.

q-bio.BM

DiAMoNDBack: Diffusion-denoising Autoregressive Model for Non-Deterministic Backmapping of C{\alpha} Protein Traces

Coarse-grained molecular models of proteins permit access to length and time scales unattainable by all-atom models and the simulation of processes that occur on long-time scales such as aggregation and folding. The reduced resolution realizes computational accelerations but an atomistic representation can be vital for a complete understanding of mechanistic details. Backmapping is the process of restoring all-atom resolution to coarse-grained molecular models. In this work, we report DiAMoNDBack (Diffusion-denoising Autoregressive Model for Non-Deterministic Backmapping) as an autoregressive denoising diffusion probability model to restore all-atom details to coarse-grained protein representations retaining only C{\alpha} coordinates. The autoregressive generation process proceeds from the protein N-terminus to C-terminus in a residue-by-residue fashion conditioned on the C{\alpha} trace and previously backmapped backbone and side chain atoms within the local neighborhood. The local and autoregressive nature of our model makes it transferable between proteins. The stochastic nature of the denoising diffusion process means that the model generates a realistic ensemble of backbone and side chain all-atom configurations consistent with the coarse-grained C{\alpha} trace. We train DiAMoNDBack over 65k+ structures from Protein Data Bank (PDB) and validate it in applications to a hold-out PDB test set, intrinsically-disordered protein structures from the Protein Ensemble Database (PED), molecular dynamics simulations of fast-folding mini-proteins from DE Shaw Research, and coarse-grained simulation data. We achieve state-of-the-art reconstruction performance in terms of correct bond formation, avoidance of side chain clashes, and diversity of the generated side chain configurational states. We make DiAMoNDBack model publicly available as a free and open source Python package.

q-bio.BM

PySAGES: flexible, advanced sampling methods accelerated with GPUs

Molecular simulations are an important tool for research in physics, chemistry, and biology. The capabilities of simulations can be greatly expanded by providing access to advanced sampling methods and techniques that permit calculation of the relevant underlying free energy landscapes. In this sense, software that can be seamlessly adapted to a broad range of complex systems is essential. Building on past efforts to provide open-source community supported software for advanced sampling, we introduce PySAGES, a Python implementation of the Software Suite for Advanced General Ensemble Simulations (SSAGES) that provides full GPU support for massively parallel applications of enhanced sampling methods such as adaptive biasing forces, harmonic bias, or forward flux sampling in the context of molecular dynamics simulations. By providing an intuitive interface that facilitates the management of a system's configuration, the inclusion of new collective variables, and the implementation of sophisticated free energy-based sampling methods, the PySAGES library serves as a general platform for the development and implementation of emerging simulation techniques. The capabilities, core features, and computational performance of this new tool are demonstrated with clear and concise examples pertaining to different classes of molecular systems. We anticipate that PySAGES will provide the scientific community with a robust and easily accessible platform to accelerate simulations, improve sampling, and enable facile estimation of free energies for a wide range of materials and processes.

physics.comp-ph

GANs and Closures: Micro-Macro Consistency in Multiscale Modeling

Sampling the phase space of molecular systems -- and, more generally, of complex systems effectively modeled by stochastic differential equations -- is a crucial modeling step in many fields, from protein folding to materials discovery. These problems are often multiscale in nature: they can be described in terms of low-dimensional effective free energy surfaces parametrized by a small number of "slow" reaction coordinates; the remaining "fast" degrees of freedom populate an equilibrium measure on the reaction coordinate values. Sampling procedures for such problems are used to estimate effective free energy differences as well as ensemble averages with respect to the conditional equilibrium distributions; these latter averages lead to closures for effective reduced dynamic models. Over the years, enhanced sampling techniques coupled with molecular simulation have been developed. An intriguing analogy arises with the field of Machine Learning (ML), where Generative Adversarial Networks can produce high dimensional samples from low dimensional probability distributions. This sample generation returns plausible high dimensional space realizations of a model state, from information about its low-dimensional representation. In this work, we present an approach that couples physics-based simulations and biasing methods for sampling conditional distributions with ML-based conditional generative adversarial networks for the same task. The "coarse descriptors" on which we condition the fine scale realizations can either be known a priori, or learned through nonlinear dimensionality reduction. We suggest that this may bring out the best features of both approaches: we demonstrate that a framework that couples cGANs with physics-based enhanced sampling techniques can improve multiscale SDE dynamical systems sampling, and even shows promise for systems of increasing complexity.

cs.LG

Reconstruction of Protein Structures from Single-Molecule Time Series

Single-molecule experimental techniques track the real-time dynamics of molecules by recording a small number of experimental observables. Following these observables provides a coarse-grained, low-dimensional representation of the conformational dynamics but does not furnish an atomistic representation of the instantaneous molecular structure. Takens' Delay Embedding Theorem asserts that, under quite general conditions, these low-dimensional time series can contain sufficient information to reconstruct the full molecular configuration of the system up to an a priori unknown transformation. By combining Takens' Theorem with tools from statistical thermodynamics, manifold learning, artificial neural networks, and rigid graph theory, we establish an approach Single-molecule TAkens Reconstruction (STAR) to learn this transformation and reconstruct molecular configurations from time series in experimentally-measurable observables such as intramolecular distances accessible to single molecule F\"orster resonance energy transfer. We demonstrate the approach in applications to molecular dynamics simulations of a C24H50 polymer chain and the artificial mini-protein Chignolin. The trained models reconstruct molecular configurations from synthetic time series data in the head-to-tail molecular distances with atomistic root mean squared deviation accuracies better than 0.2 nm. This work demonstrates that it is possible to accurately reconstruct protein structures from time series in experimentally-measurable observables and establishes the theoretical and algorithmic foundations to do so in applications to real experimental data.

physics.comp-ph

Molecular Latent Space Simulators

Small integration time steps limit molecular dynamics (MD) simulations to millisecond time scales. Markov state models (MSMs) and equation-free approaches learn low-dimensional kinetic models from MD simulation data by performing configurational or dynamical coarse-graining of the state space. The learned kinetic models enable the efficient generation of dynamical trajectories over vastly longer time scales than are accessible by MD, but the discretization of configurational space and/or absence of a means to reconstruct molecular configurations precludes the generation of continuous all-atom molecular trajectories. We propose latent space simulators (LSS) to learn kinetic models for continuous all-atom simulation trajectories by training three deep learning networks to (i) learn the slow collective variables of the molecular system, (ii) propagate the system dynamics within this slow latent space, and (iii) generatively reconstruct molecular configurations. We demonstrate the approach in an application to Trp-cage miniprotein to produce novel ultra-long synthetic folding trajectories that accurately reproduce all-atom molecular structure, thermodynamics, and kinetics at six orders of magnitude lower cost than MD. The dramatically lower cost of trajectory generation enables greatly improved sampling and greatly reduced statistical uncertainties in estimated thermodynamic averages and kinetic rates.

physics.comp-ph

Discovery of Self-Assembling $\pi$-Conjugated Peptides by Active Learning-Directed Coarse-Grained Molecular Simulation

Electronically-active organic molecules have demonstrated great promise as novel soft materials for energy harvesting and transport. Self-assembled nanoaggregates formed from $\pi$-conjugated oligopeptides composed of an aromatic core flanked by oligopeptide wings offer emergent optoelectronic properties within a water soluble and biocompatible substrate. Nanoaggregate properties can be controlled by tuning core chemistry and peptide composition, but the sequence-structure-function relations remain poorly characterized. In this work, we employ coarse-grained molecular dynamics simulations within an active learning protocol employing deep representational learning and Bayesian optimization to efficiently identify molecules capable of assembling pseudo-1D nanoaggregates with good stacking of the electronically-active $\pi$-cores. We consider the DXXX-OPV3-XXXD oligopeptide family, where D is an Asp residue and OPV3 is an oligophenylene vinylene oligomer (1,4-distyrylbenzene), to identify the top performing XXX tripeptides within all 20$^3$ = 8,000 possible sequences. By direct simulation of only 2.3% of this space, we identify molecules predicted to exhibit superior assembly relative to those reported in prior work. Spectral clustering of the top candidates reveals new design rules governing assembly. This work establishes new understanding of DXXX-OPV3-XXXD assembly, identifies promising new candidates for experimental testing, and presents a computational design platform that can be generically extended to other peptide-based and peptide-like systems.

q-bio.BM

Statistically optimal continuous free energy surfaces from biased simulations and multistate reweighting

Free energies as a function of a selected set of collective variables are commonly computed in molecular simulation and of significant value in understanding and engineering molecular behavior. These free energy surfaces are most commonly estimated using variants of histogramming techniques, but such approaches obscure two important facets of these functions. First, the empirical observations along the collective variable are defined by an ensemble of discrete observations and the coarsening of these observations into a histogram bins incurs unnecessary loss of information. Second, the free energy surface is itself almost always a continuous function, and its representation by a histogram introduces inherent approximations due to the discretization. In this study, we relate the observed discrete observations from biased simulations to the inferred underlying continuous probability distribution over the collective variables and derive histogram-free techniques for estimating this free energy surface. We reformulate free energy surface estimation as minimization of a Kullback-Leibler divergence between a continuous trial function and the discrete empirical distribution and show that this is equivalent to likelihood maximization of a trial function given a set of sampled data. We then present a fully Bayesian treatment of this formalism, which enables the incorporation of powerful Bayesian tools such as the inclusion of regularizing priors, uncertainty quantification, and model selection techniques. We demonstrate this new formalism in the analysis of umbrella sampling simulations for the $\chi$ torsion of a valine sidechain in the L99A mutant of T4 lysozyme with benzene bound in the cavity.

cond-mat.stat-mech

High-resolution Markov state models for the dynamics of Trp-cage miniprotein constructed over slow folding modes identified by state-free reversible VAMPnets

State-free reversible VAMPnets (SRVs) are a neural network-based framework capable of learning the leading eigenfunctions of the transfer operator of a dynamical system from trajectory data. In molecular dynamics simulations, these data-driven collective variables (CVs) capture the slowest modes of the dynamics and are useful for enhanced sampling and free energy estimation. In this work, we employ SRV coordinates as a feature set for Markov state model (MSM) construction. Compared to the current state of the art, MSMs constructed from SRV coordinates are more robust to the choice of input features, exhibit faster implied timescale convergence, and permit the use of shorter lagtimes to construct higher kinetic resolution models. We apply this methodology to study the folding kinetics and conformational landscape of the Trp-cage miniprotein. Folding and unfolding mean first passage times are in good agreement with prior literature, and a nine macrostate model is presented. The unfolded ensemble comprises a central kinetic hub with interconversions to several metastable unfolded conformations and which serves as the gateway to the folded ensemble. The folded ensemble comprises the native state, a partially unfolded intermediate "loop" state, and a previously unreported short-lived intermediate that we were able to resolve due to the high time-resolution of the SRV-MSM. We propose SRVs as an excellent candidate for integration into modern MSM construction pipelines.

physics.bio-ph

Capabilities and Limitations of Time-lagged Autoencoders for Slow Mode Discovery in Dynamical Systems

Time-lagged autoencoders (TAEs) have been proposed as a deep learning regression-based approach to the discovery of slow modes in dynamical systems. However, a rigorous analysis of nonlinear TAEs remains lacking. In this work, we discuss the capabilities and limitations of TAEs through both theoretical and numerical analyses. Theoretically, we derive bounds for nonlinear TAE performance in slow mode discovery and show that in general TAEs learn a mixture of slow and maximum variance modes. Numerically, we illustrate cases where TAEs can and cannot correctly identify the leading slowest mode in two example systems: a 2D "Washington beltway" potential and the alanine dipeptide molecule in explicit water. We also compare the TAE results with those obtained using state-free reversible VAMPnets (SRVs) as a variational-based neural network approach for slow modes discovery, and show that SRVs can correctly discover slow modes where TAEs fail.

stat.ML

Landmark Diffusion Maps (L-dMaps): Accelerated manifold learning out-of-sample extension

Diffusion maps are a nonlinear manifold learning technique based on harmonic analysis of a diffusion process over the data. Out-of-sample extensions with computational complexity $\mathcal{O}(N)$, where $N$ is the number of points comprising the manifold, frustrate applications to online learning applications requiring rapid embedding of high-dimensional data streams. We propose landmark diffusion maps (L-dMaps) to reduce the complexity to $\mathcal{O}(M)$, where $M \ll N$ is the number of landmark points selected using pruned spanning trees or k-medoids. Offering $(N/M)$ speedups in out-of-sample extension, L-dMaps enables the application of diffusion maps to high-volume and/or high-velocity streaming data. We illustrate our approach on three datasets: the Swiss roll, molecular simulations of a C$_{24}$H$_{50}$ polymer chain, and biomolecular simulations of alanine dipeptide. We demonstrate up to 50-fold speedups in out-of-sample extension for the molecular systems with less than 4% errors in manifold reconstruction fidelity relative to calculations over the full dataset.

stat.ML