arXiv ScienceSearch

arXiv subjects

Markus Rampp

Publications and source records attributed to Markus Rampp.

At least 19 recordsLinked to original sources

ESPResSo++: A Fast and Extensible Molecular Simulation Package for Coarse-Grained Models

ESPResSo++ is an open-source software package for molecular dynamics (MD) simulations with a particular emphasis on coarse-grained (CG) models of soft matter systems. Written in C++ with a flexible Python interface, it is designed for high-performance computing (HPC) environments and supports massively parallel simulations through MPI. The package enables simulations of polymers, membranes, colloids and complex fluids with a wide range of interaction models and advanced algorithms.

cond-mat.soft

FROST-CLUSTERS -- III. Metallicity-dependent intermediate mass black hole formation by runaway collisions in dense star clusters

We explore the formation of intermediate mass black holes (IMBHs), potential seeds for supermassive black holes (SMBHs), via runaway stellar collisions for a wide range of star cluster (surface) densities ($4\times10^3 M_\odot$ pc$^{-2} \lesssim \Sigma_\mathrm{h} \lesssim 4\times10^6 M_\odot$ pc$^{-2}$) and metallicities $(0.01 Z_\odot \lesssim Z \lesssim 1.0 Z_\odot)$. Our sample of isolated (>1400) and hierarchical (30) simulations of young, massive star clusters with up to $N=1.8\times10^6$ stars includes collisional stellar dynamics, stellar evolution, and post-Newtonian equations of motion for black holes using the BIFROST code. High stellar wind rates suppress IMBH formation at high metallicities ($Z\gtrsim0.2 Z_\odot$) and low collision rates prevent their formation at low densities ($\Sigma_\mathrm{h}\lesssim 3\times10^4 M_\odot$ pc$^{-2}$). The assumptions about stellar wind loss rates strongly affect the maximum final IMBH masses ($M_\bullet\sim 6000 M_\odot$ vs. $25000 M_\odot$). The total stellar mass loss from collisions and collisionally boosted winds before $t=3$ Myr can together reach up to $5$--$10\%$ of the final cluster mass. We present fitting formulae for IMBH masses as a function of host star cluster $\Sigma_\mathrm{h}$ and $Z$ which can be used to seed SMBHs in high resolution cosmological hydrodynamical simulations and in semi-analytic models for galaxy formation. Our results favour IMBH formation in dense low metallicity environments similar to $z\sim10$ James Webb Space Telescope (\textit{JWST}) proto globular clusters. IMBH formation is suppressed in the high metallicity and low density conditions of the local Universe.

astro-ph.GA

A high-performance and portable implementation of the SISSO method for CPUs and GPUs

SISSO (sure-independence screening and sparsifying operator) is an artificial intelligence (AI) method based on symbolic regression and compressed sensing widely used in materials science research. SISSO++ is its C++ implementation that employs MPI and OpenMP for parallelization, rendering it well-suited for high-performance computing (HPC) environments. As heterogeneous hardware becomes mainstream in the HPC and AI fields, we chose to port the SISSO++ code to GPUs using the Kokkos performance-portable library. Kokkos allows us to maintain a single codebase for both Nvidia and AMD GPUs, significantly reducing the maintenance effort. In this work, we summarize the necessary code changes we did to achieve hardware and performance portability. This is accompanied by performance benchmarks on Nvidia and AMD GPUs. We demonstrate the speedups obtained from using GPUs across the three most time-consuming parts of our code.

cs.PF

FORTE: An Open-Source System for Cost-Effective and Scalable Environmental Monitoring

Forests are an essential part of our biosphere, regulating climate, acting as a sink for greenhouse gases, and providing numerous other ecosystem services. However, they are negatively impacted by climatic stressors such as drought or heat waves. In this paper, we introduce FORTE, an open-source system for environmental monitoring with the aim of understanding how forests react to such stressors. It consists of two key components: (1) a wireless sensor network (WSN) deployed in the forest for data collection, and (2) a Data Infrastructure for data processing, storage, and visualization. The WSN contains a Central Unit capable of transmitting data to the Data Infrastructure via LTE-M and several spatially independent Satellites that collect data over large areas and transmit them wirelessly to the Central Unit. Our prototype deployments show that our solution is cost-effective compared to commercial solutions, energy-efficient with sensor nodes lasting for several months on a single charge, and reliable in terms of data quality. FORTE's flexible architecture makes it suitable for a wide range of environmental monitoring applications beyond forest monitoring. The contributions of this paper are three-fold. First, we describe the high-level requirements necessary for developing an environmental monitoring system. Second, we present an architecture and prototype implementation of the requirements by introducing our FORTE platform and demonstrating its effectiveness through multiple field tests. Lastly, we provide source code, documentation, and hardware design artifacts as part of our open-source repository.

q-bio.PE

A Study of Performance Portability in Plasma Physics Simulations

The high-performance computing (HPC) community has recently seen a substantial diversification of hardware platforms and their associated programming models. From traditional multicore processors to highly specialized accelerators, vendors and tool developers back up the relentless progress of those architectures. In the context of scientific programming, it is fundamental to consider performance portability frameworks, i.e., software tools that allow programmers to write code once and run it on different computer architectures without sacrificing performance. We report here on the benefits and challenges of performance portability using a field-line tracing simulation and a particle-in-cell code, two relevant applications in computational plasma physics with applications to magnetically-confined nuclear-fusion energy research. For these applications we report performance results obtained on four HPC platforms with server-class CPUs from Intel (Xeon) and AMD (EPYC), and high-end GPUs from Nvidia and AMD, including the latest Nvidia H100 GPU and the novel AMD Instinct MI300A APU. Our results show that both Kokkos and OpenMP are powerful tools to achieve performance portability and decent "out-of-the-box" performance, even for the very latest hardware platforms. For our applications, Kokkos provided performance portability to the broadest range of hardware architectures from different vendors.

physics.plasm-ph

3D deep learning for enhanced atom probe tomography analysis of nanoscale microstructures

Quantitative analysis of microstructural features on the nanoscale, including precipitates, local chemical orderings (LCOs) or structural defects (e.g. stacking faults) plays a pivotal role in understanding the mechanical and physical responses of engineering materials. Atom probe tomography (APT), known for its exceptional combination of chemical sensitivity and sub-nanometer resolution, primarily identifies microstructures through compositional segregations. However, this fails when there is no significant segregation, as can be the case for LCOs and stacking faults. Here, we introduce a 3D deep learning approach, AtomNet, designed to process APT point cloud data at the single-atom level for nanoscale microstructure extraction, simultaneously considering compositional and structural information. AtomNet is showcased in segmenting L12-type nanoprecipitates from the matrix in an AlLiMg alloy, irrespective of crystallographic orientations, which outperforms previous methods. AtomNet also allows for 3D imaging of L10-type LCOs in an AuCu alloy, a challenging task for conventional analysis due to their small size and subtle compositional differences. Finally, we demonstrate the use of AtomNet for revealing 2D stacking faults in a Co-based superalloy, without any defected training data, expanding the capabilities of APT for automated exploration of hidden microstructures. AtomNet pushes the boundaries of APT analysis, and holds promise in establishing precise quantitative microstructure-property relationships across a diverse range of metallic materials.

cond-mat.mtrl-sci

Efficient All-electron Hybrid Density Functionals for Atomistic Simulations Beyond 10,000 Atoms

Hybrid density functional approximations (DFAs) offer compelling accuracy for ab initio electronic-structure simulations of molecules, nanosystems, and bulk materials, addressing some deficiencies of computationally cheaper, frequently used semilocal DFAs. However, the computational bottleneck of hybrid DFAs is the evaluation of the non-local exact exchange contribution, which is the limiting factor for the application of the method for large-scale simulations. In this work, we present a drastically optimized resolution-of-identity-based real-space implementation of the exact exchange evaluation for both non-periodic and periodic boundary conditions in the all-electron code FHI-aims, targeting high-performance CPU compute clusters. The introduction of several new refined Message Passing Interface (MPI) parallelization layers and shared memory arrays according to the MPI-3 standard were the key components of the optimization. We demonstrate significant improvements of memory and performance efficiency, scalability, and workload distribution, extending the reach of hybrid DFAs to simulation sizes beyond ten thousand atoms. As a necessary byproduct of this work, other code parts in FHI-aims have been optimized as well, e.g., the computation of the Hartree potential and the evaluation of the force and stress components. We benchmark the performance and scaling of the hybrid DFA based simulations for a broad range of chemical systems, including hybrid organic-inorganic perovskites, organic crystals and ice crystals with up to 30,576 atoms (101,920 electrons described by 244,608 basis functions).

cond-mat.mtrl-sci

Roadmap on Data-Centric Materials Science

Science is and always has been based on data, but the terms "data-centric" and the "4th paradigm of" materials research indicate a radical change in how information is retrieved, handled and research is performed. It signifies a transformative shift towards managing vast data collections, digital repositories, and innovative data analytics methods. The integration of Artificial Intelligence (AI) and its subset Machine Learning (ML), has become pivotal in addressing all these challenges. This Roadmap on Data-Centric Materials Science explores fundamental concepts and methodologies, illustrating diverse applications in electronic-structure theory, soft matter theory, microstructure research, and experimental techniques like photoemission, atom probe tomography, and electron microscopy. While the roadmap delves into specific areas within the broad interdisciplinary field of materials science, the provided examples elucidate key concepts applicable to a wider range of topics. The discussed instances offer insights into addressing the multifaceted challenges encountered in contemporary materials research.

cond-mat.mtrl-sci

Machine learning-enabled tomographic imaging of chemical short-range atomic ordering

In solids, chemical short-range order (CSRO) refers to the self-organisation of atoms of certain species occupying specific crystal sites. CSRO is increasingly being envisaged as a lever to tailor the mechanical and functional properties of materials. Yet quantitative relationships between properties and the morphology, number density, and atomic configurations of CSRO domains remain elusive. Herein, we showcase how machine learning-enhanced atom probe tomography (APT) can mine the near-atomically resolved APT data and jointly exploit the technique's high elemental sensitivity to provide a 3D quantitative analysis of CSRO in a CoCrNi medium-entropy alloy. We reveal multiple CSRO configurations, with their formation supported by state-of-the-art Monte-Carlo simulations. Quantitative analysis of these CSROs allows us to establish relationships between processing parameters and physical properties. The unambiguous characterization of CSRO will help refine strategies for designing advanced materials by manipulating atomic-scale architectures.

cond-mat.mtrl-sci

Electron inertia effects in 3D hybrid-kinetic collisionless plasma turbulence

The effects of the electron inertia on the current sheets that are formed out of kinetic turbulence are relevant to understand the importance of coherent structures in turbulence and the nature of turbulence at the dissipation scales. We investigate this problem by carrying out 3D hybrid-kinetic Particle-in-Cell (PIC) simulations of decaying kinetic turbulence with our CHIEF code. The main distinguishing feature of this code is an implementation of the electron inertia without approximations. Our simulation results show that the electron inertia plays an important role in regulating and limiting the largest values of current density in both real and wavenumber Fourier space, in particular near and, unexpectedly, even above electron scales. In addition, the electric field associated to the electron inertia dominates most of the strongest current sheets. The electron inertia is thus important to accurately describe the properties of current sheets formed in turbulence at electron scales.

physics.plasm-ph

Importance of accurate consideration of the electron inertia in hybrid-kinetic simulations of collisionless plasma turbulence: 1. The 2D limit

The dissipation mechanism of the magnetic energy in turbulent collisionless space and astrophysical plasmas is still not well understood. Its investigation requires efficient kinetic simulations of the energy transfer in collisionless plasma turbulence. In this respect, hybrid-kinetic simulations, in which ions are treated as particles and electrons as an inertial fluid, have begun to attract a significant interest recently. Hybrid-kinetic models describe both ion- and electron scale processes by ignoring electron kinetic effects so that they are computationally much less demanding compared to fully kinetic plasma models. Hybrid-kinetic codes solve either the Vlasov equation for the ions (Eulerian Vlasov-hybrid codes) or the equations of motion of the ions as macro-particles (Lagrangian Particle-in-Cell (PIC)-hybrid codes). They consider the inertia of the electron fluid using different approximations. We check the validity of these approximations by employing our recently massively parallelized three-dimensional PIC-hybrid code CHIEF which considers the electron inertia without any of the common approximations. In particular we report the results of simulations of two-dimensional collisionless plasma turbulence. We conclude that the simulation results obtained using hybrid-kinetic codes which use approximations to describe the electron inertia need to be interpreted with caution. We also discuss the parallel scalability of CHIEF, to the best of our knowledge, the first PIC-hybrid code which without approximations describes the inertial electron fluid.

physics.plasm-ph

Code modernization strategies for short-range non-bonded molecular dynamics simulations

Modern HPC systems are increasingly relying on greater core counts and wider vector registers. Thus, applications need to be adapted to fully utilize these hardware capabilities. One class of applications that can benefit from this increase in parallelism are molecular dynamics simulations. In this paper, we describe our efforts at modernizing the ESPResSo++ molecular dynamics simulation package by restructuring its particle data layout for efficient memory accesses and applying vectorization techniques to benefit the calculation of short-range non-bonded forces, which results in an overall three times speedup and serves as a baseline for further optimizations. We also implement fine-grained parallelism for multi-core CPUs through HPX, a C++ runtime system which uses lightweight threads and an asynchronous many-task approach to maximize concurrency. Our goal is to evaluate the performance of an HPX-based approach compared to the bulk-synchronous MPI-based implementation. This requires the introduction of an additional layer to the domain decomposition scheme that defines the task granularity. On spatially inhomogeneous systems, which impose a corresponding load-imbalance in traditional MPI-based approaches, we demonstrate that by choosing an optimal task size, the efficient work-stealing mechanisms of HPX can overcome the overhead of communication resulting in an overall 1.4 times speedup compared to the baseline MPI version.

cs.DC

An Efficient Particle Tracking Algorithm for Large-Scale Parallel Pseudo-Spectral Simulations of Turbulence

Particle tracking in large-scale numerical simulations of turbulent flows presents one of the major bottlenecks in parallel performance and scaling efficiency. Here, we describe a particle tracking algorithm for large-scale parallel pseudo-spectral simulations of turbulence which scales well up to billions of tracer particles on modern high-performance computing architectures. We summarize the standard parallel methods used to solve the fluid equations in our hybrid MPI/OpenMP implementation. As the main focus, we describe the implementation of the particle tracking algorithm and document its computational performance. To address the extensive inter-process communication required by particle tracking, we introduce a task-based approach to overlap point-to-point communications with computations, thereby enabling improved resource utilization. We characterize the computational cost as a function of the number of particles tracked and compare it with the flow field computation, showing that the cost of particle tracking is very small for typical applications.

physics.flu-dyn

Convolutional neural network-assisted recognition of nanoscale L12 ordered structures in face-centred cubic alloys

Nanoscale L12-type ordered structures are widely used in face-centred cubic (FCC) alloys to exploit their hardening capacity and thereby improve mechanical properties. These fine-scale particles are typically fully coherent with matrix with the same atomic configuration disregarding chemical species, which makes them challenging to be characterized. Spatial distribution maps (SDMs) are used to probe local order by interrogating the three-dimensional (3D) distribution of atoms within reconstructed atom probe tomography (APT) data. However, it is almost impossible to manually analyse the complete point cloud ($>10$ million) in search for the partial crystallographic information retained within the data. Here, we proposed an intelligent L12-ordered structure recognition method based on convolutional neural networks (CNNs). The SDMs of a simulated L12-ordered structure and the FCC matrix were firstly generated. These simulated images combined with a small amount of experimental data were used to train a CNN-based L12-ordered structure recognition model. Finally, the approach was successfully applied to reveal the 3D distribution of L12-type $\delta^\prime$-Al3(LiMg) nanoparticles with an average radius of 2.54 nm in a FCC Al-Li-Mg system. The minimum radius of detectable nanodomain is even down to 5 \r{A}. The proposed CNN-APT method is promising to be extended to recognize other nanoscale ordered structures and even more-challenging short-range ordered phenomena in the near future.

cond-mat.mtrl-sci

All-electron periodic $G_0W_0$ implementation with numerical atomic orbital basis functions: algorithm and benchmarks

We present an all-electron, periodic {\GnWn} implementation within the numerical atomic orbital (NAO) basis framework. A localized variant of the resolution-of-the-identity (RI) approximation is employed to significantly reduce the computational cost of evaluating and storing the two-electron Coulomb repulsion integrals. We demonstrate that the error arising from localized RI approximation can be reduced to an insignificant level by enhancing the set of auxiliary basis functions, used to expand the products of two single-particle NAOs. An efficient algorithm is introduced to deal with the Coulomb singularity in the Brillouin zone sampling that is suitable for the NAO framework. We perform systematic convergence tests and identify a set of computational parameters, which can serve as the default choice for most practical purposes. Benchmark calculations are carried out for a set of prototypical semiconductors and insulators, and compared to independent reference values obtained from an independent $G_0W_0$ implementation based on linearized augmented plane waves (LAPW) plus high-energy localized orbitals (HLOs) basis set, as well as experimental results. With a moderate (FHI-aims \textit{tier} 2) NAO basis set, our $G_0W_0$ calculations produce band gaps that typically lie in between the standard LAPW and the LAPW+HLO results. Complementing \textit{tier} 2 with highly localized Slater-type orbitals (STOs), we find that the obtained band gaps show an overall convergence towards the LAPW+HLO results. The algorithms and techniques developed in this work pave the way for efficient implementations of correlated methods within the NAO framework.

cond-mat.mtrl-sci

NECI: N-Electron Configuration Interaction with emphasis on state-of-the-art stochastic methods

We present NECI, a state-of-the-art implementation of the Full Configuration Interaction Quantum Monte Carlo algorithm, a method based on a stochastic application of the Hamiltonian matrix on a sparse sampling of the wave function. The program utilizes a very powerful parallelization and scales efficiently to more than 24000 CPU cores. In this paper, we describe the core functionalities of NECI and recent developments. This includes the capabilities to calculate ground and excited state energies, properties via the one- and two-body reduced density matrices, as well as spectral and Green's functions for ab initio and model systems. A number of enhancements of the bare FCIQMC algorithm are available within NECI, allowing to use a partially deterministic formulation of the algorithm, working in a spin-adapted basis or supporting transcorrelated Hamiltonians. NECI supports the FCIDUMP file format for integrals, supplying a convenient interface to numerous quantum chemistry programs and it is licensed under GPL-3.0.

physics.comp-ph

Octopus, a computational framework for exploring light-driven phenomena and quantum dynamics in extended and finite systems

Over the last years extraordinary advances in experimental and theoretical tools have allowed us to monitor and control matter at short time and atomic scales with a high-degree of precision. An appealing and challenging route towards engineering materials with tailored properties is to find ways to design or selectively manipulate materials, especially at the quantum level. To this end, having a state-of-the-art ab initio computer simulation tool that enables a reliable and accurate simulation of light-induced changes in the physical and chemical properties of complex systems is of utmost importance. The first principles real-space-based Octopus project was born with that idea in mind, providing an unique framework allowing to describe non-equilibrium phenomena in molecular complexes, low dimensional materials, and extended systems by accounting for electronic, ionic, and photon quantum mechanical effects within a generalized time-dependent density functional theory framework. The present article aims to present the new features that have been implemented over the last few years, including technical developments related to performance and massive parallelism. We also describe the major theoretical developments to address ultrafast light-driven processes, like the new theoretical framework of quantum electrodynamics density-functional formalism (QEDFT) for the description of novel light-matter hybrid states. Those advances, and other being released soon as part of the Octopus package, will enable the scientific community to simulate and characterize spatial and time-resolved spectroscopies, ultrafast phenomena in molecules and materials, and new emergent states of matter (QED-materials).

physics.comp-ph

Evaluation of performance portability frameworks for the implementation of a particle-in-cell code

This paper reports on an in-depth evaluation of the performance portability frameworks Kokkos and RAJA with respect to their suitability for the implementation of complex particle-in-cell (PIC) simulation codes, extending previous studies based on codes from other domains. At the example of a particle-in-cell model, we implemented the hotspot of the code in C++ and parallelized it using OpenMP, OpenACC, CUDA, Kokkos, and RAJA, targeting multi-core (CPU) and graphics (GPU) processors. Both, Kokkos and RAJA appear mature, are usable for complex codes, and keep their promise to provide performance portability across different architectures. Comparing the obtainable performance on state-of-the art hardware, but also considering aspects such as code complexity, feature availability, and overall productivity, we finally draw the conclusion that the Kokkos framework would be suited best to tackle the massively parallel implementation of the full PIC model.

cs.DC