arXiv ScienceSearch

arXiv subjects

Peter Munch

Publications and source records attributed to Peter Munch.

At least 19 recordsLinked to original sources

Solving the (Navier-)Stokes equations with space and time adaptivity using deal.II

In this article, we solve the Stokes and Navier-Stokes equations with the deal$.$II finite-element library. In particular, we use its multigrid, adaptive-mesh, and matrix-free infrastructures to design efficient linear and nonlinear iterative solvers, respectively. We solve the stationary Stokes equations on hp-adaptive meshes with a hp-multigrid approach, the transient Stokes equations with space-time finite elements and space-time multigrid, and, finally, the stabilized incompressible Navier-Stokes equations on locally refined meshes with a monolithic multigrid solver. The selected examples underline the flexibility and modularity of the multigrid infrastructure of deal$.$II.

math.NA

Eigenvalue-based Linear Stability Analysis of Intrinsic Instabilities in Laminar Flames

Intrinsic instabilities of laminar premixed flames play an important role in the dynamics of hydrogen combustion and in the development of predictive models for reacting flows. However, determining their dispersion relations typically relies either on simplified analytical descriptions of the flame front or on computationally expensive direct numerical simulations (DNS). This work develops a generalized eigenvalue problem-based linear stability analysis (GEVP-LSA) framework that predicts the growth rates and spatial structure of intrinsic flame instabilities directly from the linearized governing equations of a 1D base flame. The approach is first validated using the classical Darrieus-Landau configuration, where the numerical results reproduce the analytical dispersion relation and eigenmode structure. The framework is then applied to a model flame of finite thickness governed by the reactive Navier-Stokes equations. The resulting dispersion relations and perturbation fields show excellent agreement with corresponding DNS results while reducing the computational effort by a factor of 1e8. The proposed method therefore provides an efficient and accurate tool for studying intrinsic flame instabilities and offers a scalable foundation for future stability analyses of more complex reacting-flow configurations relevant to combustion modeling and large-eddy simulations.

physics.flu-dyn

Revisiting the Slip Boundary Condition: Surface Roughness as a Hidden Tuning Parameter

In this paper, we investigate the effect of boundary surface roughness on numerical simulations of incompressible fluid flow past a cylinder in two and three spatial dimensions furnished with slip boundary conditions. The governing equations are approximated using a continuous finite element method, stabilized with a Galerkin least-squares approach. Through a series of numerical experiments, we demonstrate that: $(i)$ the introduction of surface roughness through numerical discretization error, or mesh distortion, makes the potential flow solution unstable; $(ii)$ when numerical surface roughness and mesh distortion are minimized by using high-order isoparametric geometry mappings, a stable potential flow is obtained in both two and three dimensions; $(iii)$ numerical surface roughness, mesh distortion and refinement level can be used as control parameters to manipulate drag and lift forces resulting in numerical values spanning more than an order of magnitude. Our results cast some doubt on the predictive capability of the slip boundary condition for wall modeling in turbulent simulations of incompressible flow.

physics.flu-dyn

An hp Multigrid Approach for Tensor-Product Space-Time Finite Element Discretizations of the Stokes Equations

We present a monolithic $hp$ space-time multigrid method for tensor-product space-time finite element discretizations of the Stokes equations. Geometric and polynomial coarsening of the space-time mesh is performed, and the entire algorithm is expressed through rigorous mathematical mappings. For the discretization, we use inf-sup stable pairs $\mathbb Q_{r+1}/\mathbb P_{r}^{\text{disc}}$ of elements in space and a discontinuous Galerkin (DG$(k)$) discretization in time with piecewise polynomials of order $k$. The key novelty of this work is the application of $hp$ multigrid techniques in space and time, facilitated and accelerated by the matrix-free capabilities of the deal$.$II library. While multigrid methods are well-established for stationary problems, their application in space-time formulations encounter unique challenges, particularly in constructing suitable smoothers. To overcome these challenges, we employ space-time cell and vertex star patch based Vanka smoothers. Extensive tests on high-performance computing platforms demonstrate the efficiency of our \( hp \) multigrid approach on problem sizes exceeding a trillion degrees of freedom (dofs), sustaining throughputs of hundreds of millions of dofs per second.

math.NA

A consistent diffuse-interface finite element approach to rapid melt--vapor dynamics with application to metal additive manufacturing

Metal additive manufacturing via laser-based powder bed fusion (PBF-LB/M) faces performance-critical challenges due to complex melt pool and vapor dynamics, often oversimplified by computational models that neglect crucial aspects, such as vapor jet formation. To address this limitation, we propose a consistent computational multi-physics mesoscale model to study melt pool dynamics, laser-induced evaporation, and vapor flow. In addition to the evaporation-induced pressure jump, we also resolve the evaporation-induced volume expansion and the resulting velocity jump at the liquid--vapor interface. We use an anisothermal incompressible Navier--Stokes solver extended by a conservative diffuse level-set framework and integrate it into a matrix-free adaptive finite element framework. To ensure accurate physical solutions despite extreme density, pressure and velocity gradients across the diffuse liquid--vapor interface, we employ consistent interface source term formulations developed in our previous work. These formulations consider projection operations to extend solution variables from the sharp liquid--vapor interface into the computational domain. Benchmark examples, including film boiling, confirm the accuracy and versatility of the model. As a key result, we demonstrate the model's ability to capture the strong coupling between melt and vapor flow dynamics in PBF-LB/M based on simulations of stationary laser illumination on a metal plate. Additionally, we show the derivation of the well-known Anisimov model and extend it to a new hybrid model. This hybrid model, together with consistent interface source term formulations, especially for the level-set transport velocity, enables PBF-LB/M simulations that combine accurate physical results with the robustness of an incompressible, diffuse-interface computational modeling framework.

cs.CE

Super-Localized Orthogonal Decomposition Method for Heterogeneous Linear Elasticity

We present the Super-Localized Orthogonal Decomposition (SLOD) method for the numerical homogenization of linear elasticity problems with multiscale microstructures modeled by a heterogeneous coefficient field without any periodicity or scale separation assumptions. Compared to the established Localized Orthogonal Decomposition (LOD) and its linear localization approach, SLOD achieves significantly improved sparsity properties through a nonlinear superlocalization technique, leading to computationally efficient solutions with significantly less oversampling - without compromising accuracy. We generalize the method to vector-valued problems and provide a supporting numerical analysis. We also present a scalable implementation of SLOD using the deal.II finite element library, demonstrating its feasibility for high-performance simulations. Numerical experiments illustrate the efficiency and accuracy of SLOD in addressing key computational challenges in multiscale elasticity.

math.NA

Matrix-free implementation of the non-nested multigrid method

Traditionally, the geometric multigrid method is used with nested levels. However, the construction of a suitable hierarchy for very fine and unstructured grids is, in general, highly non-trivial. In this scenario, the non-nested multigrid method could be exploited in order to handle the burden of hierarchy generation, allowing some flexibility on the choice of the levels. We present a parallel, matrix-free, implementation of the non-nested multigrid method for continuous Lagrange finite elements, where each level may consist of independently partitioned triangulations. Our algorithm has been added to the multigrid framework of the C++ finite-element library deal.II. Several 2D and 3D numerical experiments are presented, ranging from Poisson problems to linear elasticity. We test the robustness and performance of the proposed implementation with different polynomial degrees and geometries.

math.NA

Matrix-Free Higher-Order Finite Element Methods for Hyperelasticity

This work presents a matrix-free finite element solver for finite-strain elasticity adopting an $hp$-multigrid preconditioner. Compared to classical algorithms relying on a global sparse matrix, matrix-free solution strategies significantly reduce memory traffic by repeated evaluation of the finite element integrals. Following this approach in the context of finite-strain elasticity, the precise statement of the final weak form is crucial for performance, and it is not clear a priori whether to choose problem formulations in the material or spatial domain. With a focus on hyperelastic solids in biomechanics, the arithmetic costs to evaluate the material law at each quadrature point might favor an evaluation strategy where some quantities are precomputed in each Newton iteration and reused in the Krylov solver for the linearized problem. Hence, we discuss storage strategies to balance the compute load against memory access in compressible and incompressible neo-Hookean models and an anisotropic tissue model. Additionally, numerical stability becomes increasingly important using lower/mixed-precision ingredients and approximate preconditioners to better utilize modern hardware architectures. Application of the presented method to a patient-specific geometry of an iliac bifurcation shows significant speed-ups, especially for higher polynomial degrees, when compared to alternative approaches with matrix-based geometric or black-box algebraic multigrid preconditioners.

cs.CE

Fairness measures for biometric quality assessment

Quality assessment algorithms measure the quality of a captured biometric sample. Since the sample quality strongly affects the recognition performance of a biometric system, it is essential to only process samples of sufficient quality and discard samples of low-quality. Even though quality assessment algorithms are not intended to yield very different quality scores across demographic groups, quality score discrepancies are possible, resulting in different discard ratios. To ensure that quality assessment algorithms do not take demographic characteristics into account when assessing sample quality and consequently to ensure that the quality algorithms perform equally for all individuals, it is crucial to develop a fairness measure. In this work we propose and compare multiple fairness measures for evaluating quality components across demographic groups. Proposed measures, could be used as potential candidates for an upcoming standard in this important field.

cs.CV

A Space-Time Multigrid Method for Space-Time Finite Element Discretizations of Parabolic and Hyperbolic PDEs

We present a space-time multigrid method based on tensor-product space-time finite element discretizations. The method is facilitated by the matrix-free capabilities of the {\ttfamily deal.II} library. It addresses both high-order continuous and discontinuous variational time discretizations with spatial finite element discretizations. The effectiveness of multigrid methods in large-scale stationary problems is well established. However, their application in the space-time context poses significant challenges, mainly due to the construction of suitable smoothers. To address these challenges, we develop a space-time cell-wise additive Schwarz smoother and demonstrate its effectiveness on the heat and acoustic wave equations. The matrix-free framework of the {\ttfamily deal.II} library supports various multigrid strategies, including $h$-, $p$-, and $hp$-refinement across spatial and temporal dimensions. Extensive empirical evidence, provided through scaling and convergence tests on high-performance computing platforms, demonstrate high performance on perturbed meshes and problems with heterogeneous and discontinuous coefficients. Throughputs of over a billion degrees of freedom per second are achieved on problems with more than a trillion global degrees of freedom. The results prove that the space-time multigrid method can effectively solve complex problems in high-fidelity simulations and show great potential for use in coupled problems.

math.NA

Smoothers with localized residual computations for geometric multigrid methods

We improve the performance of multigrid solvers on many-core architectures with cache hierarchies by reorganizing operations in the smoothing step to minimize memory transfers. We focus on patch smoothers, which offer robust convergence rates with respect to the finite element degree for various equations, in the setting of multiplicative subspace correction for numerical efficiency. By combining the computation of local residuals with local solvers, we increase the locality of the problem and thus reduce data transfers. The thread-parallel implementation of this algorithm is based on coloring, which contradicts cache efficiency. We improve data locality by rearranging the loop into batches so that more data can be reused. The organization of consecutive batches prioritizes data locality.

math.NA

High-performance matrix-free unfitted finite element operator evaluation

Unfitted finite element methods, like CutFEM, have traditionally been implemented in a matrix-based fashion, where a sparse matrix is assembled and later applied to vectors while solving the resulting linear system. With the goal of increasing performance and enabling algorithms with polynomial spaces of higher degrees, this contribution chooses a more abstract approach by matrix-free evaluation of the operator action on vectors instead. The proposed method loops over cells and locally evaluates the cell, face, and interface integrals, including the contributions from cut cells and the different means of stabilization. The main challenge is the efficient numerical evaluation of terms in the weak form with unstructured quadrature points arising from the unfitted discretization in cells cut by the interface. We present design choices and performance optimizations for tensor-product elements and demonstrate the performance by means of benchmarks and application examples. We demonstrate a speedup of more than one order of magnitude for the operator evaluation of a discontinuous Galerkin discretization with polynomial degree three compared to a sparse matrix-vector product and develop performance models to quantify the performance properties over a wide range of polynomial degrees.

math.NA

Improved accuracy of continuum surface flux models for metal additive manufacturing melt pool simulations

Computational modeling of the melt pool dynamics in laser-based powder bed fusion metal additive manufacturing (PBF-LB/M) promises to shed light on fundamental mechanisms of defect generation. These processes are accompanied by rapid evaporation so that the evaporation-induced recoil pressure and cooling arise as major driving forces for fluid dynamics and temperature evolution. The magnitude of these interface fluxes depends exponentially on the melt pool surface temperature, which, therefore, has to be predicted with high accuracy. The present work utilizes a diffuse interface finite element model based on a continuum surface flux (CSF) description of interface fluxes to study dimensionally reduced thermal two-phase problems representative for PBF-LB/M in a finite element framework. It is demonstrated that the extreme temperature gradients combined with the high ratios of material properties between metal and ambient gas lead to significant errors in the interface temperatures and fluxes when classical CSF approaches, along with typical interface thicknesses and discretizations, are applied. It is expected that this finding is also relevant for other types of diffuse interface PBF-LB/M melt pool models. A novel parameter-scaled CSF approach is proposed, which is constructed to yield a smoother temperature field in the diffuse interface region, significantly increasing the solution accuracy. The interface thickness required to predict the temperature field with a given level of accuracy is less restrictive by at least one order of magnitude for the proposed parameter-scaled approach compared to classical CSF, drastically reducing computational costs. Finally, we showcase the general applicability of the parameter-scaled CSF to a 3D simulation of stationary laser melting of PBF-LB/M considering the fully coupled thermo-hydrodynamic multi-phase problem, including phase change.

cs.CE

A consistent diffuse-interface model for two-phase flow problems with rapid evaporation

We present accurate and mathematically consistent formulations of a diffuse-interface model for two-phase flow problems involving rapid evaporation. The model addresses challenges including discontinuities in the density field by several orders of magnitude, leading to high velocity and pressure jumps across the liquid-vapor interface, along with dynamically changing interface topologies. To this end, we integrate an incompressible Navier-Stokes solver combined with a conservative level-set formulation and a regularized, i.e., diffuse, representation of discontinuities into a matrix-free adaptive finite element framework. The achievements are three-fold: First, we propose mathematically consistent definitions for the level-set transport velocity in the diffuse interface region by extrapolating the velocity from the liquid or gas phase. They exhibit superior prediction accuracy for the evaporated mass and the resulting interface dynamics compared to a local velocity evaluation, especially for strongly curved interfaces. Second, we show that accurate prediction of the evaporation-induced pressure jump requires a consistent, namely a reciprocal, density interpolation across the interface, which satisfies local mass conservation. Third, the combination of diffuse interface models for evaporation with standard Stokes-type constitutive relations for viscous flows leads to significant pressure artifacts in the diffuse interface region. To mitigate these, we propose to introduce a correction term for such constitutive model types. Through selected analytical and numerical examples, the aforementioned properties are validated. The presented model promises new insights in simulation-based prediction of melt-vapor interactions in thermal multiphase flows such as in laser-based powder bed fusion of metals.

cs.CE

A highly efficient computational framework for fast scan-resolved simulations of metal additive manufacturing processes on the scale of real parts

This article proposes a novel high-performance computing approach for the prediction of the temperature field in powder bed fusion (PBF) additive manufacturing processes. In contrast to many existing approaches to part-scale simulations, the underlying computational model consistently resolves physical scan tracks without additional heat source scaling, agglomeration strategies or any other heuristic modeling assumptions. A growing, adaptively refined mesh accurately captures all details of the laser beam motion. Critically, the fine spatial resolution required for resolved scan tracks in combination with the high scan velocities underlying these processes mandates the use of comparatively small time steps to resolve the underlying physics. Explicit time integration schemes are well-suited for this setting, while unconditionally stable implicit time integration schemes are employed for the interlayer cool down phase governed by significantly larger time scales. These two schemes are combined and implemented in an efficient fast operator evaluation framework providing significant performance gains and optimization opportunities. The capabilities of the novel framework are demonstrated through realistic AM examples on the centimeter scale including the first scan-resolved simulation of the entire NIST AM Benchmark cantilever specimen, with a computation time of less than one day. Apart from physical insights gained through these simulation examples, also numerical aspects are thoroughly studied on basis of weak and strong parallel scaling tests. As potential applications, the proposed thermal PBF simulation framework can serve as a basis for microstructure and thermo-mechanical predictions on the part-scale, but also to assess the influence of scan pattern and part geometry on melt pool shape and temperature, which are important indicators for well-known process instabilities.

cs.CE

High-Order Non-Conforming Discontinuous Galerkin Methods for the Acoustic Conservation Equations

This work compares two Nitsche-type approaches to treat non-conforming triangulations for a high-order discontinuous Galerkin (DG) solver for the acoustic conservation equations. The first approach (point-to-point interpolation) uses inexact integration with quadrature points prescribed by a primary element. The second approach uses exact integration (mortaring) by choosing quadratures depending on the intersection between non-conforming elements. In literature, some excellent properties regarding performance and ease of implementation are reported for point-to-point interpolation. However, we show that this approach can not safely be used for DG discretizations of the acoustic conservation equations since, in our setting, it yields spurious oscillations that lead to instabilities. This work presents a test case in that we can observe the instabilities and shows that exact integration is required to maintain a stable method. Additionally, we provide a detailed analysis of the method with exact integration. We show optimal spatial convergence rates globally and in each mesh region separately. The method is constructed such that it can natively treat overlaps between elements. Finally, we highlight the benefits of non-conforming discretizations in acoustic computations by a numerical test case with different fluids.

math.NA

Stage-parallel fully implicit Runge-Kutta implementations with optimal multilevel preconditioners at the scaling limit

We present an implementation of a fully stage-parallel preconditioner for Radau IIA type fully implicit Runge--Kutta methods, which approximates the inverse of $A_Q$ from the Butcher tableau by the lower triangular matrix resulting from an LU decomposition and diagonalizes the system with as many blocks as stages. For the transformed system, we employ a block preconditioner where each block is distributed and solved by a subgroup of processes in parallel. For combination of partial results, we either use a communication pattern resembling Cannon's algorithm or shared memory. A performance model and a large set of performance studies (including strong scaling runs with up to 150k processes on 3k compute nodes) conducted for a time-dependent heat problem, using matrix-free finite element methods, indicate that the stage-parallel implementation can reach higher throughputs when the block solvers operate at lower parallel efficiencies, which occurs near the scaling limit. Achievable speedup increases linearly with number of stages and are bounded by the number of stages. Furthermore, we show that the presented stage-parallel concepts are also applicable to the case that $A_Q$ is directly diagonalized, which requires complex arithmetic or the solution of two-by-two blocks and sequentializes parts of the algorithm. Alternatively to distributing stages and assigning them to distinct processes, we discuss the possibility of batching operations from different stages together.

math.NA

Enhancing data locality of the conjugate gradient method for high-order matrix-free finite-element implementations

This work investigates a variant of the conjugate gradient (CG) method and embeds it into the context of high-order finite-element schemes with fast matrix-free operator evaluation and cheap preconditioners like the matrix diagonal. Relying on a data-dependency analysis and appropriate enumeration of degrees of freedom, we interleave the vector updates and inner products in a CG iteration with the matrix-vector product with only minor organizational overhead. As a result, around 90% of the vector entries of the three active vectors of the CG method are transferred from slow RAM memory exactly once per iteration, with all additional access hitting fast cache memory. Node-level performance analyses and scaling studies on up to 147k cores show that the CG method with the proposed performance optimizations is around two times faster than a standard CG solver as well as optimized pipelined CG and s-step CG methods for large sizes that exceed processor caches, and provides similar performance near the strong scaling limit.

cs.MS