arXiv ScienceSearch

arXiv · 2211.13295

Techniques, Tricks and Algorithms for Efficient GPU-Based Processing of Higher Order Hyperbolic PDEs

Abstract

GPU computing is expected to play an integral part in all modern Exascale supercomputers. It is also expected that higher order Godunov schemes will make up about a significant fraction of the application mix on such supercomputers. It is, therefore, very important to prepare the community of users of higher order schemes for hyperbolic PDEs for this emerging opportunity. We focus on three broad and high-impact areas where higher order Godunov schemes are used. The first area is computational fluid dynamics (CFD). The second is computational magnetohydrodynamics (MHD) which has an involution constraint that has to be mimetically preserved. The third is computational electrodynamics (CED) which has involution constraints and also extremely stiff source terms. Together, these three diverse uses of higher order Godunov methodology, cover many of the most important applications areas. In all three cases, we show that the optimal use of algorithms, techniques and tricks, along with the use of OpenACC, yields superlative speedups on GPUs! As a bonus, we find a most remarkable and desirable result: some higher order schemes, with their larger operations count per zone, show better speedup than lower order schemes on GPUs. In other words, the GPU is an optimal stratagem for overcoming the higher computational complexities of higher order schemes! Several avenues for future improvement have also been identified. A scalability study is presented for a real-world application using GPUs and comparable numbers of high-end multicore CPUs. It is found that GPUs offer a substantial performance benefit over comparable number of CPUs, especially when all the methods designed in this paper are used.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sethupathy Subramanian, Dinshaw S. Balsara, Deepak Bhoriya, Harish Kumar. 2023-04-05. Techniques, Tricks and Algorithms for Efficient GPU-Based Processing of Higher Order Hyperbolic PDEs. https://doi.org/10.1007/s42967-022-00235-9

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The Stability of Block Eliminations and Additive Modifications

The block elimination with additive modifications (BEAM) method was recently proposed as a alternative to LU with partial pivoting requiring less communication. Because of the novelty of BEAM, the existing theoretical analysis is lacking. To that end, we analyze both the numerical stability of the underlying block LU factorization and the effects of additive modifications. For the block LU factorization, we are able to improve the previous results of Demmel et al. from being cubic in the element growth to merely quadratic. Furthermore, we propose an alternative measure of element growth that is better aligned with block LU; this new measure of growth allows our analysis to apply to matrices that cannot be factored with pointwise LU. In the second part, we analyzed the modifications produced by BEAM and the effect they have on the condition number and growth factor. Finally, we show that BEAM will not apply any modifications in some cases that regular block LU can safely factor.

math.NA

Efficient Rigorous Continuation via Chebyshev Series Expansion I

We study the global continuation of solution manifolds arising in dynamical systems. We present a rigorous continuation method based on a Chebyshev series expansion of the solution manifold. The branch is first approximated by a high-order Chebyshev interpolation polynomial, and an explicit error bound is then obtained by verifying the contraction of a quasi-Newton operator near this approximation. The contraction is formulated on a weighted $\ell^1$ space, giving a finer control than the typical $C^0$-error bound obtained from the uniform contraction theorem. In fact, the latter follows directly from our contraction operator. Furthermore, we discuss how our strategy applies naturally to pseudo-arclength continuation, where the continuation parameter fails to provide a valid local coordinate, and extends to multi-parameter continuation. Lastly, we detail two applications in which we compute a two-parameter family of steady-states for the Cahn--Hilliard equation, and a one-parameter family of steady-states undergoing saddle-node bifurcations for the Shigesada--Kawasaki--Teramoto system.

math.NA

Efficient iterative techniques for solving tensor problems with the T-product

This paper develops two efficient iterative methods for solving tensor equations under the T-product framework. For T-symmetric positive definite tensor equations of the form $\mathcal{C} \star \mathcal{X} = \mathcal{D}$, we propose a conjugate-gradient-type algorithm that generates orthogonal residual and $\mathcal{C}$-orthogonal direction sequences, ensuring convergence within a finite number of steps. For general consistent tensor equations, we extend the method using a normal-equation transformation, and further adapt it to handle inconsistent systems by solving a least-squares minimization problem. Key advantages include direct tensor-based computations without explicit matrix expansion, rigorous finite-step convergence proofs, and the ability to obtain minimal Frobenius norm solutions. Numerical experiments on synthetic data, benchmark images, and video sequences demonstrate that the proposed algorithms achieve high precision with low computational time, confirming their practicality for large-scale multidimensional problems.

math.NA