arXiv · 2606.06386
On GPU Implementation for Multi-Precision Integer Division
Abstract
This paper presents the issues arising in implementing a fast integer division algorithm on general purpose GPUs. The algorithm uses a Newton iteration based on the shifted inverse operation, keeping all arithmetic in the integer domain and relying on data-parallel operators. The principal contribution is an efficient GPU/CUDA implementation for integer precisions from $2^{15}$ to $2^{18}$ -- sizes not supported by \cgbn{} division. We propose algorithmic refinements, define a cost model in terms of multiplications, build on prefix sums and previous work on multi-precision multiplication, and present an evaluation showing near-optimal performance relative to the model for the target precision.
Explore related subjects
Keep this discovery
Martin B. Marchioro, Aske N. Raahauge, Marc I. Løvenskjold, Cosmin E. Oancea, Stephen M. Watt. 2026-06-04. On GPU Implementation for Multi-Precision Integer Division. https://arxiv.org/abs/2606.06386
Cite the original work for its findings. Save a collection to share your selection of sources.