arXiv · 2405.11425
Enabling full-speed random access to the entire memory on the A100 GPU
Abstract
We describe some features of the A100 memory architecture. In particular, we give a technique to reverse-engineer some hardware layout information. Using this information, we show how to avoid TLB issues to obtain full-speed random HBM access to the entire memory, as long as we constrain any particular thread to a reduced access window of less than 64GB.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Alden Walker. 2024-05-19. Enabling full-speed random access to the entire memory on the A100 GPU. https://arxiv.org/abs/2405.11425
Cite the original work for its findings. Save a collection to share your selection of sources.