arXiv · 2506.16976
PUL: Pre-load in Software for Caches Wouldn't Always Play Along
Abstract
Memory latencies and bandwidth are major factors, limiting system performance and scalability. Modern CPUs aim at hiding latencies by employing large caches, out-of-order execution, or complex hardware prefetchers. However, software-based prefetching exhibits higher efficiency, improving with newer CPU generations. In this paper we investigate software-based, post-Moore systems that offload operations to intelligent memories. We show that software-based prefetching has even higher potential in near-data processing settings by maximizing compute utilization through compute/IO interleaving.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Arthur Bernhardt, Sajjad Tamimi, Florian Stock, Andreas Koch, Ilia Petrov. 2025-06-20. PUL: Pre-load in Software for Caches Wouldn't Always Play Along. https://arxiv.org/abs/2506.16976
Cite the original work for its findings. Save a collection to share your selection of sources.