arXiv · 2010.12114
The nanoPU: Redesigning the CPU-Network Interface to Minimize RPC Tail Latency
Abstract
The nanoPU is a new networking-optimized CPU designed to minimize tail latency for RPCs. By bypassing the cache and memory hierarchy, the nanoPU directly places arriving messages into the CPU register file. The wire-to-wire latency through the application is just 65ns, about 13x faster than the current state-of-the-art. The nanoPU moves key functions from software to hardware: reliable network transport, congestion control, core selection, and thread scheduling. It also supports a unique feature to bound the tail latency experienced by high-priority applications. Our prototype nanoPU is based on a modified RISC-V CPU; we evaluate its performance using cycle-accurate simulations of 324 cores on AWS FPGAs, including real applications (MICA and chain replication).
Explore related subjects
Keep this discovery
Stephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen, Muhammad Shahbaz, Nick McKeown, Changhoon Kim. 2020-10-23. The nanoPU: Redesigning the CPU-Network Interface to Minimize RPC Tail Latency. https://arxiv.org/abs/2010.12114
Cite the original work for its findings. Save a collection to share your selection of sources.