arXiv · 2609.34334
HOCCL: Offloading Collective Communication from GPU Cores to Accelerate Distributed Training
Abstract
Large language model training involves massive computation on GPU streaming multiprocessors (SMs), the primary compute units of GPUs. Since SMs host specialized accelerators such as Tensor Cores, their efficient utilization is critical to training efficiency. Unfortunately, existing collective communication systems compete with computation for SMs, as they consume SMs for communication-related data movement and synchronization operations. We observe that communication can, in principle, be driven by DMA engines, thereby eliminating SM involvement in communication. Based on this insight, we propose HOCCL, a zero-SM collective communication framework consisting of three components: a stream manager, a point-to-point (P2P) executor, and a collective scheduler. The stream manager preserves operator-level temporal ordering with other GPU kernels. The P2P executor enables zero-SM point-to-point communication, while the collective scheduler orchestrates P2P transfers to maximize bandwidth. Experiments show that HOCCL preserves near-peak communication performance, achieving within 3% of the state of the art on average, while eliminating communication occupancy on nearly 10% of total GPU SMs. By freeing SM resources for computation, HOCCL improves end-to-end training throughput by up to 5%.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yao Fei, Gongming Zhao, Hongli Xu, Jin Fang, Jiacheng Zhu, Shuo Xu, Kun Huang, Zhuolong Yu. 2026-09-28. HOCCL: Offloading Collective Communication from GPU Cores to Accelerate Distributed Training. https://arxiv.org/abs/2609.34334
Cite the original work for its findings. Save a collection to share your selection of sources.