arXiv · 2503.16672
Accelerating Transformer Inference and Training with 2:4 Activation Sparsity
Abstract
In this paper, we demonstrate how to leverage 2:4 sparsity, a popular hardware-accelerated GPU sparsity pattern, to activations to accelerate large language model training and inference. Crucially we exploit the intrinsic sparsity found in Squared-ReLU activations to provide this acceleration with no accuracy loss. Our approach achieves up to 1.3x faster Feed Forward Network (FFNs) in both the forwards and backwards pass. This work highlights the potential for sparsity to play a key role in accelerating large language model training and inference.
Explore related subjects
Keep this discovery
Daniel Haziza, Timothy Chou, Dhruv Choudhary, Luca Wehrstedt, Francisco Massa, Jiecao Yu, Geonhwa Jeong, Supriya Rao, Patrick Labatut, Jesse Cai. 2025-03-20. Accelerating Transformer Inference and Training with 2:4 Activation Sparsity. https://arxiv.org/abs/2503.16672
Cite the original work for its findings. Save a collection to share your selection of sources.