arXiv · 2609.22131
Correlation-Aware Structured Pruning for Large Language Models
Abstract
Structured pruning is a promising approach for reducing the substantial inference costs of Large Language Models (LLMs) while maintaining hardware efficiency. Many existing methods assess the importance of prunable units (e.g., channels or heads) in isolation, implicitly assuming that pruning errors are additive. This independence assumption is often invalidated by the non-orthogonality of model weights and strong correlations between unit activations, potentially leading to performance degradation. To address this, we propose a Correlation-Aware Structured Pruning method. We formulate the pruning objective as a cardinality-constrained binary quadratic program that explicitly models cross-unit dependencies in the reconstruction error. Since this binary quadratic program is NP-hard and difficult to solve exactly, we develop a greedy interaction algorithm based on dependency-aware marginal costs to optimize unit selection. Furthermore, we incorporate a gradient-based strategy to achieve adaptive layer-wise sparsity allocation across the entire model. Extensive experiments on mainstream LLMs demonstrate that incorporating correlation information yields competitive accuracy-efficiency trade-offs compared to representative structured pruning baselines.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sicheng Xu, Hao Shi, Wei Zhang, Haoran Pang, Zhenyu Ming, Hao Wu, Zhongyi Huang, Xin Yao, Gong Zhang. 2026-08-23. Correlation-Aware Structured Pruning for Large Language Models. https://arxiv.org/abs/2609.22131
Cite the original work for its findings. Save a collection to share your selection of sources.