arXiv · 2610.05105
CommuteProp: Decoupled Training for Communication Bound Split LLM Fine-Tuning
Abstract
Split learning has emerged as a promising paradigm for privacy-preserving LLM fine-tuning, yet its practical deployment is severely hindered by the sequential communication-computation bottleneck. In conventional synchronous pipelines, clients remain idle while waiting for server-side gradients, resulting in substantial training inefficiency. We propose CommuteProp, an asynchronous split-learning algorithm that decouples the training process into two concurrent phases: a cross-block forward-backward pass and an in-block weight update. By overlapping computation with communication, CommuteProp reduces the marginal cycle time from a sum of all stage latencies to the dominant computational bottleneck. We provide a rigorous asynchronous error and convergence analysis. Moreover, we derive an NS preconditioner method based on our analysis to further mitigate staleness-induced noise. Comprehensive experiments indicate that our algorithm achieves substantial throughput gains while maintaining accuracy comparable to synchronous methods; furthermore, it functions as a plug-and-play module that not only accommodates but actively enhances existing privacy enhancement methods.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
CHEN Ding, LUO Haochen, LIU Chen. 2026-10-04. CommuteProp: Decoupled Training for Communication Bound Split LLM Fine-Tuning. https://arxiv.org/abs/2610.05105
Cite the original work for its findings. Save a collection to share your selection of sources.