arXiv · 2607.01701
Arachne: Orchestrating Cascades for Efficient Text-to-Video Model Training
Abstract
The rising demand for AI-generated videos is fueled by advances in large-scale Text-to-Video (T2V) models, trained on extensive datasets of video clips spanning diverse resolutions and durations. To address this data heterogeneity, current training methods often use a bucketing strategy that groups samples into discrete buckets for efficiency. However, this approach struggles to scale with compute and data volumes under static parallelism schemes, such as data and sequence parallelism, leading to significant workload imbalances and hardware under-utilization. In this paper, we present Arachne, a novel training framework for efficient T2V model training at scale. Arachne decomposes the training process into fine-grained computational units, called \textit{cascades}, orchestrating their distributed execution and synchronization across the cluster through coordinated spatial and temporal optimization. Our comprehensive evaluation demonstrates that Arachne reduces iteration time by up to 65\% over leading frameworks, exhibiting a positive scaling trend where its performance advantages amplify as training scale grows.
Explore related subjects
Keep this discovery
Peng Yu, Yuankai Fan, Yang Qiu, Tian Li, Bihuan Chen, Yin Chen, Qizhen Weng. 2026-07-02. Arachne: Orchestrating Cascades for Efficient Text-to-Video Model Training. https://arxiv.org/abs/2607.01701
Cite the original work for its findings. Save a collection to share your selection of sources.