arXiv · 2506.07136
MotionStrata: Hierarchical Motion Latents for Compact Video Autoencoding
Abstract
First-frame-conditioned video autoencoders represent a clip with persistent content and a compact motion code. Although this removes much of the appearance redundancy, the remaining motion is typically compressed with a homogeneous latent geometry. Such representations use the same temporal support for broad scene evolution and fine-grained, frame-specific details. We introduce MotionStrata, which organizes a fixed motion budget into Global Motion and Detailed Motion. Temporally compressed Global queries summarize broad evolution, whereas frame-aligned Detailed queries preserve fine-grained structures whose configuration varies across frames. Frequency-guided routing and coarse-to-fine training establish this hierarchy without increasing motion dimensionality. Experiments show that MotionStrata maintains high reconstruction quality under aggressive compression and outperforms uniform and alternative grouped representations. Additional experiments evaluate hierarchical representation, downstream generation, and decoding cost. These results support hierarchical motion organization as a useful design principle for compact video autoencoding.
Explore related subjects
Keep this discovery
Wenzhang Sun, Huaize Liu, Chunfeng Wang, Biao Gong, Hao Li, Changqing Zou. 2025-06-08. MotionStrata: Hierarchical Motion Latents for Compact Video Autoencoding. https://arxiv.org/abs/2506.07136
Cite the original work for its findings. Save a collection to share your selection of sources.