arXiv · 2507.06590
MOST: Motion Diffusion Model for Rare Text via Temporal Clip Banzhaf Interaction
Abstract
We introduce MOST, a novel motion diffusion model via temporal clip Banzhaf interaction, aimed at addressing the persistent challenge of generating human motion from rare language prompts. While previous approaches struggle with coarse-grained matching and overlook important semantic cues due to motion redundancy, our key insight lies in leveraging fine-grained clip relationships to mitigate these issues. MOST's retrieval stage presents the first formulation of its kind - temporal clip Banzhaf interaction - which precisely quantifies textual-motion coherence at the clip level. This facilitates direct, fine-grained text-to-motion clip matching and eliminates prevalent redundancy. In the generation stage, a motion prompt module effectively utilizes retrieved motion clips to produce semantically consistent movements. Extensive evaluations confirm that MOST achieves state-of-the-art text-to-motion retrieval and generation performance by comprehensively addressing previous challenges, as demonstrated through quantitative and qualitative results highlighting its effectiveness, especially for rare prompts.
Explore related subjects
Keep this discovery
Yin Wang, Mu li, Zhiying Leng, Frederick W. B. Li, Xiaohui Liang. 2025-07-09. MOST: Motion Diffusion Model for Rare Text via Temporal Clip Banzhaf Interaction. https://arxiv.org/abs/2507.06590
Cite the original work for its findings. Save a collection to share your selection of sources.