arXiv · 2607.11081
Controlling Motion Transfer in Diffusion Transformers via Attention Heads
Abstract
Diffusion Transformers (DiTs) have advanced video generation with high-quality, temporally coherent results. However, extending them to motion transfer, which requires following reference motion while aligning with a target prompt, remains challenging due to limited understanding of motion and structure representations within DiTs. We analyze video DiTs at the attention-head level and identify distinct heads specialized for motion and spatial structure. Based on this insight, we propose a head-aware controllable motion transfer framework that requires no parameter updates. Our method refines motion cues from motion-specialized heads via semantic correspondence guidance and preserves structure through selective feature injection. This head-level control not only enables accurate motion transfer but also provides an interpretable foundation for controllable video generation with DiTs.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sunyoung Jung, Jiwoo Park, Yoonseok Choi, Kyobin Choo, Ming-Hsuan Yang, Seong Jae Hwang. 2026-07-13. Controlling Motion Transfer in Diffusion Transformers via Attention Heads. https://arxiv.org/abs/2607.11081
Cite the original work for its findings. Save a collection to share your selection of sources.