arXiv · 2609.10405
Frequency-Conditioned Flow Matching for Vision-Language-Action Models
Abstract
Robot actions are temporally correlated trajectories whose frequency components encode motion at different scales with highly non-uniform energy distributions. Yet Flow Matching--based vision-language-action (VLA) models typically generate actions in temporal coordinates, without explicitly modeling or systematically leveraging this frequency heterogeneity. We introduce \emph{FreqFM}, a frequency-conditioned Flow Matching framework for VLA models. It raises action frequency from an implicit trajectory property to an explicit conditioning dimension that spans the entire generation pipeline. Concretely, in DCT frequency coordinates, FreqFM constructs a spectrum-matched source distribution, adaptively balances the objective across frequencies, and constrains per-frequency guidance residuals using the corresponding reference transport scales. FreqFM integrates into existing Flow Matching action experts without changing the VLA backbone. Across LIBERO, LIBERO-Plus, and VLA-Arena, FreqFM consistently improves performance, including a 9.3-point gain on LIBERO-Plus, and further demonstrates its effectiveness on six real-robot tasks.
Explore related subjects
Keep this discovery
Haochen Niu, Shengye Dong, Hao Liu, Peiwen Lin, Wang Chuang. 2026-09-09. Frequency-Conditioned Flow Matching for Vision-Language-Action Models. https://arxiv.org/abs/2609.10405
Cite the original work for its findings. Save a collection to share your selection of sources.