arXiv · 2608.30913
Decoupled Latent Flow Matching for Few-Step Joint Vocal-Accompaniment Separation
Abstract
Generative modeling provides a flexible way to model mixture-conditioned source distributions, but iterative diffusion and flow matching models are costly for long music signals. This paper studies joint vocal-accompaniment separation through latent flow matching, where a pretrained variational autoencoder (VAE) maps mixtures and sources into a compact latent space and a flow matching model generates vocal and accompaniment latents jointly. The proposed framework decouples semantic separation from acoustic velocity prediction through a Separation Encoder and a Velocity Decoder. To reduce sampling cost, we further apply latent adversarial post-training inspired by Flow2GAN for few-step generation. Experiments show that latent adversarial refinement can improve perceptual and separation metrics under a reduced sampling budget.
Explore related subjects
Keep this discovery
Lishi Zuo, Youzhi Tu, Lu Yi, Zezhong Jin, Chongxin Gan, Man-Wai Mak, KongAik Lee. 2026-08-31. Decoupled Latent Flow Matching for Few-Step Joint Vocal-Accompaniment Separation. https://arxiv.org/abs/2608.30913
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.