arXiv · 2609.14344
AURA: Unified Multimodal Framework for Conversational Music Editing
Abstract
Instruction-guided music editors typically process each request independently, limiting their ability to support workflows in which users progressively refine a track. We introduce AURA, a unified multimodal framework for conversational music editing. AURA uses a multimodal large language model to interpret the complete dialogue history, an optional image, and reference audio, distilling the editing intent into compact concept tokens. A concept-to-audio module injects these tokens and frame-aligned reference features into a frozen MusicGen backbone, enabling precise edits while preserving unaffected content. AURA optimizes only 91M parameters while retaining 1.9B frozen backbone parameters. Experiments on Slakh2100 and MoisesDB demonstrate substantial improvements in edit correctness and content preservation over existing instruction-guided methods, including a 4-5 times reduction in FAD for out-of-domain addition and removal.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Quoc-Huy Trinh, Minh-Van Nguyen, Debesh Jha. 2026-09-13. AURA: Unified Multimodal Framework for Conversational Music Editing. https://arxiv.org/abs/2609.14344
Cite the original work for its findings. Save a collection to share your selection of sources.