arXiv · 2409.09866
S2Cap: A Benchmark and a Baseline for Singing Style Captioning
Abstract
Singing voices contain much richer information than common voices, including varied vocal and acoustic properties. However, current open-source audio-text datasets for singing voices capture only a narrow range of attributes and lack acoustic features, leading to limited utility towards downstream tasks, such as style captioning. To fill this gap, we formally define the singing style captioning task and present S2Cap, a dataset of singing voices with detailed descriptions covering diverse vocal, acoustic, and demographic characteristics. Using this dataset, we develop an efficient and straightforward baseline algorithm for singing style captioning. The dataset is available at https://zenodo.org/records/15673764.
Explore related subjects
Keep this discovery
Hyunjong Ok, Jaeho Lee. 2024-09-15. S2Cap: A Benchmark and a Baseline for Singing Style Captioning. https://arxiv.org/abs/2409.09866
Cite the original work for its findings. Save a collection to share your selection of sources.