arXiv · 2402.01298
Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations
Abstract
We propose a framework to learn semantics from raw audio signals using two types of representations, encoding contextual and phonetic information respectively. Specifically, we introduce a speech-to-unit processing pipeline that captures two types of representations with different time resolutions. For the language model, we adopt a dual-channel architecture to incorporate both types of representation. We also present new training objectives, masked context reconstruction and masked context prediction, that push models to learn semantics effectively. Experiments on the sSIMI metric of Zero Resource Speech Benchmark 2021 and Fluent Speech Command dataset show our framework learns semantics better than models trained with only one type of representation.
Explore related subjects
Keep this discovery
Jaeyeon Kim, Injune Hwang, Kyogu Lee. 2024-02-02. Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations. https://arxiv.org/abs/2402.01298
Cite the original work for its findings. Save a collection to share your selection of sources.