arXiv · 1611.09524
Understanding Audio Pattern Using Convolutional Neural Network From Raw Waveforms
Abstract
One key step in audio signal processing is to transform the raw signal into representations that are efficient for encoding the original information. Traditionally, people transform the audio into spectral representations, as a function of frequency, amplitude and phase transformation. In this work, we take a purely data-driven approach to understand the temporal dynamics of audio at the raw signal level. We maximize the information extracted from the raw signal through a deep convolutional neural network (CNN) model. Our CNN model is trained on the urbansound8k dataset. We discover that salient audio patterns embedded in the raw waveforms can be efficiently extracted through a combination of nonlinear filters learned by the CNN model.
Explore related subjects
Keep this discovery
Shuhui Qu, Juncheng Li, Wei Dai, Samarjit Das. 2016-11-29. Understanding Audio Pattern Using Convolutional Neural Network From Raw Waveforms. https://arxiv.org/abs/1611.09524
Cite the original work for its findings. Save a collection to share your selection of sources.