arXiv · 2102.02417
Audio Adversarial Examples: Attacks Using Vocal Masks
Abstract
We construct audio adversarial examples on automatic Speech-To-Text systems . Given any audio waveform, we produce an another by overlaying an audio vocal mask generated from the original audio. We apply our audio adversarial attack to five SOTA STT systems: DeepSpeech, Julius, Kaldi, wav2letter@anywhere and CMUSphinx. In addition, we engaged human annotators to transcribe the adversarial audio. Our experiments show that these adversarial examples fool State-Of-The-Art Speech-To-Text systems, yet humans are able to consistently pick out the speech. The feasibility of this attack introduces a new domain to study machine and human perception of speech.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kai Yuan Tay, Lynnette Ng, Wei Han Chua, Lucerne Loke, Danqi Ye, Melissa Chua. 2021-02-06. Audio Adversarial Examples: Attacks Using Vocal Masks. https://arxiv.org/abs/2102.02417
Cite the original work for its findings. Save a collection to share your selection of sources.