arXiv · 2103.00676
Token-Modification Adversarial Attacks for Natural Language Processing: A Survey
Abstract
Many adversarial attacks target natural language processing systems, most of which succeed through modifying the individual tokens of a document. Despite the apparent uniqueness of each of these attacks, fundamentally they are simply a distinct configuration of four components: a goal function, allowable transformations, a search method, and constraints. In this survey, we systematically present the different components used throughout the literature, using an attack-independent framework which allows for easy comparison and categorisation of components. Our work aims to serve as a comprehensive guide for newcomers to the field and to spark targeted research into refining the individual attack components.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Tom Roth, Yansong Gao, Alsharif Abuadbba, Surya Nepal, Wei Liu. 2021-03-01. Token-Modification Adversarial Attacks for Natural Language Processing: A Survey. https://arxiv.org/abs/2103.00676
Cite the original work for its findings. Save a collection to share your selection of sources.