arXiv · 2108.10069
An Interpretable Approach to Hateful Meme Detection
Abstract
Hateful memes are an emerging method of spreading hate on the internet, relying on both images and text to convey a hateful message. We take an interpretable approach to hateful meme detection, using machine learning and simple heuristics to identify the features most important to classifying a meme as hateful. In the process, we build a gradient-boosted decision tree and an LSTM-based model that achieve comparable performance (73.8 validation and 72.7 test auROC) to the gold standard of humans and state-of-the-art transformer models on this challenging task.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Tanvi Deshpande, Nitya Mani. 2021-08-09. An Interpretable Approach to Hateful Meme Detection. https://arxiv.org/abs/2108.10069
Cite the original work for its findings. Save a collection to share your selection of sources.