arXiv · 2108.10724
How Hateful are Movies? A Study and Prediction on Movie Subtitles
Abstract
In this research, we investigate techniques to detect hate speech in movies. We introduce a new dataset collected from the subtitles of six movies, where each utterance is annotated either as hate, offensive or normal. We apply transfer learning techniques of domain adaptation and fine-tuning on existing social media datasets, namely from Twitter and Fox News. We evaluate different representations, i.e., Bag of Words (BoW), Bi-directional Long short-term memory (Bi-LSTM), and Bidirectional Encoder Representations from Transformers (BERT) on 11k movie subtitles. The BERT model obtained the best macro-averaged F1-score of 77%. Hence, we show that transfer learning from the social media domain is efficacious in classifying hate and offensive speech in movies through subtitles.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Niklas von Boguszewski, Sana Moin, Anirban Bhowmick, Seid Muhie Yimam, Chris Biemann. 2021-08-19. How Hateful are Movies? A Study and Prediction on Movie Subtitles. https://arxiv.org/abs/2108.10724
Cite the original work for its findings. Save a collection to share your selection of sources.