arXiv · 1907.11065
DropAttention: A Regularization Method for Fully-Connected Self-Attention Networks
Abstract
Variants dropout methods have been designed for the fully-connected layer, convolutional layer and recurrent layer in neural networks, and shown to be effective to avoid overfitting. As an appealing alternative to recurrent and convolutional layers, the fully-connected self-attention layer surprisingly lacks a specific dropout method. This paper explores the possibility of regularizing the attention weights in Transformers to prevent different contextualized feature vectors from co-adaption. Experiments on a wide range of tasks show that DropAttention can improve performance and reduce overfitting.
Explore related subjects
Keep this discovery
Lin Zehui, Pengfei Liu, Luyao Huang, Junkun Chen, Xipeng Qiu, Xuanjing Huang. 2019-07-25. DropAttention: A Regularization Method for Fully-Connected Self-Attention Networks. https://arxiv.org/abs/1907.11065
Cite the original work for its findings. Save a collection to share your selection of sources.