AAAI 2020
Multi-Scale Self-Attention for Text Classification
Abstract
In this paper, we introduce the prior knowledge, multi-scale structure, into self-attention modules. We propose a Multi- Scale Transformer which uses multi-scale multi-head selfattention to capture features from different scales. Based on the linguistic perspective and the analysis of pre-trained Transformer (BERT) on a huge corpus, we further design a strategy to control the scale distribution for each layer. Results of three different kinds of tasks (21 datasets) show our Multi-Scale Transformer outperforms the standard Transformer consistently and significantly on small and moderate size datasets.
Authors
Keywords
No keywords are indexed for this paper.
Context
- Venue
- AAAI Conference on Artificial Intelligence
- Archive span
- 1980-2026
- Indexed papers
- 28718
- Paper id
- 249041277437086547