arXiv · 2010.12776
Improved Synthetic Training for Reading Comprehension
Abstract
Automatically generated synthetic training examples have been shown to improve performance in machine reading comprehension (MRC). Compared to human annotated gold standard data, synthetic training data has unique properties, such as high availability at the possible expense of quality. In view of such differences, in this paper, we explore novel applications of synthetic examples to MRC. Our proposed pre-training and knowledge distillation strategies show significant improvements over existing methods. In a particularly surprising discovery, we observe that synthetic distillation often yields students that can outperform the teacher model.
Explore related subjects
Keep this discovery
Yanda Chen, Md Arafat Sultan, Vittorio Castelli. 2020-10-24. Improved Synthetic Training for Reading Comprehension. https://arxiv.org/abs/2010.12776
Cite the original work for its findings. Save a collection to share your selection of sources.