arXiv · 1808.04164
Automatic Reference-Based Evaluation of Pronoun Translation Misses the Point
Abstract
We compare the performance of the APT and AutoPRF metrics for pronoun translation against a manually annotated dataset comprising human judgements as to the correctness of translations of the PROTEST test suite. Although there is some correlation with the human judgements, a range of issues limit the performance of the automated metrics. Instead, we recommend the use of semi-automatic metrics and test suites in place of fully automatic metrics.
Explore related subjects
Keep this discovery
Liane Guillou, Christian Hardmeier. 2018-08-13. Automatic Reference-Based Evaluation of Pronoun Translation Misses the Point. https://arxiv.org/abs/1808.04164
Cite the original work for its findings. Save a collection to share your selection of sources.