arXiv · 2111.08808
User Response and Sentiment Prediction for Automatic Dialogue Evaluation
Abstract
Automatic evaluation is beneficial for open-domain dialog system development. However, standard word-overlap metrics (BLEU, ROUGE) do not correlate well with human judgements of open-domain dialog systems. In this work we propose to use the sentiment of the next user utterance for turn or dialog level evaluation. Specifically we propose three methods: one that predicts the next sentiment directly, and two others that predict the next user utterance using an utterance or a feedback generator model and then classify its sentiment. Experiments show our model outperforming existing automatic evaluation metrics on both written and spoken open-domain dialogue datasets.
Explore related subjects
Keep this discovery
Sarik Ghazarian, Behnam Hedayatnia, Alexandros Papangelis, Yang Liu, Dilek Hakkani-Tur. 2021-11-16. User Response and Sentiment Prediction for Automatic Dialogue Evaluation. https://arxiv.org/abs/2111.08808
Cite the original work for its findings. Save a collection to share your selection of sources.