arXiv · 2609.11131
Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting
Abstract
Human simultaneous interpreting (SI) is commonly assessed with analytic rubrics separating meaning transfer, delivery quality, and temporal synchrony, yet no automatic metric is designed for rubric-aligned segment-level SI evaluation. We construct a professionally annotated corpus of 1,101 SI segments with scores for meaning transfer (LQ), delivery quality (EXP), and perceived latency (LAT). We show that structured LLM prompting and scalar supervision collapse rubric dimensions, yielding near-zero correlation with human ratings and strong cross-dimension coupling. To isolate supervision structure under identical backbone capacity, we introduce dual regression heads on a LoRA-adapted COMET-KIWI encoder. On a held-out talk-level test set, the model achieves Pearson correlations of 0.388 (LQ) and 0.301 (EXP), improving over frozen COMET-KIWI. Given low absolute rater agreement, we interpret results relative to human consistency and target stable ranking signals for formative assessment.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ziyu Zhang, Satoshi Nakamura. 2026-09-10. Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting. https://arxiv.org/abs/2609.11131
Cite the original work for its findings. Save a collection to share your selection of sources.