arXiv · 2610.11389
Fact over Fiction: Detection of Pathological Hallucinations in Sinhala-to-English Neural Machine Translation
Abstract
Neural Machine Translation (NMT) models, while capable of producing highly fluent outputs, remain vulnerable to hallucinations, which are translations that are natural yet semantically unrelated to the source. This vulnerability is acute in low-resource settings like Sinhala-to-English, where weak cross-lingual alignment leads to hallucinations. This paper introduces a framework for reference-free hallucination detection in this language pair. We present a 45,000-sample synthetic dataset generated through a probabilistic chain of five linguistically motivated corruption strategies, with a semantic rescue mechanism that uses character-level similarity to distinguish hallucinations from morphological variants. We fine-tune mDeBERTa-v3 for token-level sequence labelling, reaching a token-level F1 of 0.841 +/- 0.001 over three seeds on a source-disjoint test set, and study a three-signal ensemble integrating neural risk scores, sequence log-probabilities, and cross-lingual semantic embeddings (LaBSE). A source-ablation control shows that the detector relies on the Sinhala source rather than on surface artefacts of the corruption process: shuffling or removing the source reduces sentence-level AUROC from 0.970 to chance. We benchmark eight NMT systems spanning five model families and find that detector firings vary by an order of magnitude across architectures.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Navam Obeysekara, Nevidu Jayatilleke. 2026-10-08. Fact over Fiction: Detection of Pathological Hallucinations in Sinhala-to-English Neural Machine Translation. https://arxiv.org/abs/2610.11389
Cite the original work for its findings. Save a collection to share your selection of sources.