arXiv ScienceSearch

arXiv · 2609.04173

Last Translation Benchmark

Vilém Zouhar·Niyati Bafna·Mukund Choudhary·Maike Züfle·Sara Rajaee·Pinzhen Chen·Jannis Vamvas·Sara Papi·Ona de Gibert·Bhavitvya Malik·Eliya Habba·Orfeas Menis Mastromichalakis·Patrícia Schmidtová·Michelle Wastl·Sheriff Issaka·Leshem Choshen·Stella Biderman·Antonis Anastasopoulos·Jan Niehues·Rico Sennrich·Mrinmaya Sachan·Ondřej Bojar·Kenton Murray·Jörg Tiedemann·Alham Fikri Aji·Philipp Koehn·Christof Monz·Alexandra Birch·Sowmya Vajjala·Chalamalasetti Kranti·Cristina España-Bonet·Nobin Sarwar·David Kaczér·Shunta Asano·Malik Marmonier·Daban Q. Jaff·Vaisakhi Mishra·Hend Al- Khalifa·Gabriele Sarti·Sourajit Saha·Nils Rehlinger·Juan Daniel Cuervo Villa·Jonathan Tonglet·Saugata Purkayastha·Dominik Macháček·Jagannathan Ramanujam·Heejin Do·Zuzana Nadova·Fred Philippy·Fabian Retkowski·Maria Lymperaiou·Silvia Casola·Hanna Yukhymenko·Shubhashis Roy Dipta·Sangwon Ryu·Andrés Jerez·Ron Keinan·Shuaib Shuaib Yusuf·Avantica Vempati·Maria Carmen Staiano·Sukannya Purkayastha·Adrian Cosma·Vitalii Babenko·Erivan Inan·Aviral Nigam·Wafa Aissa·Fatima Haouari·Venkata Prasanth Kumar Gummadi·Mehdi Jafarzadeh·Valentin Scourneau·Lukas Edman·Kaiser Sun·Shaomu Tan·Mohammad Sadegh Gholizadeh·Johannes-Rudolf David·Dipankar Srirag·Javier García Gilabert·Ruta Binkyte·Manar Ali·Ana-Maria Bucur·Sabry E. Farrag·Youssef Saber·Yihong Liu·Jean Maillard·Cojocaru Nicoleta·Xiaochuang Yuan·Sina Ahmadi·Philipp Mondorf·Kaustubh Dhole·Roman Wixinger·Shenbin Qian·Manuel Tuor·Sergey Troshin·Jonathan Yahav·Fida Mohammad Thoker·Amir Arsalan Rezapour·Lance Calvin Lim Gamboa·Manon Reusens·Kätriin Kukk·Koel Dutta Chowdhury

Abstract

For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, vulnerable to reward-hacking, and provide unactionable assessments. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking objective progress in the field and identifying pathways for improvement. We introduce the Last Translation Benchmark, a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break leading machine translation models. We also present a new evaluation approach: each example comes with handcrafted verification rules describing concrete failure cases on that example, therefore allowing reliable and actionable future evaluation. The Last Translation Benchmark is a live dataset that accepts ongoing contributions. The latest version is LTBv1, containing accepted contributions prior to September 1st 2026, with future releases planned as new data is continuously collected.

Explore related subjects

Keep this discovery

BibTeXRIS

Vilém Zouhar, Niyati Bafna, Mukund Choudhary, Maike Züfle, Sara Rajaee, Pinzhen Chen, Jannis Vamvas, Sara Papi, Ona de Gibert, Bhavitvya Malik, Eliya Habba, Orfeas Menis Mastromichalakis, Patrícia Schmidtová, Michelle Wastl, Sheriff Issaka, Leshem Choshen, Stella Biderman, Antonis Anastasopoulos, Jan Niehues, Rico Sennrich, Mrinmaya Sachan, Ondřej Bojar, Kenton Murray, Jörg Tiedemann, Alham Fikri Aji, Philipp Koehn, Christof Monz, Alexandra Birch, Sowmya Vajjala, Chalamalasetti Kranti, Cristina España-Bonet, Nobin Sarwar, David Kaczér, Shunta Asano, Malik Marmonier, Daban Q. Jaff, Vaisakhi Mishra, Hend Al- Khalifa, Gabriele Sarti, Sourajit Saha, Nils Rehlinger, Juan Daniel Cuervo Villa, Jonathan Tonglet, Saugata Purkayastha, Dominik Macháček, Jagannathan Ramanujam, Heejin Do, Zuzana Nadova, Fred Philippy, Fabian Retkowski, Maria Lymperaiou, Silvia Casola, Hanna Yukhymenko, Shubhashis Roy Dipta, Sangwon Ryu, Andrés Jerez, Ron Keinan, Shuaib Shuaib Yusuf, Avantica Vempati, Maria Carmen Staiano, Sukannya Purkayastha, Adrian Cosma, Vitalii Babenko, Erivan Inan, Aviral Nigam, Wafa Aissa, Fatima Haouari, Venkata Prasanth Kumar Gummadi, Mehdi Jafarzadeh, Valentin Scourneau, Lukas Edman, Kaiser Sun, Shaomu Tan, Mohammad Sadegh Gholizadeh, Johannes-Rudolf David, Dipankar Srirag, Javier García Gilabert, Ruta Binkyte, Manar Ali, Ana-Maria Bucur, Sabry E. Farrag, Youssef Saber, Yihong Liu, Jean Maillard, Cojocaru Nicoleta, Xiaochuang Yuan, Sina Ahmadi, Philipp Mondorf, Kaustubh Dhole, Roman Wixinger, Shenbin Qian, Manuel Tuor, Sergey Troshin, Jonathan Yahav, Fida Mohammad Thoker, Amir Arsalan Rezapour, Lance Calvin Lim Gamboa, Manon Reusens, Kätriin Kukk, Koel Dutta Chowdhury. 2026-09-03. Last Translation Benchmark. https://arxiv.org/abs/2609.04173

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

CyrillicQA: The Influence of Phonetically Encoded Secret Language on LLM Performance

Due to the selection of their training data, large language models (LLMs) perform best on standard-language inputs from languages using the Latin alphabet with large speaker populations, while disadvantaging other language varieties. Nevertheless, they can also be a versatile tool for preserving precisely such endangered languages. But do they also possess the necessary creativity and capacity for abstraction to decode phonetically encoded language the same way humans do?

cs.CL

Density Matrices for Metaphor Understanding

In physics, density matrices are used to represent mixed states, i.e. probabilistic mixtures of pure states. This concept has previously been used to model lexical ambiguity. In this paper, we consider metaphor as a type of lexical ambiguity, and examine whether metaphorical meaning can be effectively modelled using mixtures of word senses. We find that modelling metaphor is significantly more difficult than other kinds of lexical ambiguity, but that our best-performing density matrix method outperforms simple baselines as well as some neural language models.

cs.CL

Informational Antilocality and the Locality Bias in LLMs

We consider the ability of transformer-based language models (LLMs) to learn what we call k-antilocal languages, i.e., languages that have no mutual information across any span of $k$ contiguous symbols. We construct such languages with increasing $k$, finding that LLMs trained on them achieve comparable cross-entropy loss regardless of antilocality, but converge more slowly on more antilocal languages. Our findings support the idea that non-local dependencies are more difficult to learn, but the evidence for this bias comes from learning speed rather than learning success.

cs.CL