Header

Suche

Research Blog: A Global Challenge for Machine Translation – Introducing the Last Translation Benchmark

Michelle Wastl, PhD Candidate in Computational Linguistics at the University of Zurich

Contributors from our department: Jannis Vamvas, Michelle Wastl, Rico Sennrich, Javier García Gilabert, Sina Ahmadi, Manuel Tuor, Marius Huber, Andrianos Michail, Sophia Conrad, Not Battesta Soliva, Juri Opitz 

Our department has collaborated with ETH Zurich, Johns Hopkins University, Mohamed bin Zayed University of Artificial Intelligence, and many other institutions to create the Last Translation Benchmark. The Last Translation Benchmark is a new large-scale translation evaluation project, in which difficult-to-translate examples were collected in a global community effort. Every input is human-authored and peer-reviewed by other contributors before it enters the dataset. 

What makes an example “difficult to translate” is determined by how many current translation systems get it wrong. Instead of pairing each input with a reference translation, contributors write handcrafted verification rules: short instructions specifying what a correct translation has to contain (see example below). 

The first version covers 3456 difficult-to-translate examples across 109 languages, out of which our department has contributed to English, German, Catalan, Kurdish, Swiss German, Croatian, Romansh, and Cypriot Greek.

How do current systems perform on our difficult examples?

We evaluate translation into English and German from the languages contributed by our department and vice versa, using four general-purpose commercial large language models (LLMs), the open Swiss LLM Apertus 1.5 (70B), and two translation-specific models, NLLB 3.3B and Tower+ . The translations are evaluated on the challenging and diverse subset LTBv1-eval, which excludes Romansh and Cypriot Greek due to low sample count. 

 

 

 

 

 

While the newest and largest corporate LLMs achieve a considerable pass rate when translating into English and German, the systems still fall short of the human reference and perform even weaker when having to translate into the non-English variety. For the other two commercial LLMs, DeepSeek V4 Pro and Claude Sonnet 4.5, this picture reverses, both translate better into the non-English variety. Translation-specific systems like NLLB 3.3B and Tower+ trail the commercial LLMs slightly, while the open Apertus v1.5 70B model seems to struggle with translation in general compared to other LLMs, with one exception: Swiss German. 

All systems produce fluent output and capture the surface meaning; what they miss is the specific point each item tests: a pun, a false friend, a dialect-specific word. Take the example above, where a single source item is translated into three different varieties, and the models need to disambiguate the meaning based on an emoji. The outcome depends on both the model and the target language. The strongest models can capture the intended meaning and, in the case of Swiss German, the specific dialect. However, even strong models can miss the intended distinction in particular languages.

Bear in mind that these items were designed to break machine translation, and that the benchmark includes only examples that most systems failed. The results should therefore be interpreted as a measure of performance on difficult cases, rather than as a measure of overall translation quality. 

You can still contribute to the Last Translation Benchmark

The Last Translation Benchmark is an ongoing community effort, so you can still be part of it by contributing your own difficult-to-translate examples. 

Right now, most items translate into English or German rather than out of them, so examples running the other way are especially valuable. The language pairs themselves could also be more diverse than those involving English. Switzerland’s language landscape alone is more diverse: alongside official French and Italian, the Federal Statistical Office lists Albanian and Portuguese as the most common non-official first languages after English.

There is also room for more nuanced varieties. Swiss German, Kurdish, and Romansh are oftentimes bucketed as a single variety in the benchmark, even though all three are umbrella terms covering varieties that can be mutually unintelligible. This bucketing reflects low sample counts and, in some cases, limited awareness of the finer distinctions, both of which more contributions can help fix.

Finding items that genuinely challenge current frontier models takes some creativity. Longer texts, audio, and image translation are all worth exploring. And as an added incentive, contributors with ten or more accepted submissions will be listed as authors in future versions of the report.

Contribute here: https://last-translation-benchmark.vilda.net/ 

Thanks to Jannis Vamvas, Rico Sennrich, Vilém Zouhar, Marius Huber and Javier García Gilabert for helpful input on this blog post.

Unterseiten