DeepL AI Labs

One model to translate, one model to evaluate: How we’re building Translation Quality Evaluation for DeepL

The problem with AI translations is that they look perfect, even when they’re not. Because errors aren’t obvious at first glance, you need to take the AI model’s word that everything is correct. The only alternative is to ask language professionals to go through the work, line-by-line, hunting for errors, incomplete translations and subtle changes of meaning that a non-expert would most likely miss.

This endless requirement for checking is one of the most common complaints that localization teams have about AI translations. It raises a fundamental question for anyone engaged in AI research for language. How can machine translations earn the trust of the organizations relying on them? How can they avoid asking human reviewers to shadow the work they’ve already done, just to make sure?

Building an independent TQE AI model for DeepL

It turns out that AI can solve this problem, but it can’t do so through the same AI model that created the translation in the first place. Quality testing has shown that, when LLMs check translations, they favor the results of their own model, and miss its potential weaknesses. When it comes to AI, you can’t earn trust by asking a model to check its own homework. However, you can do so by building and training a different model that’s wholly focused on identifying any potential errors or misplaced meanings.

That’s what DeepL’s Senior Product Manager Marta Danylenko, Senior Research Scientist Martin Berger and the rest of the team behind our Translation Quality Evaluation tool have been doing. They’ve developed an independent TQE AI model with one specific purpose: to stress-test translations, and categorize any errors or potential errors that customers should know about.

How TQE works in practice

TQE’s expert AI editor goes through translations segment by segment, displaying source and translated text alongside each other. In a separate column, it highlights any potential issues with the translation, with a severity category of Critical, Major or Minor. Human reviewers are able to jump to each potential issue, review the segment, edit the translation directly, mark as resolved, or alternatively, just dismiss the flag.

This raises an obvious question, and one that Marta and Martin get to answer a lot:  If TQE knows the mistakes are there, why can’t it just go ahead and fix them? Why involve human reviewers at all? The reason comes down to the nature of the mistakes themselves. And that’s a direct result of the nature of AI.

AI earns trust best by identifying the errors it can’t fix

AI models are fundamentally probabilistic. They make highly aggregated, statistical guesses with incredible accuracy, but they’re still best guesses – and sometimes they’re wrong. When crucial context is missing, that likelihood increases. Another model can spot the divergence and the potential ambiguity, but can’t fix it without making another guess. It’s far better to get a human involved.

Sometimes, the lack of context that leads to mistaken guesses comes down to implicit knowledge: something the AI can’t know that a human translator would. Sometimes it results from the nature of the source text itself. It might contain grammatical mistakes that confuse the AI model. It might even contain deliberate grammatical errors that represent a style choice. An AI translator has to fill in the gaps, guessing at intention and meaning. A model that’s purpose-built for quality evaluation can detect this divergence from the original and flag it.

The precision/recall trade-off: How strict should a TQE model be?

For Martin, Marta and the team, the really interesting discussions started with the question of how to categorize these “errors” in translation, and how much severity to attach to each one. They started by using the long-established Multi-dimensional Quality Metrics (MQM) framework for translation quality, which focuses on areas like grammar, spelling, readability, as well as whether a translation accurately matches what’s in the source text. 

Working with language experts, though, the team quickly realized that the types of errors shift when dealing with a translation produced by AI rather than a human. There are no typos in machine translation. However, there are plenty of occasions when lack of implicit knowledge is exposed.

That’s why the goal of TQE has never been to replace human reviewers. By combining a translating model with an evaluating one, we’re ensuring that AI is able to detect its own limitations and loop in those reviewers wherever they’re needed, saving them crucial time. They remain an essential part of the process. The goal is to enable them to work with AI in a more satisfying way.

This requires a balance: between flagging every type of potential error, but creating more work for reviewers, and issuing alerts only for those that have serious implications. That’s why the discussions around what constitutes a Critical, Major or Minor translation error were some of the most interesting in the development of TQE. 

Working with customers to build the best version of TQE

Calibrating how strict TQE should be is one of many areas in which we’re working with customers on refining the model. We’re integrating DeepL’s customization features like glossaries and style rules to give the model more contextual knowledge. We’ve made TQE available as an API, and we’re also adding suggestions for how to fix errors, and options to bulk accept and reject them as well.

Many of these developments come from insights shared by customers working with our TQE testing tool. If you’d like to be part of developing this next step in AI translation, then please reach out to request access to the testing environment. We’d love to hear how TQE works for you.

Interested in gaining access to our TQE testing environment? Email marta.danylenko@deepl.com.

Dela

Håll kontakten

Få en förhandsvisning av våra senaste AI-innovationer.