AI Post-Editing at Scale: What a 71,262-Segment Study Reveals About the Future of Translation 

What is the most effective way to apply large language models to professional translation?

6 Minutes

As organizations continue integrating AI into enterprise translation workflows, one question remains at the forefront of localization strategy: What is the most effective way to apply large language models (LLMs) to professional translation? 

Some are experimenting with using LLMs to translate content directly from the source language, while others are incorporating them into existing workflows to refine machine translation output. Both approaches promise improvements in quality and efficiency, but until recently there has been little production-scale evidence comparing how they perform in real-world localization environments. 

Those questions are explored in “AI Post-Editing in Production: A 71,262-Segment Evaluation Across Five Domains, Ten Languages and Five Systems,” presented at the 2026 European Association for Machine Translation (EAMT) Conference by Mara Nunziatini and Mercedes Speroni of Welocalize. The research, which was published in Volume 2 of the proceedings, evaluates AI post-editing (AIPE) across more than 71,000 production translation segments spanning ten target languages and five content domains. Rather than relying on benchmark datasets or laboratory conditions, the study measures how different AI translation approaches perform in the complex environments that enterprise localization teams manage every day. 

The findings suggest that the greatest value does not necessarily come from replacing traditional translation workflows with LLMs. Instead, the strongest results emerged when AI was used to enhance existing translation processes with the help of enterprise knowledge, terminology, and linguistic guidance. 

Evaluating AI Under Real Production Conditions

Many published studies evaluate AI translation using relatively small datasets or a limited number of language pairs. Enterprise localization presents a far more complicated challenge. Content spans multiple domains, translation memories of varying quality, specialized terminology, and diverse business requirements. 

To reflect those realities, the researchers evaluated 71,262 production translation segments, representing more than 553,000 words translated from English into ten target languages. The dataset covered five distinct content domains, including Product and Service, Marketing, Support, and Medical Devices. In addition to automatic evaluation metrics, the researchers conducted a human assessment using 60 professional translators who reviewed more than 6,600 translation segments using the DQF-MQM quality framework and blind preference rankings. 

The evaluation compared five different systems: Production AIPE, Generic MT combined with AIPE, direct LLM translation, Google Translate, and DeepL.  

Generic MT + AIPE 


Generic MT translates the text, then an LLM (gpt-4o) post-edits it using style guides and approved past translations.  

Production AIPE


Mirrors a real production pipeline: TMs are applied first, down to 75% fuzzy matches, then MT (generic or customized) fills the rest. The LLM post-edits this combined pre-translation using the same style guides and approved past translations.

LLMT


AI Direct Translator. The same LLM translates from scratch, with no MT injection or TM leverage. It still gets the same style guides and past translations, at about 60% lower cost per segment. 

Google Translate & DeepL


Off-the-shelf generic MT, with no client-specific resources. 

This approach allowed the researchers to compare complete translation workflows rather than isolated models, providing a more realistic view of how organizations deploy AI in production

AI Post-Editing Produced the Strongest Results 

Across every automatic evaluation metric, the two AI post-editing workflows consistently outperformed the other systems. 

Production AIPE achieved the highest chrF, BLEU, and COMET scores while also requiring the fewest edits compared with professionally reviewed translations. Generic machine translation followed by AIPE performed nearly as well, while direct LLM translation produced only modest improvements over generic machine translation. 

These results highlight an important distinction. Rather than asking an LLM to generate every translation from scratch, AIPE starts with an existing translation and refines it using additional context. The workflow incorporates retrieval-augmented generation (RAG), language-specific style guides, human-reviewed bilingual examples, and enterprise terminology to improve accuracy, consistency, and brand voice before the content reaches final review. 

The findings suggest that, in production environments, the combination of machine translation and targeted AI refinement currently delivers stronger results than relying solely on direct LLM translation. 

Human Evaluation Reinforced the Findings 

Automatic evaluation metrics provide an important measure of translation quality, but professional linguists remain the ultimate benchmark for production content. 

To validate the automated findings, professional translators independently ranked the outputs without knowing which system had produced them. Their assessments closely matched the automatic metrics. 

Generic machine translation followed by AIPE received the highest percentage of first-place rankings, followed by Production AIPE. Google Translate and direct LLM translation occupied the middle tier, while DeepL ranked lowest overall. The consistency between automated measurements and expert human evaluation provides additional confidence that the observed improvements represent meaningful quality gains rather than statistical artifacts. 

For localization leaders evaluating AI investments, this agreement between human reviewers and automated scoring strengthens the case for incorporating AI post-editing into production workflows. 

Better Inputs Lead to Better Outputs 

One of the more interesting observations from the study involves the relationship between the quality of the initial translation and the effectiveness of AI post-editing. 

Although both AIPE workflows relied on the same post-editing engine, Production AIPE generated a higher proportion of major and critical errors than Generic MT combined with AIPE.  The researchers hypothesize that this difference may be due to the nature of the pre-translation inputs. 

Production AIPE processes a combination of machine translation and translation memory matches, including fuzzy matches. While fuzzy translation memory segments represented only 19 percent of the translated content, they accounted for 33 percent of all severe errors identified during evaluation. This suggests that inconsistent or lower-quality source material may limit the extent to which AI can improve the final translation. 

The finding serves as a reminder that AI is only one component of a larger localization ecosystem. The quality of translation memories, terminology resources, and other enterprise assets continues to influence overall translation performance. 

Content and Language Influence Performance 

The research also demonstrates that AI performance is not uniform across every type of content. 

Medical device translations consistently showed a higher concentration of major and critical errors than other domains, reflecting the precision required for regulated content. Marketing, product, and support content generally produced fewer severe errors across the evaluated systems. 

Similarly, although the overall ranking of systems remained consistent across all ten target languages, the severity and distribution of errors varied by locale. Some languages produced relatively few total errors but a greater concentration of major issues, reinforcing the importance of evaluating translation quality beyond simple error counts. Organizations deploying AI across multiple markets should consider both content type and language when determining where automation can be introduced most effectively. 

The Future of AI Translation Is Workflow-Oriented 

Perhaps the most important takeaway from the research is that successful AI translation depends on more than selecting the newest or most capable language model. 

The highest-performing systems integrated multiple technologies into a coordinated workflow. Machine translation, retrieval-augmented generation, translation memories, bilingual examples, style guides, and AI post-editing each contributed to the final result.  Rather than replacing existing localization infrastructure, the LLM served as an intelligent refinement layer that leveraged enterprise knowledge to improve quality and consistency. 

As organizations continue adopting AI for multilingual content, workflow design may become an even greater competitive advantage than individual model selection. The research by Nunziatini and Speroni suggests that the future of enterprise localization lies not in choosing between machine translation and large language models, but in orchestrating them effectively within a production-ready workflow. By combining established translation technologies with AI-driven refinement, organizations can achieve higher quality while maintaining the scalability and governance that enterprise localization demands. 

Research Sources

AI Post-Editing in Production: A 71,262-Segment Evaluation Across Five Domains, Ten Languages and Five Systems, by Mara Nunziatini and Mercedes Speroni, presented at the 2026 European Association for Machine Translation (EAMT) Conference.