Improving Translation Quality Without Starting Over
Why the next generation of AI should build on your existing linguistic assets rather than replace them.
The Value Already Built into Enterprise Translation Workflows
For most enterprise localization programs, machine translation is not the first step in the workflow. Before any automated translation occurs, content typically moves through a pre-translation process that pulls from translation memory and term bases to create an initial draft. These assets represent years of accumulated organizational knowledge, capturing approved terminology, preferred phrasing, and language choices that have been refined over thousands of prior projects. Because of that investment, organizations expect these pre-translations to provide a strong starting point for human review and, at the very least, maintain an established quality baseline.
This existing linguistic infrastructure is often one of the most valuable assets within a mature localization program. It reflects years of decisions about terminology, brand voice, consistency standards, and reviewer preferences. Any AI system introduced into that workflow should build on those assets rather than risk undermining them.
Why Traditional Machine Translation Creates a Quality Tradeoff
Generic machine translation (MT) engines such as Google Translate and DeepL generate output by translating directly from the source text, effectively disregarding any existing translations already present within the workflow. In doing so, they overwrite linguistic decisions that may already align closely with client preferences or previously approved human translations. Rather than improving the quality of what already exists, these systems often introduce unnecessary variation and inconsistency, replacing known assets with output generated from broad multilingual training data that lacks awareness of a specific client’s historical language decisions.
Most localization teams understand this limitation, which is why established workflows frequently apply machine translation only to segments without translation memory matches while preserving pre-translated content exactly as it exists. While this protects quality, it also limits where automation can contribute value. Segments containing existing translations are effectively excluded from further optimization, forcing organizations to choose between preserving existing quality and expanding AI-driven efficiency across the broader workflow.
What the Benchmark Set Out to Measure
To better understand how different systems handle this tradeoff, the benchmark evaluated performance using ChrF delta, a metric that measures character-level similarity between system output and a human reference translation. The goal was to determine whether a system improved the quality of a segment that had already undergone pre-translation or whether additional processing actually moved the translation further away from what a human linguist would ultimately produce.
The benchmark focused specifically on content that had already passed through the pre-translation stage, allowing direct comparison between the quality of the existing segment and the quality of the output after machine processing. This made it possible to isolate an important question that most localization teams rarely test directly: can AI improve work that is already partially complete, or does additional automation simply introduce unnecessary degradation?
What the Benchmark Revealed
The results demonstrated a clear distinction between conventional machine translation systems and Opal’s optimization-based architecture. Across all ten locales and five content types included in the benchmark, Opal was the only system that consistently improved the existing pre-translation. Its output produced a positive ChrF delta, meaning the processed segment moved closer to the human reference translation than the starting version already present in the workflow.
By contrast, every other system tested, including Google Translate, DeepL, and large language model translation workflows, generated negative deltas. In those cases, the system output moved further away from the human reference than the original pre-translation itself. More notably, this pattern remained consistent across every cell in the evaluation matrix, suggesting the outcome was not isolated to specific languages or content types but reflected a fundamental difference in how the systems approached translation processing.
Why Opal Performs Differently
The difference comes down to architecture. Traditional MT engines are designed to generate entirely new target text from a source segment, treating every translation request as a standalone generation task. They do not account for whether an existing translation already exists, nor do they incorporate a client’s accumulated linguistic assets as part of the translation process. Their objective is replacement.
Opal operates differently by treating the existing translation as the foundation for improvement rather than something to discard. Instead of generating new output from scratch, the system optimizes the pre-translation using the client’s own translation memory, terminology preferences, and historical quality patterns to refine what is already there. The existing translation remains intact as the starting point, allowing the system to improve quality without sacrificing consistency.
This distinction translated into a significant performance advantage during testing. Opal outperformed the next closest system by more than 16 ChrF points, a margin that often represents the kind of leap typically associated with an entirely new generation of MT technology. What makes this result particularly notable is that the improvement came not from replacing the translation engine itself, but from introducing an optimization layer capable of improving quality while preserving the value of existing linguistic assets.
What This Changes for Enterprise Localization Programs
For enterprise localization programs, the implications extend well beyond benchmark performance. Historically, organizations have had to protect pre-translated segments from machine translation in order to preserve quality, limiting automation to only portions of the content pipeline. With an optimization-based approach, those same protected segments can now move safely through AI processing while maintaining alignment with established terminology and translation history. This expands the percentage of content that can benefit from automation while reducing the quality risks traditionally associated with direct machine translation replacement.
Production data reflects these operational gains. In a six-month enterprise deployment, one program maintained quality levels above its 99.4% target for two consecutive quarters while reducing review complexity enough for a single-linguist workflow to achieve output quality previously requiring multiple reviewers. For reviewers, the practical benefit is equally important. Content arrives closer to final quality, reducing unnecessary rewriting and allowing linguists to focus attention on the smaller percentage of segments where expert judgment creates the greatest value.
Organizations that rely heavily on translation memory and established linguistic assets should pay close attention to these findings. As AI adoption accelerates across localization, the systems that deliver the greatest value will not necessarily be those that replace existing workflows, but those capable of improving the assets organizations have already spent years building.
The full benchmark results are available in the Opal in Production report here.




Dan O’Brien
Erin Wynn
Chris Grebisz
Christy Conrad
Matt Grebisz
Siobhan Hanna
Kimberly Olson
Nicole Sheehan