+16 Points Closer: What ChrF Means for Your Workflow

Why translation quality metrics can reveal hidden efficiency gains long before they appear in production workflows. 

3 Minutes

If you manage a localization program, you probably track quality through LQA scores, error rates, and reviewer feedback. You might not track ChrF. Most ops teams do not. But in our blind evaluation of five MT systems, ChrF turned out to be one of the most telling measures. 

What ChrF Measures 

ChrF stands for character F-score. It compares a machine translation against a human reference at the character level, measuring overlap. Higher ChrF means the MT output is closer to what a human translator would produce. It catches differences that word-level metrics miss: morphological variations, minor reorderings, and inflection differences. 

ChrF does not tell you whether a translation is correct. It tells you how close it is to what a qualified human would write. That closeness correlates strongly with how much editing a reviewer needs to do. 

What 16 Points Looks Like

In the benchmark, Opal scored +16 ChrF points above the next closest system. That gap typically separates one generation of MT engine from the next, like the jump from phrase-based statistical MT to neural MT around 2016-2017. It is a generational difference, not an incremental one. 

Higher ChrF means less rewriting. Fewer changed characters per segment, less time per segment, and the accumulated effect across thousands of segments is measurable in hours saved. 

The benchmark confirmed this through a second metric: edit distance. Opal cut post-editing workload in half compared to Google Translate. ChrF and edit distance measure different things (one compares to a reference, the other measures actual edits made), but they pointed in the same direction. 

Why ChrF Reveals More Than Error Counts 

A segment can have zero formal errors and still require significant editing. The terminology might be acceptable but not match the client’s terms. The register might be correct for general use but wrong for the audience. Reviewers fix all of these, and none show up in standard LQA. 

This reviewer work goes uncounted. ChrF surfaces it because it measures distance from the human reference. Opal’s 16-point advantage means that even for zero-error segments, its output is closer to what your reviewer would write. 

Half the Post-Editing Workload 

Linguists editing Opal output made roughly half the character-level changes compared to editing Google Translate output. Post-editing is where most of a localization program’s variable cost sits. Linguist time is typically the largest line item in program budgets. 

If each segment takes half as long to review, the per-word cost drops proportionally. The reviewer stays busy but processes more segments in the same time. For an LSP running production at scale, this translates directly to margin: the same review team handles more volume, or the same volume takes fewer hours. 

The caveats 

ChrF and edit distance are averages across a large dataset. Your results will vary by language pair and content type. Programs with high-quality custom MT may see a smaller gap. Programs on generic engines will likely see a larger one. ChrF also does not capture reviewer satisfaction or calendar turnaround time. The metric is informative, not comprehensive.