A translation score is a single number between 0 and 100. High is good. That is the whole reason it exists: an error rate of 3 % needs a moment of thought, a score of 87 does not. This article explains how the Glotari Translation Score is built, so that you can decide how much weight to give it.
- Sample
- based on 87 fields
- Language pair
- English → German
- Checked
- checked 2 min ago
One number, built from findings that already exist
The score triggers no model call of its own. It is an aggregation of findings the Judge has already produced for each field: category, severity, reasoning, kept for 90 days. Calculating the score costs nothing.
For any set of checked fields, a job, a language, a whole shop or a period, the score is 100 × (1 − Σ gᵢ / n). Here n is the number of checked fields and gᵢ is the weight of the most severe finding in field i. A field without findings has gᵢ = 0. Several findings in one field count once, with the highest weight, so a single broken field cannot dominate the result.
The weights
Four of the five heaviest finding types come from deterministic checks that cost nothing. A quick check without any model call is therefore already a solid base for the score. The Judge refines it.
Weights are versioned. A score always states which version produced it. Changing a weight creates a new version, old scores are not recalculated.
What the score cannot tell you
It cannot tell you whether a translation sells. Tone and style carry a weight of 0,15 because they are a matter of taste, not an error. It cannot replace a native reader for legal texts. And it is only as good as the sample it was calculated on.
Sample size is part of the number
A score from 12 fields is not the same as one from 12.000. Below 30 fields Glotari shows the score in grey, marked preliminary, with no band. Between 30 and 199 fields the band appears together with the sample size. From 200 fields the addition disappears. No confidence interval, the sample size is enough.
| FINDING | ORIGIN | WEIGHT |
|---|---|---|
| Translation missing | deterministic | 1,0 |
| Identical to source | deterministic | 1,0 |
| Numbers differ | deterministic | 1,0 |
| HTML structure broken | deterministic | 1,0 |
| Never-translate term translated | deterministic | 0,8 |
| Meaning flagged | Judge | 0,7 |
| Grammar flagged | Judge | 0,3 |
| Tone or style flagged | Judge | 0,15 |
| No verdict, source outdated, not checkable | status | not counted in n |
A second opinion is a check, not proof of correctness. Different model families disagree, and that disagreement is the signal.
Bands
- 95 to 100
- Ready to publish
- 85 to 94
- Good, a few fields to fix
- 70 to 84
- Review before publishing
- 0 to 69
- Not ready
An example: 87 from 87 fields
Three findings in the diff below, plus a handful of grammar and tone notes across the sample.
Source: Glotari test shop, judge findings, September 2026
What you do in the app
follows
Three numbered steps once the app screens are final.