Typed errors, located
Mistranslation, omission, addition, terminology, grammar, spelling, register and locale format, each tied to the characters it applies to.
The same source through every engine, model or vendor you are considering, scored by one stack and ranked against your own reference. You get the numbers, the evidence behind them and a note on where they are weak.
Eval is in early access. Ask for it, we run a first bake-off with you, and switch it on for your workspace.
LLM B is the judge model, so it was graded by the designated fallback judge instead. Neural metrics are not computed yet and the column says so.
Each metric answers a narrow question. We lead with the one that holds up on morphologically rich languages, and we say plainly what a number can and cannot tell you.
Character n-grams with a word component. It degrades gracefully on Turkish, Arabic and other morphologically rich targets, where word-level metrics punish a correct translation for inflecting.
Kept because procurement asks for them and because a second view is cheap. TER is reported as an error rate turned into a goodness score, so every column reads in the same direction.
COMET-22 against a reference and MetricX-24-QE without one are planned as a pluggable scoring worker. An eval run does not compute them yet. The report shows the column as not computed rather than quietly leaving it out.
Some of the strongest published metrics are licensed for research only. We do not ship them inside a commercial product, and we would rather say so than let you assume they are in the stack.
When you supply human scores for a set, the judge's ranking is checked against them with Kendall tau-b and pairwise accuracy, so you know how far to trust the automatic column before a decision rests on it.
Alongside the metrics, an MQM-style judge reads each segment and returns typed errors with the exact target span. That is what turns a leaderboard into something you can defend in a meeting.
Mistranslation, omission, addition, terminology, grammar, spelling, register and locale format, each tied to the characters it applies to.
If the span the judge quotes cannot be found in the target, the finding is dropped rather than pinned to an approximate position.
One judge model is fixed for the whole run. When a candidate is that same model, it is graded by a designated second model from another family or skipped and recorded as skipped.
Temperature zero, so a re-run of the same test set gives the same reading. Repeated sampling is available when you want a consensus view.
The judge is a real cost per segment per candidate, so it is opt-in per run rather than on by default.
If a provider is unavailable the run still completes on the metric and rule scores, with the judge column empty instead of invented. An engine that drops a placeholder or breaks an inline tag in any segment, sampled or not, is excluded from the recommendation. The gate checks every segment, including segments outside the audit sample.
A real delivery or a sample from your own content. Pasted vendor columns count as candidates too, so a human vendor can be compared with an engine.
Each engine or model translates the same source. Cost is recorded alongside quality, because the decision is rarely quality alone.
Metrics, the deterministic rules and optionally the MQM judge, applied identically to every candidate.
A ranked view you can open segment by segment, with the errors highlighted in place. Re-run it when an engine ships a new version.
Reference-based metrics measure similarity to one approved translation, not quality in the abstract. A good candidate that phrases things differently will score lower, which is why the evidence view exists and why a human reads the close calls.
An engine that wins on support articles can lose on campaign copy. Run the bake-off on the content you actually ship, and re-run it per content type rather than once per year.
Send a sample and tell us which engines, models or vendors are in the running. We set up the run and walk you through where the numbers are solid and where they are not.