003Eval · Engine and vendor bake-offEarly access

Which engine
should you
actually ship?

The same source through every engine, model or vendor you are considering, scored by one stack and ranked against your own reference. You get the numbers, the evidence behind them and a note on where they are weak.

Eval is in early access. Ask for it, we run a first bake-off with you, and switch it on for your workspace.

chrF++ leads, not BLEU No model grades itself Evidence per segment
Test set
2,400 words · EN to TR
Reference
Your approved human
Primary metric
chrF++
Judge
Fixed for the run
Candidates · chrF++
Engine A74
LLM B71
Engine C66
Custom58
Notes on the run

LLM B is the judge model, so it was graded by the designated fallback judge instead. Neural metrics are not computed yet and the column says so.

EKRanked, with the caveats attachedillustrative figures
01Metrics

Metrics chosen for
your languages, not for a chart

Each metric answers a narrow question. We lead with the one that holds up on morphologically rich languages, and we say plainly what a number can and cannot tell you.

  • 01

    chrF++, the primary metric

    Character n-grams with a word component. It degrades gracefully on Turkish, Arabic and other morphologically rich targets, where word-level metrics punish a correct translation for inflecting.

  • 02

    BLEU, chrF and TER

    Kept because procurement asks for them and because a second view is cheap. TER is reported as an error rate turned into a goodness score, so every column reads in the same direction.

  • 03

    Neural metrics, planned

    COMET-22 against a reference and MetricX-24-QE without one are planned as a pluggable scoring worker. An eval run does not compute them yet. The report shows the column as not computed rather than quietly leaving it out.

  • 04

    What we will not bundle

    Some of the strongest published metrics are licensed for research only. We do not ship them inside a commercial product, and we would rather say so than let you assume they are in the stack.

  • 05

    Calibration against your own reviewers

    When you supply human scores for a set, the judge's ranking is checked against them with Kendall tau-b and pairwise accuracy, so you know how far to trust the automatic column before a decision rests on it.

02The MQM judge

A number tells you
which, not why

Alongside the metrics, an MQM-style judge reads each segment and returns typed errors with the exact target span. That is what turns a leaderboard into something you can defend in a meeting.

01

Typed errors, located

Mistranslation, omission, addition, terminology, grammar, spelling, register and locale format, each tied to the characters it applies to.

02

No fabricated spans

If the span the judge quotes cannot be found in the target, the finding is dropped rather than pinned to an approximate position.

03

No self-grading

One judge model is fixed for the whole run. When a candidate is that same model, it is graded by a designated second model from another family or skipped and recorded as skipped.

04

Deterministic by default

Temperature zero, so a re-run of the same test set gives the same reading. Repeated sampling is available when you want a consensus view.

05

Off unless you ask

The judge is a real cost per segment per candidate, so it is opt-in per run rather than on by default.

06

Fails safe

If a provider is unavailable the run still completes on the metric and rule scores, with the judge column empty instead of invented. An engine that drops a placeholder or breaks an inline tag in any segment, sampled or not, is excluded from the recommendation. The gate checks every segment, including segments outside the audit sample.

03How a run works

Same source, same rubric,
one set of numbers

01

Pick the set

A real delivery or a sample from your own content. Pasted vendor columns count as candidates too, so a human vendor can be compared with an engine.

02

Run the candidates

Each engine or model translates the same source. Cost is recorded alongside quality, because the decision is rarely quality alone.

03

Score once

Metrics, the deterministic rules and optionally the MQM judge, applied identically to every candidate.

04

Read the evidence

A ranked view you can open segment by segment, with the errors highlighted in place. Re-run it when an engine ships a new version.

04Worth knowing

What a bake-off
cannot settle

A

A reference shapes the answer

Reference-based metrics measure similarity to one approved translation, not quality in the abstract. A good candidate that phrases things differently will score lower, which is why the evidence view exists and why a human reads the close calls.

B

Content type decides more than the engine

An engine that wins on support articles can lose on campaign copy. Run the bake-off on the content you actually ship, and re-run it per content type rather than once per year.

005Early access

Test the engines on
your own content

Send a sample and tell us which engines, models or vendors are in the running. We set up the run and walk you through where the numbers are solid and where they are not.