126 of 131 WMT26 metric runs score the masculine translation higher

A WMT 2026 paper posted on 18 September 2026 scored every submitted translation-quality evaluator on 1,308 sentence pairs per language that differ only in whether an occupation is rendered masculine or feminine. Of the 131 system-language runs with a statistically detectable preference, 126 scored the masculine translation higher. The error-span systems tilt further: Gemma 4 calls the masculine Russian variant error-free 27.8 percentage points more often, and both Lexicala submissions on Arabic are at 30 points. The evaluator with the smallest average movement, COMETKiwi22 at 1.14% of its range, is also the one that picks masculine most consistently, on 78.1% of pairs.
What happened on 18 September 2026
A team from Instituto de Telecomunicações, the National Technical University of Athens and the University of Amsterdam posted arXiv:2609.21490 on 18 September 2026, accepted at the Eleventh Conference on Machine Translation (WMT26, co-located with EMNLP 2026). They ran the WMT 2026 Automated Translation Quality Evaluation submissions over an occupation-balanced subset of GAMBIT+: 1,308 masculine/feminine translation pairs per target language, three source texts for each of the 436 four-digit ISCO-08 occupational groups, seven English-source pairs into Arabic, Czech, German, Greek, Icelandic, Russian and Ukrainian.
Each pair is the same sentence twice. The source says The Member of Parliament delivered a compelling speech and leaves the gender open; one target realises it masculine, the other feminine. A metric that is doing its job should not care. After Benjamini-Hochberg correction at q = 0.05, 131 of the 148 system-language runs showed a directional preference. 126 pointed masculine. The five that pointed feminine were Vertical_870257 on Arabic, Vertical_870554 on Arabic and German, and fluency2-gemini35-esa on German and Ukrainian.
German is new in this paper. The six other pairs were inherited from the 2025 GAMBIT+ release; the German targets were generated with LongCat-2.0 and spot-checked by a named reviewer.
Thirteen evaluators, three numbers each, and they disagree
The paper's central methodological point is that one aggregate is not enough, and its own Table 2 shows why. We reproduced the table for the thirteen systems scored on all seven language pairs and computed the macro-average across languages, plus the cancellation ratio 1 - |S| / A, which says how much of an evaluator's gender sensitivity disappears when you average signed differences.
| Evaluator | What it is | Signed S% | Absolute A% | Cancellation | Source |
|---|---|---|---|---|---|
| COMETKiwi22 | trained encoder | +0.79 | 1.14 | 0.311 | Table 2, our macro-average |
| AEGIS | not disclosed | +0.40 | 1.27 | 0.683 | Table 2, our macro-average |
| gemba-poly | LLM prompt | +0.31 | 1.31 | 0.763 | Table 2, our macro-average |
| Vertical_870257 | not disclosed | +0.45 | 1.74 | 0.743 | Table 2, our macro-average |
| Vertical_870554 | not disclosed | +0.52 | 2.51 | 0.793 | Table 2, our macro-average |
| bytepop_pro9 | not disclosed | +2.42 | 3.83 | 0.369 | Table 2, our macro-average |
| bytepop_pro10 | not disclosed | +2.55 | 3.97 | 0.357 | Table 2, our macro-average |
| FACET_869543 | not disclosed | +1.90 | 4.27 | 0.554 | Table 2, our macro-average |
| FACET_869546 | not disclosed | +1.95 | 4.42 | 0.560 | Table 2, our macro-average |
| Cohere CAT+ ensemble | LLM ensemble | +2.76 | 4.79 | 0.423 | Table 2, our macro-average |
| MQM-LLM | LLM prompt | +3.50 | 5.24 | 0.331 | Table 2, our macro-average |
| MQM-LLM CA | LLM prompt | +3.81 | 5.83 | 0.347 | Table 2, our macro-average |
| fluency2-gemini35-esa | LLM prompt | -0.02 | 13.20 | 0.998 | Table 2, our macro-average |
S% is the mean signed difference, masculine minus feminine, as a percent of that evaluator's own observed score range. A% is the mean absolute difference on the same scale. Both are our macro-averages over the seven languages of Table 2 in arXiv:2609.21490. "What it is" is our label, not the paper's; several submissions do not disclose an architecture.
Two checks confirm the reconstruction. Our per-language mean A% over these thirteen systems is 3.14 for German and 4.64 for Icelandic, and the paper states exactly those bounds. Our per-language mean S% is 2.72 for Russian and 0.26 for German, which are the two values the paper names.
Rank the systems by absolute magnitude and by signed magnitude and you get different orders. The Spearman correlation between the two columns is 0.467 across the thirteen systems. fluency2-gemini35-esa moves its score by 13.20% of its range on the average pair and its signed mean is -0.02: 99.8% of its gender sensitivity cancels. COMETKiwi22 moves by 1.14% and 31% cancels, which is the lowest cancellation in the panel. Smallest effect, most consistent direction.
The preference frequencies make that concrete. COMETKiwi22 scores the masculine variant higher on 78.1% of pairs, the feminine variant higher on 20.8%, and ties on 1.1%. Its average movement is a rounding error. Its direction is not. gemba-poly goes the other way: 61.8% of its pairs are exact ties, so its low numbers partly measure a coarse output scale rather than an even hand.
The error-span task shows the same tilt, at bigger numbers
Task 1 asks systems to mark error spans rather than assign a score. Count how often each variant is predicted completely error-free, subtract, and the gap is in percentage points rather than fractions of a range. Gemma 4 calls the masculine Russian translation error-free 27.8 points more often than the feminine one; the reference-based variant reaches 28.4. Both Lexicala submissions on Arabic sit at 30.0 and 30.3 points. cuni-v14 on Czech is 21.3 and cuni-cat-v4 is 21.1. The fluency2-gemini35-spans family is the counterexample, favouring feminine variants in Czech, German and Ukrainian by up to 3.5 points.
That is the number a localization team should care about, because an error-free flag is what a pipeline acts on. A quality gate built on one of these annotators will send the feminine rendering back for review substantially more often than the masculine one, on sentences where the source gives no basis for either.
Which occupations move, and by how much
| ISCO-08 | Occupation | Mean S% over 7 languages | Favoured variant | Source |
|---|---|---|---|---|
| 2222 | Midwifery Professionals | -7.73 | feminine | Table 3 |
| 3222 | Midwifery Associate Professionals | -4.97 | feminine | Table 3 |
| 5311 | Child Care Workers | -3.41 | feminine | Table 3 |
| 5241 | Fashion and Other Models | -3.21 | feminine | Table 3 |
| 5151 | Cleaning and Housekeeping Supervisors | -2.94 | feminine | Table 3 |
| 7126 | Plumbers and Pipe Fitters | +6.96 | masculine | Table 3 |
| 6224 | Hunters and Trappers | +6.73 | masculine | Table 3 |
| 4414 | Scribes and Related Workers | +6.16 | masculine | Table 3 |
| 4213 | Pawnbrokers and Money-lenders | +6.11 | masculine | Table 3 |
| 2636 | Religious Professionals | +5.84 | masculine | Table 3 |
The five strongest feminine-preferring and five strongest masculine-preferring ISCO-08 groups, mean normalised signed difference across the seven target languages over the 13 systems with full coverage. Source: arXiv:2609.21490, Table 3.
Midwifery professionals draw the largest feminine preference at -7.73, and plumbers the largest masculine preference at +6.96. The paper is careful here and so are we: each occupation has three examples per language, so these are descriptive extremes, not stable estimates. The one that does not fit the stereotype story is Fashion and Other Models (ISCO 5241), which nets feminine at -3.21 overall while scoring positive in Arabic (+4.41) and Czech (+0.81). The sign flips by language.
What we found in the released data
We pulled the six publicly released language pairs from ailsntua/gambit-plus on 21 September 2026 and counted. Three things came out of it.
The German extension is not published yet. The repository was last modified on 23 September 2025 and gambit_plus.en-de.jsonl returns Entry not found. Every German number in the paper, including the German column of Table 2 and the 3.14 that anchors the low end of the sensitivity range, currently rests on data nobody outside the team can inspect.
The full benchmark size checks out. Each released file holds 8,771 pairs. Thirty-three source-target combinations at 8,771 gives 289,443, which is the "nearly 290,000 paired instances" the paper cites as its reason for subsampling. The ISCO code is in the meta_domain field: 436 distinct values, between 4 and 105 examples each, so a three-per-group subset is a real subsample and the specific draw is not in the release.
The pairs are not minimal. The paper's Table 1 example differs in one word, Der versus Die. In the released data the median pair differs by 2 whitespace tokens in Icelandic, 3 in Greek and Russian, 4 in Czech and Ukrainian, and 5 in Arabic. In Arabic 71.0% of pairs differ by more than three tokens; in Icelandic only 17.8% do. Much of that is legitimate agreement morphology propagating through verbs, adjectives and participles, so a token count overstates how much semantic content changed. It still means that for most Arabic pairs the two strings a metric compares are not one word apart.
| Target language | Signed S% | Absolute A% | Median tokens that differ | Pairs differing by >3 tokens | Source |
|---|---|---|---|---|---|
| Arabic | +1.75 | 4.27 | 5 | 71.0% | Table 2; our count on the released jsonl |
| Czech | +1.12 | 3.85 | 4 | 62.9% | Table 2; our count on the released jsonl |
| German | +0.26 | 3.14 | not reported | not reported | Table 2; not in the public release |
| Greek | +2.03 | 4.40 | 3 | 44.5% | Table 2; our count on the released jsonl |
| Icelandic | +2.15 | 4.64 | 2 | 17.8% | Table 2; our count on the released jsonl |
| Russian | +2.72 | 4.42 | 3 | 43.8% | Table 2; our count on the released jsonl |
| Ukrainian | +1.47 | 4.08 | 4 | 61.5% | Table 2; our count on the released jsonl |
Per-language signed and absolute differences are our macro-averages over the 13 fully covered systems in Table 2 of arXiv:2609.21490. Token differences are our own count with Python's difflib.SequenceMatcher over whitespace tokens on the released jsonl files, 8,771 pairs per language.
The obvious hypothesis is that metrics react more where the surface changes more. It does not hold. Across the six released pairs the Spearman correlation between mean tokens changed and mean A% is -0.829: Arabic changes the most text and sits mid-range on sensitivity, while Icelandic changes the least and draws the highest sensitivity of any language. With six points this is weak evidence, and it is a negative result rather than a mechanism. What it rules out is the comfortable explanation that these metrics are simply reacting to string length.
What a localization team should do with this
Do not rank your MT systems on a single QE score in a gendered target language without checking the direction. The two evaluators in this panel with near-zero average movement, COMETKiwi22 and gemba-poly, arrive there by opposite routes: one is consistently tilted and small, the other ties most of the time. Before you trust either, run the paired test on your own content.
An error-free flag is riskier than a score. A 20 to 30 point gap in error-free rate is a routing decision, not a measurement artefact. If your review queue is populated by one of these annotators, the feminine rendering gets the extra human pass and the masculine one ships.
Report three numbers, not one. Signed mean, mean absolute, and win rate. The paper's own recommendation, and our table shows the ranking changes depending on which you pick.
Ambiguity is a third option that this benchmark does not test. The paper says so in its limitations: only masculine and feminine variants are scored, gender-neutral and gender-inclusive alternatives are out of scope. If your style guide prescribes neutral forms in Greek or Czech, none of these numbers tell you how a metric will treat them.
Check it yourself
Everything in the previous section runs on a laptop in about a minute.
# the six released pairs (en-de returns "Entry not found" as of 21 Sep 2026)
for L in ar cs el is ru uk; do
curl -sL "https://huggingface.co/datasets/ailsntua/gambit-plus/resolve/main/gambit_plus.en-$L.jsonl" \
-o "en-$L.jsonl"
done
wc -l en-*.jsonl # 8771 each
python3 - <<'EOF'
import json, collections, difflib, statistics
for lp in ['ar','cs','el','is','ru','uk']:
diffs, isco = [], collections.Counter()
for line in open(f'en-{lp}.jsonl'):
d = json.loads(line)
isco[d['meta_domain']] += 1
a, b = d['masc'].split(), d['fem'].split()
sm = difflib.SequenceMatcher(None, a, b, autojunk=False)
diffs.append(sum(max(i2-i1, j2-j1)
for tag,i1,i2,j1,j2 in sm.get_opcodes() if tag != 'equal'))
big = 100*sum(1 for d in diffs if d > 3)/len(diffs)
print(lp, 'n=', len(diffs), 'ISCO groups=', len(isco),
'median=', statistics.median(diffs),
'mean=', round(statistics.mean(diffs), 2),
f'>3 tokens={big:.1f}%')
EOF
# ar 8771 436 5 5.81 71.0% ... is 8771 436 2 2.35 17.8%
The paper itself is arXiv:2609.21490; Table 2 carries the S%/A% pairs we macro-averaged, Table 3 the occupations, Table 4 the error-free rates. The shared task it scores is reported separately in Lavie et al., 2026, "Findings of the WMT26 shared task on automated translation quality evaluation".
What would prove this wrong
The German release. If ailsntua/gambit-plus publishes gambit_plus.en-de.jsonl and an independent count gives a per-system A% materially different from the 0.57 to 8.93 range in Table 2, our reproduction of the German column is wrong and so is the 3.14 low bound. Prediction, testable the day the file appears: the German pairs will have the lowest median token difference of the seven, below Icelandic's 2, because German gender is carried mostly by the article.
The direction under re-runs. The paper observes one returned run per system and says so. If the organizers publish repeated runs and COMETKiwi22's 78.1% masculine win rate moves outside roughly 70 to 85% across seeds, the consistency claim weakens even though the magnitude claim survives.
The surface correlation. Our -0.829 rests on six points and excludes German. Adding German, or running the same count on the full 33-pair GAMBIT+, could flip it. We would treat a positive correlation on eleven target languages as showing that the effect is partly string length after all.
Sources
- Orfeas Menis Mastromichalakis, Giorgos Filandrianos, Wafaa Mohammed, Giuseppe Attanasio, Chrysoula Zerva, Benchmarking Gender Bias in Machine Translation Evaluation Metrics across Occupations, arXiv:2609.21490, 18 September 2026. Accepted at WMT26 / EMNLP 2026.
- AILS NTUA, ailsntua/gambit-plus, Hugging Face dataset, last modified 23 September 2025. Downloaded and counted 21 September 2026.
- George Filandrianos et al., GAMBIT+: A challenge set for evaluating gender bias in machine translation quality estimation metrics, Proceedings of the Tenth Conference on Machine Translation, 2025, pages 314-326.
- International Labour Organization, ISCO-08 classification of occupations, the 436 four-digit groups used to index the benchmark.
Related BLOMEGA guides: Multilingual LLM judges and translationese bias · LLM judges rank machines above human translators · COMETKiwi on real localisation post-edits
FAQ
Do machine translation quality metrics prefer masculine translations?
In the WMT 2026 evaluation shared task, yes, for most systems. Across 148 system-language runs on an occupation-balanced GAMBIT+ subset, 131 showed a statistically detectable directional preference after Benjamini-Hochberg correction, and 126 of those favoured the masculine variant. The five exceptions were Vertical_870257 on Arabic, Vertical_870554 on Arabic and German, and fluency2-gemini35-esa on German and Ukrainian.
Is COMETKiwi biased on gender?
It has the smallest average movement of the thirteen fully covered systems, 1.14% of its observed score range, but it scores the masculine variant higher on 78.1% of pairs against 20.8% feminine and 1.1% ties. Small magnitude and consistent direction are different properties, and COMETKiwi22 has the lowest cancellation ratio in the panel at 0.311.
What is GAMBIT+?
A challenge set of paired machine translations that differ only in the gender realisation of an occupational reference whose gender the English source leaves open. It indexes examples by the 436 four-digit ISCO-08 occupational groups. The public release at ailsntua/gambit-plus holds 8,771 pairs per language pair; the WMT26 paper uses an occupation-balanced subset of 1,308 pairs, three per ISCO group.
Which occupations show the strongest gender preference in MT metrics?
Averaged across the seven target languages, midwifery professionals draw the strongest feminine preference at -7.73 and plumbers and pipe fitters the strongest masculine preference at +6.96. Hunters and trappers, scribes, pawnbrokers and religious professionals follow on the masculine side; child care workers and fashion models on the feminine side. With three examples per occupation per language these are descriptive extremes, not stable estimates.
Can I reproduce these numbers?
Partly. The six non-German language pairs are downloadable from ailsntua/gambit-plus and the ISCO code sits in the meta_domain field, so occupation-level splits reproduce. German was added in this paper and is not in the public repository as of 21 September 2026. The specific three-per-occupation draw used in the paper is not published either.