BLOMEGA

Meta's 8B Omnilingual MT beats Llama 3 70B into mid-resource languages and loses into zero-resource ones

Lab note · 15 September 2026 · BLOMEGA

Abstract terraced strata thinning toward the edge of the frame, dense at one end and sparse at the other

Translating out of English on Meta's BOUQuET benchmark, OMT-LLaMA 8B scores 45.8 chrF++ into mid-resource languages against 37.2 for Llama 3 70B, a lead of 8.6 points, and 30.8 against 23.7 into low-resource ones. Into zero-resource languages, with under 1,000 parallel documents, the order flips: Llama 3 70B scores 14.3 and OMT-LLaMA 8B 12.6. No model in the paper's Table 9.2 exceeds 14.3 there. Specialization buys a lot where some parallel data exists, and nothing where none does.

What Meta released on 17 March 2026, and what it did not

17 March 2026. Meta's FAIR team posted Omnilingual MT: Machine Translation for 1,600 Languages (arXiv:2603.16309), revised to v3 on 7 May 2026. The paper describes two model families built on Llama 3: OMT-LLaMA, decoder-only, at 1B, 3B and 8B parameters, and OMT-NLLB, a 3B encoder-decoder on the OmniSONAR embedding space. It reports "non-trivial performance when translating from 1,600 and into about 1,200 languages" and says the number of languages modern models "understand sufficiently well" doubles from about 200 to over 400.

The paper also ships evaluation infrastructure: BOUQuET, a multilingual evaluation set built from scratch, Met-BOUQuET with human quality judgements across 161 language directions, the BLASER 3 reference-free quality estimator and the OmniTOX toxicity classifier. BOUQuET and Met-BOUQuET are published on Hugging Face with a leaderboard.

What we could not find is the models. On 15 September 2026, Hugging Face searches for "omnilingual" and "OMT-LLaMA" returned Omnilingual ASR models and community conversions of them, and no OMT translation weights. GitHub's facebookresearch organization has an omnilingual-asr repository and no MT counterpart. The paper's conclusion encourages the community to use the "OMT-LLaMA, OMT-NLLB, BLASER 3 and OmniTOX recipes". That is a recipe release plus an evaluation release, as far as we can verify.

This matters because the speech side is open. Omnilingual ASR, released in November 2025, is open source, covers 1,600+ languages including more than 500 never before supported by ASR, and ships models from 300M to 7B parameters. The paper suggests cascading it with Omnilingual MT for speech translation. A localization team can run the first half of that cascade today and not the second.

Where does specialization help, by resource tier?

Headline averages hide the shape. Table 9.2 of the paper splits BOUQuET chrF++ by the resource level of the non-English language, in both directions, for 18 systems. The rows below are the ones a localization team would actually weigh.

Table 1. BOUQuET chrF++ by resource level of the non-English language, from Table 9.2 of arXiv:2603.16309 v3. En-YY is translation out of English. Model sizes from the paper's Table 9.1. Bold marks the best value in each En-YY column among these rows.
SystemSizeEn-YY highmidlowv. lowzeroEn-YY totalXX-En zeroSource
OMT-LLaMA8B60.745.830.818.812.632.821.6Table 9.2
OMT-NLLB3B61.147.627.317.511.531.923.2Table 9.2
OMT-LLaMA1B56.542.225.114.912.329.119.9Table 9.2
NLLB-2003B62.946.824.617.513.231.322.7Table 9.2
GPT-OSS120B63.643.726.016.013.430.824.3Table 9.2
Gemma 327B62.640.924.014.811.628.924.6Table 9.2
TranslateGemma27B59.542.022.115.512.428.524.9Table 9.2
Llama 370B60.237.223.716.214.328.224.0Table 9.2
Tiny Aya Global3B58.529.513.39.010.321.020.6Table 9.2

Four readings, all from that table.

At the top, size wins and OMT does not. Into high-resource languages GPT-OSS 120B (63.6), NLLB-200 (62.9) and Gemma 3 27B (62.6) all beat OMT-LLaMA 8B (60.7). If your launch locales are Spanish, German and Japanese, this paper is not about you.

The gain lives in the middle. Against Llama 3 70B out of English, OMT-LLaMA 8B leads by 0.5 points on high-resource languages, 8.6 on mid, 7.1 on low and 2.6 on very-low. The even smaller OMT-LLaMA 1B, at 29.1 total, beats the 70B model's 28.2.

At zero resource, every model is near the floor. The En-YY zero column runs from 10.3 (Tiny Aya Global) to 14.3 (Llama 3 70B). Into English from zero-resource languages, TranslateGemma leads at 24.9 and OMT-LLaMA 8B scores 21.6, 2.4 points behind Llama 3 70B.

The "1B to 8B match or exceed a 70B baseline" claim is true on totals, not on every tier. The abstract's statement holds for the average. The zero-resource column is where it does not.

Out of English, by resource level of the target language BOUQuET chrF++ · orange = OMT-LLaMA 8B · grey = Llama 3 70B · blue = NLLB-200 3B 0 20 40 60 chrF++ 60.7 60.2 62.9 high 45.8 37.2 46.8 mid +8.6 30.8 23.7 24.6 low +7.1 18.8 16.2 17.5 very low +2.6 12.6 14.3 13.2 zero -1.7 Green and red figures: OMT-LLaMA 8B minus Llama 3 70B. At the high tier the difference is +0.5.
The 8B specialist and the 3B NLLB-200 track each other within 2.2 points everywhere except the low tier, where the 8B model leads by 6.2.

The long-tail figures in section 9.1.3 tell the same story at a coarser grain. On a Bible benchmark of 1,560 languages translated into English, with MetricX mapped to an estimated human XSTS+R+P score, OMT-LLaMA 8B passes the 2.5 "passable" threshold for 440 languages, OMT-NLLB for 416 and NLLB-200 for 221. At the 3.5 "good" threshold all three sit at around 130. The expansion is in passable, not in good.

Languages passing a quality bar, of 1,560 in the Bible benchmark translation into English · estimated XSTS+R+P from MetricX · section 9.1.3 OMT-LLaMA 8B 440 passable OMT-NLLB 3B 416 passable NLLB-200 3B 221 passable about 130 good about 130 good about 130 good orange or grey bar: passable, score of 2.5 or more · green bar: good, 3.5 or more (the paper says "around 130" for all three) Out of English into the long tail: baselines fall to near-random at about 300 to 400 languages; OMT holds for about 1,200. The paper's own feature table (Table 9.5) says OMT-LLaMA generates "around 1000" languages and OMT-NLLB "around 250".
Doubling the passable count while leaving the good count flat is the paper's result in one picture. It widens what a model can roughly understand far more than what it can write well.

There is an internal inconsistency worth flagging. The abstract and section 9.1.3 say generation holds for "about 1,200" languages. Table 9.5, comparing the two model families, lists OMT-LLaMA as generating "around 1000" languages and OMT-NLLB "around 250". The gap is probably a difference between "above random" and "supported", but the paper does not reconcile them.

Why the gains stop at the zero tier

The paper defines its tiers by parallel documents from primary sources, not mined or synthetic. It reports a clear quality shift above 1 million parallel documents and another qualitative change near 40,000, "comparable to that of the Bible, supplemented by at least one additional source of parallel training data". Everything OMT adds works by manufacturing or amplifying parallel signal, and each technique needs something to amplify.

Every OMT technique amplifies in-language data that has to exist first tier thresholds from section 3 · deltas from Table 9.2 · retrieval gains from Table 6.2 (56 BOUQuET directions) high > 50M docs +0.5 mid > 1M docs +8.6 low 40K to 1M docs +7.1 very low 1K to 40K docs +2.6 zero < 1K docs -1.7 OMT-LLaMA 8B minus Llama 3 70B, chrF++ out of English, by tier of the target language Backtranslation, mining needs monolingual web text and a language ID model; GlotLID covers 1,880 nothing to mine below the floor Retrieval in the prompt OMT-LLaMA 8B, Table 6.2: +2.30 with 30K+ samples +0.52 with fewer gain scales with what you can retrieve MeDLEy seed data manually created, 109 low-resource languages, 92 not in SMOL authors: gains "generally small"
The zero tier is defined by the absence of the thing every technique on the bottom row consumes. A bigger general model does marginally better there because it transfers from related languages; nothing in the specialist recipe replaces the missing text.

The retrieval numbers are the cleanest illustration. In Table 6.2, across 56 BOUQuET directions at sentence level, adding retrieved examples to OMT-LLaMA 8B lifts chrF++ from 39.83 to 42.13 on the 31 directions with at least 30,000 retrieval samples, and from 31.04 to 31.56 on the 25 with fewer. The same retrieval on the 70B Llama model (the section specifies LLaMA 3.3 70B) adds 3.51 and 0.75. More in-language examples, more gain, for both models.

The manual seed data result is candid. The authors write that improvements from adding MeDLEy to existing seed datasets are "generally small, indicating the challenges of making significant improvements for LRLs via manual collection of data at the scale of a few thousands of sentences", and that scores "remain low in general, especially in the en-xx direction". Their extension experiments in Table 10.1 add that fine-tuning on targeted parallel data improves translation out of English but hurts translation into English, which the authors attribute to degeneration after training on repetitive English outputs.

Tokenization is one fix that does not need more data. OMT's extended tokenizer (256K vocabulary against Llama 3's 128K in the paper's ablation) averages 44.8 tokens per sentence over the 212 FLORES+ languages against 80.7 for the original Llama 3 tokenizer. That is 44.5% fewer tokens for the same sentence, which lowers inference cost for long-tail locales before any quality question arises. Our arithmetic on the paper's figures.

What to do with this if you ship long-tail locales

Bucket every target locale by parallel documents before choosing a model. The paper's thresholds (1M and 40K) are the most useful operational numbers in it. Above 50M (high), large general models and NLLB-200 match or beat the specialist. Between 1M and 50M (mid) and between 40K and 1M (low), OMT-LLaMA 8B leads Llama 3 70B out of English by 8.6 and 7.1 chrF++. Between 1K and 40K the lead shrinks to 2.6. Below 1K, nothing in the table is usable for publishing without a human writing the target text.

Do not plan around OMT weights you cannot download. As of 15 September 2026 we could not find them. NLLB-200 3B, which is released, is within 1.5 chrF++ of OMT-LLaMA 8B on the En-YY total (31.3 against 32.8) and ahead of it at the high and mid tiers. For many teams it remains the practical baseline, and the BOUQuET leaderboard lets you check any candidate on the same test set.

Direction matters more than vendors admit. Among the systems in Table 1, translation into English from zero-resource languages scores 19.9 to 24.9. Out of English into the same tier, 10.3 to 14.3. A product that reads user input in a long-tail language and responds in English is a very different engineering problem from one that writes in that language. Price and staff them differently.

For the zero and very-low tiers, the budget line is in-language text, not a bigger model. Retrieval gains more than quadruple when a direction has 30,000+ examples to draw on (2.30 against 0.52 for OMT-LLaMA 8B). Our judgement: the cheapest quality improvement for a very-low-resource locale is collecting enough consented, domain-matched parallel sentences to cross that retrieval threshold, and the MeDLEy result says a few thousand hand-built sentences will not get you there on their own.

Keep humans on the output side. The good-quality count stays at about 130 languages no matter which model you use. Beyond those, a native reviewer is not a nice-to-have. The model's output is a draft that can be understood, not text that can be shipped.

Check it yourself

Table 9.2 is in the HTML version of the paper, the evaluation data and leaderboard are public, and the release status takes two API calls.

# 1. the table: open section 9.1.2, "Performance on BOUQuET by the language resource level"
#    https://arxiv.org/html/2603.16309v3

# 2. the evaluation data and the leaderboard
pip install -U "huggingface_hub[cli]"
huggingface-cli download facebook/bouquet --repo-type dataset --local-dir bouquet
#    leaderboard: https://huggingface.co/spaces/facebook/bouquet

# 3. release status of the MT models (2026-09-15: no OMT translation weights found)
curl -s "https://huggingface.co/api/models?search=OMT-LLaMA&limit=20"
curl -s "https://huggingface.co/api/models?search=omnilingual&limit=50" | python3 -c \
  "import sys,json;print([m['id'] for m in json.load(sys.stdin)])"
curl -s "https://api.github.com/search/repositories?q=omnilingual+org:facebookresearch" \
  | python3 -c "import sys,json;print([r['full_name'] for r in json.load(sys.stdin)['items']])"

# 4. the tier deltas in this note, from Table 9.2 (En-YY, then XX-En)
python3 - <<'PY'
tiers = ["high","mid","low","v.low","zero","total"]
omt8  = {"en-yy":[60.7,45.8,30.8,18.8,12.6,32.8], "xx-en":[65.1,53.5,38.7,27.6,21.6,40.6]}
l70   = {"en-yy":[60.2,37.2,23.7,16.2,14.3,28.2], "xx-en":[65.0,47.6,33.3,26.2,24.0,37.6]}
for d in omt8:
    print(d, {t: round(a-b,1) for t,a,b in zip(tiers, omt8[d], l70[d])})
PY
# en-yy {'high': 0.5, 'mid': 8.6, 'low': 7.1, 'v.low': 2.6, 'zero': -1.7, 'total': 4.6}
# xx-en {'high': 0.1, 'mid': 5.9, 'low': 5.4, 'v.low': 1.4, 'zero': -2.4, 'total': 3.0}

To place your own locale in a tier, count the sentence-level parallel data you can actually license for it, excluding machine-translated and mined pairs, and compare against 1,000, 40,000 and 1 million documents.

What would prove this wrong

The claim under test is that specialized translation recipes do not improve generation into zero-resource languages, because they amplify in-language data that those languages lack. It is wrong if, by 31 March 2027, any model on the public BOUQuET leaderboard or in a peer-reviewed paper reports an average chrF++ of 20 or more translating out of English into BOUQuET's zero-resource languages, without adding human-created parallel data for those languages. The best in Table 9.2 today is 14.3.

A second prediction, marked as judgement: Meta will not publish downloadable OMT-LLaMA or OMT-NLLB translation weights before 31 March 2027, having released the ASR models and the evaluation sets but not the MT models in the first six months. If the weights appear on Hugging Face or GitHub before that date, this reading was wrong.

FAQ

How many languages can Meta's Omnilingual MT translate?

The paper reports non-trivial performance from about 1,600 languages and into about 1,200. On a 1,560-language Bible benchmark into English, OMT-LLaMA 8B passes a "passable" quality threshold for 440 languages, OMT-NLLB for 416 and NLLB-200 for 221. At the "good" threshold all three cover around 130. The paper's Table 9.5 lists OMT-LLaMA as generating around 1,000 languages and OMT-NLLB around 250.

Is a specialized translation model better than a large general LLM for low-resource languages?

For mid-, low- and very-low-resource languages on BOUQuET, yes. OMT-LLaMA 8B scores 45.8 chrF++ out of English into mid-resource languages against 37.2 for Llama 3 70B, and 30.8 against 23.7 into low-resource ones. Into zero-resource languages Llama 3 70B scores 14.3 and OMT-LLaMA 8B 12.6, and no model in the table exceeds 14.3.

What counts as a low-resource language in machine translation?

In the Omnilingual MT paper: high resource above 50 million parallel document pairs, mid above 1 million, low from 40,000 to 1 million, extremely low from 1,000 to 40,000, zero below 1,000. The authors observe quality shifts at about 1 million and about 40,000.

Are Omnilingual MT model weights released?

We could not find them on 15 September 2026. Hugging Face and GitHub searches returned Omnilingual ASR models and repositories only. BOUQuET and Met-BOUQuET are published at huggingface.co/datasets/facebook/bouquet with a leaderboard, and the paper encourages use of its recipes.

Sources

  1. The Omnilingual MT Team (Meta FAIR), Omnilingual MT: Machine Translation for 1,600 Languages, arXiv:2603.16309, v1 17 March 2026, v3 7 May 2026, HTML. Section 3 (resource tiers), 4.3 (MeDLEy, 109 languages, 92 not in SMOL), 5 (tokenizer, 44.8 vs 80.7 tokens per sentence over 212 FLORES+ languages), Table 6.2 (retrieval), Table 9.1 (systems and sizes), Table 9.2 (BOUQuET chrF++ by resource level), 9.1.3 (Bible long-tail counts), Table 9.5 (family comparison), Table 10.1 (extension).
  2. Meta AI, Omnilingual MT research publication page, 17 March 2026.
  3. facebook/bouquet dataset and BOUQuET leaderboard space on Hugging Face, both returning HTTP 200 on 15 September 2026.
  4. Omnilingual ASR Team, Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages, arXiv:2511.09690, 12 November 2025, and facebookresearch/omnilingual-asr on GitHub.
  5. Release-status observations: Hugging Face model API searches for "omnilingual" and "OMT-LLaMA", and GitHub repository search for "omnilingual" in facebookresearch, run by BLOMEGA on 15 September 2026. Commands in "Check it yourself".

Related BLOMEGA guides: A 30-trillion-token corpus buys Basque a 160-million-parameter model · Your multilingual LLM judge prefers the machine translation · Consented AI training data providers.

BLOMEGA collects consented, domain-matched parallel text and speech for very-low-resource locales, with native-speaker review. Contact [email protected].