BLOMEGA

CVSS-X clones 9,851 Common Voice volunteers into 28 languages, and its files contradict its paper twice

Lab note · 20 September 2026 · BLOMEGA

Abstract technical illustration of one waveform splitting into many parallel coloured waveforms fanning out across a dark grid, cyan and orange accents

CVSS-X, posted to arXiv on 11 September 2026 by a Federal University of Goiás team with Ermis.ai, is 16,070 hours of synthetic English-to-28-language speech in which one variant clones the voices of Common Voice volunteers into every target language. In the released English-to-Portuguese training manifest, 9,851 volunteers supply 222,349 clips and 365 of them supply half. The files also contradict the paper twice: Portuguese was translated by TranslateGemma 4B, not NLLB-200, and 98.2% of test clips get the male canonical voice because 87.2% of test speakers have no gender label.

What was released between 8 and 17 September 2026

The dataset repository lgris/XVSS-X was created on Hugging Face on 8 September 2026 and last modified on 17 September. The paper, arXiv:2609.13413 (Gris, Ferreira, de Oliveira, da Rosa, Ferro Filho, Galvão Filho, Soares), went up on 11 September with a v2 on 15 September, and is accepted, non-archival, at the second SALMA workshop at EMNLP 2026. Funding is Brazil's AKCIT immersive-technology centre under an EMBRAPII grant; Huglabs and Ermis.ai are acknowledged. The code is at github.com/ErmisAI/XVSS-X.

It reverses Google's 2022 CVSS corpus. CVSS translated 21 languages into English; CVSS-X takes 240,192 English Common Voice 17 recordings (91.0% of the 264,037 CVSS used) and produces Portuguese, Spanish, French, Italian, Romanian, Catalan, German, Dutch, Swedish, Danish, Norwegian, Russian, Polish, Czech, Ukrainian, Chinese, Japanese, Korean, Finnish, Hungarian, Hindi, Persian, Greek, Hebrew, Turkish, Thai, Indonesian and Vietnamese. That is 6,725,376 translated pairs, eight times the hours of CVSS, under CC BY-NC 4.0.

There are two variants. CVSS-X-C uses two synthetic voices per language, one male and one female, designed with ElevenLabs' voice design feature. CVSS-X-T (9,340 of the 16,070 hours) uses OmniVoice zero-shot cloning conditioned on each English clip, so each volunteer's timbre speaks 28 languages they did not record. About 3.4% of clips were skipped for insufficient signal. Mean speaker similarity is 0.607 on ECAPA-TDNN, 0.648 for Germanic targets and 0.567 for Slavic.

The released manifests disagree with the paper on the translation model and the voice split

Each language folder ships a JSON manifest with the source sentence, the target text, the Common Voice client_id hash, the volunteer's self-reported gender, age and accent, and a translation.model field. We downloaded all 28 test manifests and the Portuguese train and dev manifests on 20 September.

Table 2. What the paper says against what the released files say. Manifests from huggingface.co/datasets/lgris/XVSS-X, fetched 20 September 2026 (dataset last modified 17 September 2026).
ItemPaper (arXiv:2609.13413v2)Released filesHow we got it
Translation modelNLLB-200 distilled 600M for all 28 (Section 3.1)NLLB-200 600M in 27 manifests; google/translategemma-4b-it in en-ptManifest translation.model field
Why CC BY-NC"inherited from NLLB-200" (Limitations 3)Dataset card: "due to the underlying facebook/nllb-200-distilled-600M model license"README, License
Canonical voice split81.4% male, 18.6% femaleTrain labels: male 55.5%, none 23.1%, other 2.9%, female 18.6%Our count, manifest_train.json
Test split voicenot reported87.2% of test clips carry no gender label; 98.2% get the male voiceOur count, manifest_test.json
Speakersnot reported9,851 distinct Common Voice client_ids in en-pt train; top one 6,147 clipsOur count, manifest_train.json
Test sources per language7,8437,843, identical sentence and speaker list in all 28 manifestsOur check, 28 manifest_test.json files

Portuguese is a different translation system. The paper says every target was translated with the distilled 600M NLLB-200, chosen after an LLM-judge comparison of seven systems on 100 English-to-Portuguese samples. Twenty-seven manifests agree. The en-pt manifest names google/translategemma-4b-it, which the paper lists only as future work ("regenerating translations with TranslateGemma-12B would ... enable Apache 2.0 licensing"). We cannot tell from the files whether the en-pt audio was synthesized from the TranslateGemma text or whether only the manifest was regenerated. Either way, Portuguese, the language used to pick the translator, is not the language the paper describes, and the Romance-family scores in Table 1 below include it.

The 81.4% male figure is a routing rule, not a population. The paper assigns the male canonical voice to gender=male, empty, or other. Our count of the Portuguese train manifest reproduces the paper's 18.6% female exactly (41,308 of 222,349), and shows where the rest comes from: 55.5% male, 23.1% unlabeled, 2.9% "other". Over a quarter of the training clips get a male voice because the volunteer declined to answer or did not identify as male or female.

Every unlabeled or 'other' speaker is voiced by the male canonical voice Common Voice gender labels in the CVSS-X en-pt manifests, share of utterances. Our count, 20 Sep 2026. Train n = 222,349 male 55.5% no label 23.1% female 18.6% Too thin to label: other 2.9% Canonical voice assigned: male 81.4%, female 18.6% Test n = 7,843 no label 87.2% Too thin to label: male 10.9%, other 0.2%, female 1.8% Canonical voice assigned: male 98.2%, female 1.8%
Gender labels behind the canonical voice split. The test split is the extreme case: 6,836 of 7,843 test clips have no label, so a system evaluated on CVSS-X-C test output is evaluated almost entirely on the male voice. Source: our count from the en-pt manifests of lgris/XVSS-X.
Table 1. CVSS-X quality by language family, as printed in arXiv:2609.13413v2 Tables 4 and 5. C = canonical voices, T = cloned (timbre-transferred) voices. WER/CER and ASR-BLEU come from Whisper large-v3 transcripts scored against the machine-translated prompt text, on 200 dev utterances per language. UTMOS is a predicted naturalness score, 1 to 5. The CVSS row is the English-target original, re-scored by the CVSS-X authors with the same pipeline.
FamilyWER/CER CWER/CER TASR-BLEU CASR-BLEU TUTMOS CUTMOS TSource
Romance (6)5.88.290.3883.543.27Table 4
Germanic (5)9.512.285.381.13.623.27Table 4
Slavic (4)77.786.985.93.483.17Table 4
CJK (3), CER510.275.267.93.563.2Table 4
Uralic (2)10.314.38478.33.543.21Table 4
Indo-Iranian (2)22.52363.7603.613.2Table 4
Other (6)24.824.86564.93.533.14Table 4
CVSS-X average12.114.182.479.43.553.21Table 4
CVSS (X to EN), re-scored3.5494.293.84.433.61Table 5

Read Table 1 for what it measures. The abstract says translation quality is "comparable to CVSS". The numbers are 82.4 against 94.2 ASR-BLEU, 12.1% against 3.5% WER and 3.55 against 4.43 UTMOS, and none of them compares a translation against a human translation. Hebrew scores 46.1 ASR-BLEU and 37.4% WER; Thai scores 77.0% WER, which the authors attribute to Whisper dropping tone diacritics.

ASR-BLEU checks that the audio says the machine translation, so a fluent mistranslation scores perfectly

The evaluation transcribes 200 synthesized dev clips per language with Whisper large-v3 and computes WER and BLEU against the translated prompt text. That is a text-to-speech round trip. If NLLB-200 renders a sentence wrongly and OmniVoice speaks the wrong sentence clearly, WER is zero and ASR-BLEU is 100.

The quality loop closes on the machine translation, not on a human reference Common Voice 17 EN 240,192 clips matched 9,851 speakers (en-pt train) Machine translation NLLB-200 600M: 27 langs TranslateGemma 4B: pt CVSS-X-C 2 ElevenLabs voices per language CVSS-X-T OmniVoice clones each volunteer 28 languages ~16,070 hours 6,725,376 pairs Whisper large-v3 ASR on 200 dev clips per language ASR-BLEU 82.4 (C) / 79.4 (T) scored against the MT text not a human translation reference = MT output What the loop can catch: TTS that mispronounces or drops words (Hebrew ASR-BLEU 46.1, Thai WER 77.0%). What it cannot catch: a wrong translation spoken clearly. Mistranslations from NLLB-200 score as perfect. The authors say so in Limitations (2): a check against human references (CoVoST 2 EN to 15 languages) "remains an ongoing benchmark". The abstract still calls translation quality "comparable to CVSS".
Where each number in the paper is measured. The dashed arrow is the problem: the reference for every quality score is the machine translation itself. Source: arXiv:2609.13413v2, Sections 3 and 4; translation model per language from the released manifests.

The cloning branch has a second property the paper does not discuss: concentration. Common Voice is crowd-recorded, so a few volunteers read a lot. In the Portuguese train manifest the most frequent speaker, a volunteer labelled male, supplies 6,147 clips.

365 of 9,851 volunteers supply half the training clips 0 25 50 75 100 Cumulative share of en-pt training utterances (%). Our count from manifest_train.json, 20 Sep 2026. Top 1 speaker 2.76% 6,147 clips Top 10 speakers 12.39% 27,538 clips Top 365 speakers 50% half of train All 9,851 speakers 100% 222,349 clips
Speaker concentration in CVSS-X en-pt train. All 28 test manifests carry the identical sentence and speaker list, and the paper gives 240,192 samples for every language, so the same volunteers very likely recur across all 28 languages; we verified this only for the test split. Source: our count from manifest_train.json.

If the train list is shared the way the test list is, that one volunteer's cloned voice appears in up to 172,116 clips (6,147 times 28), or about 166,000 after the 3.4% cloning skip rate. We have not downloaded the audio, so this is an upper bound from the manifests, not a count of files.

For dubbing and speech-translation teams: use it to pretrain, not to evaluate or to ship

Licence first. The synthetic pairs are CC BY-NC 4.0, so CVSS-X cannot go into a commercial speech-to-speech or dubbing model. The audio sources are Common Voice, which is CC0; the non-commercial restriction comes from the translation model, and the dataset card names NLLB-200 as the reason even for the one language where the manifest names a different model.

Consent is a judgement call the paper does not make. Common Voice's dataset terms ask users "to not attempt to determine the identity of speakers". Cloning a volunteer's timbre into 28 languages does not identify anyone, and CC0 permits it. It is still a use the volunteers were not asked about when they read sentences to improve speech recognition, and the manifests keep the client_id on every row, which makes it trivial to pull every synthetic clip of one person. A judgement, marked as one: for a commercial dubbing pipeline, a voice you did not get consent to clone is a liability whatever the licence says, and that is the gap consented voice data exists to fill.

Do not evaluate on the test split as-is. The canonical test audio is 98.2% one male voice per language, and the targets are machine translations. A model that learns NLLB-200's output style will score well against it. Evaluate on human references such as CoVoST 2 or FLEURS instead.

Treat Hebrew and Thai as unverified. At 46.1 ASR-BLEU and 77.0% WER respectively, the round trip itself is failing, so either the synthesis or the ASR is wrong, and the paper cannot say which.

Pin the Portuguese version. If you train on en-pt, record the manifest's model field and the dataset commit you pulled. Results on en-pt may not be comparable with the other 27 languages or with a later corrected release.

Check it yourself

# 1. Translation model per language (28 small files, about 150 MB in total)
for l in pt es fr it ro ca de nl sv da no ru pl cs uk zh ja ko fi hu hi fa el he tr th id vi; do
  curl -sL "https://huggingface.co/datasets/lgris/XVSS-X/resolve/main/data/en-$l/manifest_test.json" -o mt_$l.json
done
python3 - <<'EOF'
import json, glob, collections
m = collections.Counter()
for f in glob.glob("mt_*.json"):
    d = json.loads(open(f).read().replace("NaN", "null"))
    m[d["translation"]["model"]] += 1
    if "translategemma" in d["translation"]["model"]: print(f, d["translation"])
print(m)
EOF
# printed on 20 Sep 2026:
# mt_pt.json {'model': 'google/translategemma-4b-it', 'source_lang': 'en', 'target_lang': 'pt'}
# Counter({'facebook/nllb-200-distilled-600M': 27, 'google/translategemma-4b-it': 1})

# 2. Gender labels and speaker concentration (train manifest is 170 MB)
curl -sL "https://huggingface.co/datasets/lgris/XVSS-X/resolve/main/data/en-pt/manifest_train.json" -o tr.json
python3 - <<'EOF'
import json, collections
S = json.loads(open("tr.json").read().replace("NaN", "null"))["samples"]
print(collections.Counter(str(s["gender"]) for s in S).most_common())
c = collections.Counter(s["client_id"] for s in S).most_common()
acc = k = 0
for _, v in c:
    acc += v; k += 1
    if acc >= len(S) / 2: break
print(len(S), len(c), c[0][1], k)
EOF
# [('male', 123364), ('None', 51289), ('female', 41308), ('other', 6388)]
# 222349 9851 6147 365

Common Voice's speaker-identity term is quoted on the Common Voice dataset card under "Personal and Sensitive Information". The paper's Table 4 and 5 values are on page 3 of the PDF at https://arxiv.org/pdf/2609.13413v2.

What would prove this wrong

Our claim is that CVSS-X's reported quality measures synthesis, not translation. The authors list a check against human CoVoST 2 references as ongoing. If that check, once published, gives English-to-X BLEU against human references within 5 points of the 82.4 ASR-BLEU average, the gap we describe is small and our advice to avoid it for evaluation is too cautious. We will look for it in the repository and on arXiv by 31 March 2027.

Narrower: if the en-pt manifest's translation.model field turns out to be a labelling error, the next dataset revision will change it to NLLB-200 without changing the target texts. If instead the Portuguese target texts change, the Portuguese audio was generated from a model the paper does not describe. Either outcome is visible in the Hugging Face commit history.

Sources

  1. Lucas Rafael Stefanel Gris et al. (Federal University of Goiás, Federal University of Mato Grosso, São Paulo State University), CVSS-X: A Multilingual Speech-to-Speech Translation Corpus for 28 Languages, arXiv:2609.13413, v1 11 September 2026, v2 15 September 2026. Tables 1 to 5, Sections 3.1 to 3.2, Limitations.
  2. lgris, XVSS-X dataset on Hugging Face, created 8 September 2026, last modified 17 September 2026. Manifests data/en-*/manifest_*.json and dataset card licence section, fetched 20 September 2026.
  3. ErmisAI, XVSS-X code repository, GitHub, accessed 20 September 2026.
  4. Ye Jia et al., CVSS Corpus and Massively Multilingual Speech-to-Speech Translation, LREC 2022, arXiv:2201.03713.
  5. Mozilla Common Voice, dataset card (CC0; "You agree to not attempt to determine the identity of speakers in the Common Voice dataset"), accessed 20 September 2026.
  6. Han Zhu et al., OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models, arXiv:2604.00688, 2026.

Related BLOMEGA guides: Cross-lingual voice cloning: the identity tradeoff at IWSLT 2026 · Consented AI training data providers · Data provenance and chain of title

FAQ

What is CVSS-X?

A synthetic speech-to-speech translation corpus released in September 2026 (arXiv:2609.13413): 240,192 English Common Voice 17 recordings machine-translated into 28 languages and synthesized with OmniVoice, about 16,070 hours in total, under CC BY-NC 4.0. It has a canonical-voice variant and a variant that clones each English speaker's voice.

Can CVSS-X be used to train a commercial dubbing or speech translation model?

No. The synthetic pairs are licensed CC BY-NC 4.0, which forbids commercial use. The underlying Common Voice audio is CC0, but the dataset card attributes the non-commercial restriction to the NLLB-200 translation model.

Did Common Voice volunteers consent to having their voices cloned?

They released their recordings under CC0, which permits it legally. Common Voice's dataset terms ask users not to try to identify speakers; they do not mention voice synthesis. The CVSS-X paper does not discuss consent. In the English-to-Portuguese training manifest, 9,851 volunteers are cloned and the most frequent one supplies 6,147 clips.

Is CVSS-X translation quality comparable to CVSS?

The paper's evidence does not show that. Its ASR-BLEU of 82.4 (versus 94.2 for CVSS) scores Whisper transcripts against the machine translation itself, so it measures whether the audio says the prompt text, not whether the translation is correct. A comparison against human references is listed as ongoing.

Which translation model did CVSS-X use?

The paper says NLLB-200 distilled 600M for all languages. The released manifests say NLLB-200 for 27 languages and google/translategemma-4b-it for English to Portuguese, as of 20 September 2026.