# BLOMEGA - full text of every article 32 articles. Newest: 2026-09-16. Source of truth for company facts: https://blomega.com/api/company/facts.json Each article below is also available on its own at the URL given in its header. ============================================================================== URL: https://blomega.com/guides/cross-lingual-voice-cloning-identity-tradeoff-2026/ Published: 2026-09-16 | Updated: 2026-09-16 ============================================================================== --- title: "In IWSLT 2026's first voice-cloning track, the best clone of your speaker got the script most wrong" url: https://blomega.com/guides/cross-lingual-voice-cloning-identity-tradeoff-2026/ published: 2026-09-16 updated: 2026-09-16 source: BLOMEGA (https://blomega.com/) --- # In IWSLT 2026's first voice-cloning track, the best clone of your speaker got the script most wrong Lab note · 16 September 2026 · BLOMEGA Rank the five submissions to the first IWSLT Cross-Lingual Voice Cloning shared task by Chinese word error rate, then rank them by speaker similarity. You get the same order. The system with the cleanest script, IIT-Patna at **0.042 WER**, had the least recognisable voice at **0.499** similarity. The system with the most recognisable voice, SIT-TCD at **0.789**, misread nearly one word in five at **0.191 WER**. Spearman rank correlation between the two: **1.0** for Chinese, 0.8 for French. Whoever tells you a 2026 voice clone gives you both is quoting one column of a two-column table. ## The task ran for the first time in 2026, with 12 real speakers per language **July 2026, San Diego.** The 23rd IWSLT evaluation campaign added a Cross-Lingual Voice Cloning track. Five teams entered: HW-TSC, IIT-Patna, KIT, Langswap and SIT-TCD. The setup is deliberately narrow, which is what makes the results readable. Systems get English reference audio from a source speaker and target-language text in Arabic, Chinese or French, and must synthesise that text in that speaker's voice. No translation step, no source transcript, no ground-truth target audio. Results are in [Speech Translation and Metrics in 2026](https://aclanthology.org/2026.iwslt-1.39/), Track VIII, Tables 13 and 14. The reference material is unusually generous. Twelve reference audio files per target language, each a different source speaker, extracted from 12 ACL 2023 presentations in English with diverse speaker accents, most of them around five minutes long. The organisers note they deliberately did not trim them, since voice cloning normally needs a few seconds, and say some submissions benefited from the extra length. Target text came from bilingual journals: arXiv (Chinese, French), TAL (French), Journal of Software (Chinese), and An-Najah, Palestine Ahliya and Princess Sumaya University journals (Arabic), filtered by LaBSE semantic similarity. The reference text is 49 lines in Arabic, 112 in Chinese and 99 in French, which is exactly the 588, 1,344 and 1,188 evaluation samples the paper reports, at 12 speakers each. Metrics, all automatic. Content consistency: WER and CER from faster-whisper large-v3 with beam size 5 and voice-activity filtering, after NFKC normalisation, lowercasing, punctuation removal and jieba segmentation for Chinese. Speaker similarity: cosine similarity of ECAPA-TDNN embeddings in SpeechBrain. Prosody similarity: cosine similarity over mean and standard deviation of fundamental frequency plus mean energy, extracted with Praat-Parselmouth. No human listening test was reported. One thing the findings paper does not describe is a consent or release process for the 12 source speakers whose voices were cloned into three languages by five research teams. The presentations are public; that is not the same fact. For anyone building a commercial dubbing pipeline, the gap between "the recording is public" and "the speaker agreed to this use of their voice" is the entire compliance surface. ## All 14 published results, both columns, side by side The findings paper splits content consistency (Table 13) from speaker and prosody similarity (Table 14). They belong next to each other, because neither is interpretable alone. | Target | Team | Backbone | Adaptation | WER | CER | Speaker sim. | Prosody sim. | Source | | --- | --- | --- | --- | --- | --- | --- | --- | --- | | Arabic | IIT-Patna | VoxCPM2 | zero-shot | 0.219 | 0.135 | 0.669 | 0.990 | Tables 13, 14 | | Arabic | KIT | Fish Audio S2 Pro | GRPO | 0.157 | 0.055 | 0.637 | 0.982 | Tables 13, 14 | | Arabic | Langswap | OmniVoice | zero-shot | 0.160 | 0.063 | 0.785 | 0.996 | Tables 13, 14 | | Arabic | SIT-TCD | OmniVoice | LoRA | **0.132** | **0.050** | **0.786** | **0.997** | Tables 13, 14 | | Chinese | HW-TSC | Qwen3-TTS 1.7B | zero-shot | 0.043 | 0.047 | 0.580 | 0.989 | Tables 13, 14 | | Chinese | IIT-Patna | Qwen3-TTS 1.7B | zero-shot | **0.042** | **0.046** | 0.499 | 0.989 | Tables 13, 14 | | Chinese | KIT | Fish Audio S2 Pro | GRPO | 0.113 | 0.099 | 0.609 | 0.981 | Tables 13, 14 | | Chinese | Langswap | Qwen3-TTS 0.6B | zero-shot | 0.189 | 0.171 | 0.686 | 0.989 | Tables 13, 14 | | Chinese | SIT-TCD | OmniVoice | LoRA | 0.191 | 0.181 | **0.789** | **0.993** | Tables 13, 14 | | French | HW-TSC | Qwen3-TTS 1.7B | zero-shot | **0.050** | **0.010** | 0.582 | 0.987 | Tables 13, 14 | | French | IIT-Patna | Qwen3-TTS 1.7B | zero-shot | 0.051 | **0.010** | 0.479 | 0.987 | Tables 13, 14 | | French | KIT | Fish Audio S2 Pro | GRPO | 0.063 | 0.017 | 0.602 | 0.980 | Tables 13, 14 | | French | Langswap | Qwen3-TTS 0.6B | zero-shot | 0.133 | 0.052 | 0.761 | 0.991 | Tables 13, 14 | | French | SIT-TCD | OmniVoice | LoRA | 0.069 | 0.020 | **0.813** | **0.996** | Tables 13, 14 | Three readings, all from that table. **Chinese is the clean case.** Five systems, and the WER order and the speaker-similarity order match exactly. Spearman rank correlation 1.0. That is not a trend, it is a line. **Arabic breaks the pattern, and is also the hardest.** SIT-TCD leads Arabic on all four columns at once: 0.132 WER, 0.050 CER, 0.786 speaker, 0.997 prosody. The Arabic rank correlation between WER and speaker similarity is -0.4 over four systems. Arabic is still where the errors live: the best Arabic WER, 0.132, is worse than the worst French WER except one, and three times the best Chinese result. **The backbone decides the content column.** Every system that put Qwen3-TTS 1.7B behind Chinese or French landed under 0.052 WER. The 0.6B version of the same family, used by Langswap, landed at 0.189 and 0.133. Parameter count inside one model family moved WER by a factor of four. _Up and to the right means a more recognisable voice reading a less accurate script. Sources: IWSLT 2026 findings, Tables 13 and 14; correlations computed by BLOMEGA, code below._ ## The tradeoff is written into the selection objectives the teams chose This is not an emergent property of neural speech. Three of the five teams published the exact rule they used to pick which of several generated takes to submit, and the rule predicts where they landed. _SIT-TCD wrote the tradeoff into its scoring function at 50/50 and finished at the identity end in all three languages. Source: IWSLT 2026 findings, Track VIII system descriptions._ HW-TSC is the instructive failure. It generated many takes and kept the one with the highest timbre cosine similarity against three random segments of the reference. It optimised for exactly the quantity it was about to be graded on, and finished 4th of 5 in Chinese speaker similarity at 0.580. The selection ran against its own timbre feature vectors; the grading ran against ECAPA-TDNN embeddings in SpeechBrain. Optimising against your own embedder does not transfer to someone else's. **Prosody similarity carried no information.** Across all 14 results it ran from 0.980 to 0.997, a total spread of 0.017. Speaker similarity over the same 14 ran 0.479 to 0.813, a spread of 0.334, roughly twenty times wider. The prosody feature vector is mean and standard deviation of fundamental frequency plus mean energy, three summary statistics of a five-minute recording, compared by cosine similarity. Any two speech signals of similar loudness and register will score above 0.98 on that. If you are writing an acceptance test for a dubbing vendor, do not put this metric in it. _The prosody bar is drawn to scale. Source: IWSLT 2026 findings, Table 14._ ## What this changes for anyone buying or building AI dubbing **Specify both numbers or you will be sold one.** A demo clip that sounds uncannily like the original actor is a speaker-similarity demo. A demo that reads a technical script flawlessly is a WER demo. On the 2026 evidence, in Chinese, those were opposite ends of the same ranking. Write both into acceptance criteria with thresholds, and state which one loses when they conflict. **Pick the language before you pick the vendor.** Best Arabic WER in the track was 0.132. Best French was 0.050. Same teams, same models, same reference audio. An Arabic dubbing plan priced off a French pilot is priced wrong, and the error shows up as human correction hours. **Content errors in cloned speech are invisible to the listener who needs them most.** A 0.191 WER in a cloned Chinese voice is not a robotic tone, it is words that were not in the script, delivered in a familiar voice at conversational confidence. That is the failure mode to review for, and it is the reason a native-speaker QA pass is not optional at these error rates. **Fine-tuning moves identity; the backbone moves accuracy.** The two fine-tuned systems took the top speaker-similarity slot in all three languages between them. The lowest error rates came from zero-shot use of a larger backbone. If you want both, that suggests fine-tuning a large backbone rather than choosing between the two, which no 2026 submission did. **Get the release before you get the reference audio.** Five minutes of a person's voice is now enough for five separate teams to produce their voice in three languages they may not speak. The consent artefact needs to name the languages, the uses and the retention period, because the technical constraint that used to limit the use has gone. ## Check it yourself The tables are in one open PDF, the models are on Hugging Face, and the correlations are four lines of Python. ``` # 1. the source tables (Track VIII, Tables 13 and 14) curl -sL -o iwslt2026.pdf https://aclanthology.org/2026.iwslt-1.39.pdf python3 -c " import pypdf t='\n'.join(p.extract_text() or '' for p in pypdf.PdfReader('iwslt2026.pdf').pages) i=t.find('Track VIII'); print(t[i:i+14000])" # 2. the correlations quoted above python3 - <<'PY' from scipy.stats import spearmanr zh = [("IIT",0.042,0.499),("HW-TSC",0.043,0.580),("KIT",0.113,0.609), ("Langswap",0.189,0.686),("SIT-TCD",0.191,0.789)] fr = [("HW-TSC",0.050,0.582),("IIT",0.051,0.479),("KIT",0.063,0.602), ("SIT-TCD",0.069,0.813),("Langswap",0.133,0.761)] ar = [("SIT-TCD",0.132,0.786),("KIT",0.157,0.637),("Langswap",0.160,0.785), ("IIT",0.219,0.669)] for name, rows in [("Chinese",zh),("French",fr),("Arabic",ar)]: r = spearmanr([x[1] for x in rows],[x[2] for x in rows]) print(f"{name:8s} n={len(rows)} rho={r.statistic:+.2f}") allsim = [x[2] for x in zh+fr+ar] prosody = [.989,.989,.981,.989,.993,.987,.987,.980,.996,.991,.997,.982,.996,.990] print("speaker range", round(max(allsim)-min(allsim),3), " prosody range", round(max(prosody)-min(prosody),3)) PY # Chinese n=5 rho=+1.00 # French n=5 rho=+0.80 # Arabic n=4 rho=-0.40 # speaker range 0.334 prosody range 0.017 # 3. the models, all public # https://hf.co/Qwen/Qwen3-TTS-12Hz-1.7B-Base https://hf.co/k2-fsa/OmniVoice # https://hf.co/openbmb/VoxCPM2 https://hf.co/fishaudio/s2-pro https://hf.co/ResembleAI/chatterbox # 4. the recommended in-domain data # https://hf.co/datasets/ymoslem/acl-6060 (ACL 60/60, translations of ACL 2022 talks) ``` To replicate the evaluation on your own content: transcribe the generated audio with faster-whisper large-v3 at beam size 5, normalise with NFKC plus lowercasing plus punctuation removal (and jieba for Chinese), score WER and CER with Hugging Face `evaluate`, and take cosine similarity of SpeechBrain ECAPA-TDNN embeddings against the reference. That is the whole protocol. ## What would prove this wrong The claim under test is that in cross-lingual voice cloning, speaker similarity and content accuracy are currently in tension, and that the tension is a property of the systems rather than of one benchmark. It is wrong if, at IWSLT 2027 or in a peer-reviewed evaluation using the same protocol, a single system posts a Chinese WER at or below **0.060** together with an ECAPA-TDNN speaker similarity at or above **0.780**. In 2026 the best on each axis separately were 0.042 and 0.789, and they were different systems at opposite ends of the ranking. SIT-TCD's Arabic result, best on all four columns, is the existence proof that the tension can break, so a second edition could settle this quickly. A second prediction, marked as judgement: the 2027 edition will drop or replace the prosody similarity metric, because a metric with a 0.017 observed range across 14 systems cannot rank anything. The organisers' own future-work section names code-switching and terminology consistency as additions and does not mention prosody. If prosody similarity is reported unchanged in 2027, this reading was wrong. ## FAQ ### Is there really a tradeoff between clone accuracy and speaker identity? In IWSLT 2026, yes. Ranking the five Chinese submissions by WER and by speaker similarity gives the identical order, a Spearman correlation of 1.0. French is 0.8. Arabic is the exception at -0.4, where SIT-TCD led on both. ### Which target language is hardest? Arabic. Best Arabic WER was 0.132, against 0.042 for Chinese and 0.050 for French. French was the most tractable, with most systems under 0.07 WER and 0.02 CER. ### Does prosody similarity tell you anything? Not as measured here. It ran 0.980 to 0.997 across all 14 results, a spread of 0.017, against 0.334 for speaker similarity. It is a cosine over three summary statistics and does not separate systems. ### Does fine-tuning help cross-lingual voice cloning? For identity, clearly. The two fine-tuned systems, SIT-TCD (LoRA) and KIT (GRPO), took the top speaker-similarity result in three and one of three languages. The lowest error rates came from zero-shot use of Qwen3-TTS 1.7B. ## Sources - IWSLT 2026 organisers (60 authors), [Speech Translation and Metrics in 2026: Findings of the IWSLT Campaign](https://aclanthology.org/2026.iwslt-1.39/), Proceedings of the 23rd International Conference on Spoken Language Translation, San Diego, July 2026. Track VIII (Cross-Lingual Voice Cloning): task setup, evaluation data, metrics, system descriptions, Table 13 (WER and CER) and Table 14 (speaker and prosody similarity). [PDF](https://aclanthology.org/2026.iwslt-1.39.pdf). - Elizabeth Salesky et al., ACL 60/60 evaluation set, [ymoslem/acl-6060](https://huggingface.co/datasets/ymoslem/acl-6060) on Hugging Face. The in-domain development data recommended by the organisers, built from ACL 2022 presentations. - Model cards for the backbones named in the system descriptions: [Qwen3-TTS 1.7B](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-Base), [OmniVoice](https://huggingface.co/k2-fsa/OmniVoice), [VoxCPM2](https://huggingface.co/openbmb/VoxCPM2), [Fish Audio S2 Pro](https://huggingface.co/fishaudio/s2-pro), [Chatterbox](https://huggingface.co/ResembleAI/chatterbox). - Evaluation components: [faster-whisper](https://github.com/SYSTRAN/faster-whisper) (Whisper large-v3, beam 5), [SpeechBrain](https://speechbrain.github.io/) ECAPA-TDNN speaker embeddings, [Praat-Parselmouth](https://github.com/YannickJadoul/Parselmouth) for F0 and energy, [Hugging Face evaluate](https://github.com/huggingface/evaluate) for WER and CER. - Correlation and range calculations in this note computed by BLOMEGA on 16 September 2026 from Tables 13 and 14. Code in "Check it yourself". Related BLOMEGA guides: [EU AI Act Article 50 and AI dubbing disclosure](https://blomega.com/guides/eu-ai-act-article-50-ai-dubbing-watermarks/) · [The cost of a dubbed minute in 2026](https://blomega.com/guides/cost-of-a-dubbed-minute-2026/) · [Consented AI training data providers](https://blomega.com/guides/consented-ai-training-data-providers/). ============================================================================== URL: https://blomega.com/guides/japanese-subtitle-scores-segmentation-artifact-2026/ Published: 2026-09-16 | Updated: 2026-09-16 ============================================================================== --- title: "Japanese subtitles scored 12.19 chrF at IWSLT 2026. Re-scored without the word segmenter, 28.18" url: https://blomega.com/guides/japanese-subtitle-scores-segmentation-artifact-2026/ published: 2026-09-16 updated: 2026-09-16 source: BLOMEGA (https://blomega.com/) --- # Japanese subtitles scored 12.19 chrF at IWSLT 2026. Re-scored without the word segmenter, 28.18 Lab note · 16 September 2026 · BLOMEGA AppTek's primary Japanese subtitles on the IWSLT 2026 ITV test set scored 12.19 chrF against 28.14 for the same team's Chinese, a 16-point gap that looks like a translation failure. The organisers re-scored the same files as single long character strings, bypassing Japanese word-level re-segmentation, and chrF rose to **28.18** for Japanese and 35.52 for Chinese. The gap fell from 16 points to **7**. Separately, the human reference Japanese subtitles in that data meet Japan's 4 characters-per-second reading-speed rule **44.89%** of the time, while the machine output met it 91.88% of the time. Two of the numbers people quote about CJK subtitling measure the measuring apparatus. ## The 2026 subtitling track added Japanese, Chinese and a YouTube domain **July 2026, San Diego.** The IWSLT automatic subtitling track, running since 2023, expanded to five target languages (Arabic, Chinese, German, Japanese, Spanish) from English audio-visual content across three domains: ITV entertainment series, Asharq Business with Bloomberg economic news, and audio from the YODAS YouTube dataset, which is new this year. Three teams entered: AppTek, the MT unit of Fondazione Bruno Kessler (FBK), and Huawei Translation Service Center (HW-TSC). Results are in [Speech Translation and Metrics in 2026](https://aclanthology.org/2026.iwslt-1.39/), Track IV, Tables 31 to 33. The constraints came from published industry style guides, not from the organisers' preference. Maximum reading speed of 21 characters per second for Arabic, German and Spanish from [TED's subtitling tips](https://www.ted.com/participate/translate/subtitling-tips), 4 for Japanese and 9 for Chinese from Netflix's [Japanese](https://partnerhelp.netflixstudios.com/hc/en-us/articles/215767517-Japanese-Timed-Text-Style-Guide) and [Chinese Simplified](https://partnerhelp.netflixstudios.com/hc/en-us/articles/215986007-Chinese-Simplified-Timed-Text-Style-Guide) timed text style guides. Line length 42, 13 and 16 characters respectively, whitespace included. Two lines per subtitle, everywhere. Half-width characters count as 0.5 in Japanese. Scoring is three-sided. Subtitle quality is SubER, the primary ranking metric. Translation quality is BLEU, chrF and BLEURT. Compliance is measured as the rate of subtitles inside the CPS limit, the rate of lines inside the CPL limit, and the rate of subtitles at or under 2 lines. Before any metric is computed, automatic subtitles are realigned to the reference with `mweralign`, a variant of the AS-WER algorithm. That realignment step is where the Japanese problem starts. ## The published Japanese numbers and the re-scored ones Table 1 is the 2026 test set for the ITV entertainment domain, primary systems only, one row per language, plus the organisers' own re-scoring of the two CJK rows. | Target | Team | SubER | BLEU | chrF | chrF, re-scored | BLEURT | CPS | CPL | LPB | Source | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | Japanese | AppTek | 78.81 | 7.61 | **12.19** | **28.18** | .2521 | 91.88 | 100.00 | 99.86 | Table 33 + Track IV text | | Chinese | AppTek | 57.13 | 33.27 | 28.14 | 35.52 | .5681 | 99.53 | 100.00 | 99.26 | Table 33 + Track IV text | | German | AppTek | 74.01 | 18.62 | 47.83 | not reported | .5408 | 90.48 | 100.00 | 94.67 | Table 33 | | Spanish | AppTek | 62.86 | 23.34 | 48.86 | not reported | .5735 | 94.06 | 100.00 | 97.84 | Table 33 | | Japanese | FBK | 91.44 | 7.21 | 14.06 | not reported | .2845 | 61.60 | 73.89 | 100.00 | Table 33 | | Chinese | FBK | 64.78 | 24.94 | 21.47 | not reported | .4931 | 96.57 | 93.38 | 100.00 | Table 33 | | Chinese | HW-TSC | 70.87 | 24.18 | 21.77 | not reported | .4978 | 84.22 | 99.89 | 100.00 | Table 33 | | Japanese (human reference) | ITV dev26 | n/a | n/a | n/a | n/a | n/a | **44.89** | not reported | not reported | Track IV text | Read the last row against the first. The human reference Japanese subtitles comply with the 4 characters-per-second limit 44.89% of the time. The machine complies 91.88% of the time. The benchmark's own ground truth violates the benchmark's own rule more than half the time, and a system is penalised on SubER and BLEURT for matching the rule rather than the reference. The organisers did the arithmetic on the relaxation too. Moving the Japanese threshold from 4 characters per second to 6 raises the reference compliance rate from 44.89% to **95.74%**. Two characters per second is the difference between a rule almost nobody follows and a rule almost everybody follows. _The grey bar is the ground truth. Four of the six system runs beat it. Source: IWSLT 2026 findings, Track IV text and Table 33._ ## The Japanese penalty is applied twice, and only once on purpose Automatic subtitles do not line up with reference subtitles block for block, so the scorer realigns them first. For languages with spaces, realignment operates on tokens the text already provides. For Japanese it has to invent them. _The orange box is a preprocessing step, not a property of the translation. Source: IWSLT 2026 findings, Track IV, Results section._ The organisers are explicit about the logic: chrF is sensitive to word-level segmentation during hypothesis re-segmentation but not during evaluation itself, so comparing the two scoring paths isolates the segmentation effect. Chinese still gained 7.38 points, because realignment still ran, but Chinese evaluation is character-level and therefore largely insulated. Japanese gained 15.99. The paper's conclusion, in its own words: "the underlying translation quality is not as divergent as the official scores would indicate." This does not make the Japanese output good. BLEURT for AppTek Japanese is .2521 against .5681 for Chinese and .5735 for Spanish, and BLEURT does not run through `mweralign` word tokens the same way. The honest reading is that roughly half of the visible Japanese deficit is instrumentation, and the rest is real. ## On a frozen test set, translation quality really did move The track keeps re-scoring old submissions on old test sets, which makes a clean four-year series possible. ITV tst23, German and Spanish, primary runs only. _Same audio, same references, four years of systems. Source: IWSLT 2026 findings, Table 33 (legacy tst23 rows re-evaluated)._ The compliance columns tell the opposite story. On German tst26, AppTek's contrastive run hit 99.26% reading-speed compliance and the primary run hit 90.48%, and the primary run scored better BLEURT. On the Asharq-Bloomberg German section, the organisers note AppTek's 2026 system improved BLEURT from .6020 to .6234 while reading-speed compliance fell from 92.44 to 73.50. Their summary: the optimal tradeoff between translation quality and subtitling spatiotemporal constraints "has yet to be conclusively resolved." The 2026 systems reveal how that tradeoff is now negotiated. AppTek runs a hybrid production ASR plus Transformer Big MT with genre, speaker gender and length-class metadata, uses its Intelligent Line Segmentation to place breaks, iteratively shortens translations with the length parameter until CPS, CPL and LPB are all satisfied, and post-edits with gpt-4o in roughly 20-sentence chunks. For Japanese and some German runs it _disables_ strict CPS enforcement, explicitly because it avoids a large drop in MT metrics and matches the low compliance of the human references. FBK runs an entirely open stack: SpeechBrain VAD, Whisper large-v3, MADLAD-400-10B-MT, with a second pass that re-transcribes aggregated segments using Voxtral's 32k-token context and realigns via `mweralign`. HW-TSC compresses non-compliant blocks with Qwen3-32B in two passes, greedy at temperature 0 removing auxiliaries and conjunctions, then temperature 0.3 for deeper compression, always retaining proper nouns. ## What to do with this if you buy subtitling **Never accept a cross-language metric comparison in CJK without asking how it was segmented.** A vendor reporting "chrF 12 in Japanese versus 28 in Chinese" may be reporting the same translation quality through two different scorers. Ask for the tokenisation, and ask for a character-level re-score alongside. **Write the reading-speed limit you will actually enforce into the brief.** If your reference material is 44.89% compliant at 4 cps and 95.74% at 6, then 4 is an aspiration and 6 is your operating rule. Say which one the vendor is being scored against, because at 4 cps a system that complies is punished by every reference-based metric you are also using. **Score compliance and quality separately and weight them yourself.** Every 2026 result shows them trading against each other: AppTek turned CPS enforcement off in Japanese to protect its MT scores, and lost 18.94 points of Asharq German compliance to gain 0.0214 BLEURT. That is your decision to make, not your vendor's. **An open stack is now within a few points of a production stack.** On German tst26 ITV, FBK's fully open pipeline scored SubER 75.51 and BLEURT .5252 against AppTek's 74.01 and .5408, and FBK beat AppTek outright on German in the Asharq-Bloomberg domain. If your constraint is that no audio leaves your infrastructure, the penalty for that in 2026 is small and measurable. **Budget separately for Japanese.** Even after correcting for the segmenter, AppTek's Japanese BLEURT of .2521 sits less than half of its Spanish .5735. Japanese needs the reviewer hours, and the 13-character line limit with half-width counting is a real production constraint that no length-control parameter fully absorbs. ## Check it yourself ``` # 1. the tables and the re-scoring paragraph (Track IV) curl -sL -o iwslt2026.pdf https://aclanthology.org/2026.iwslt-1.39.pdf python3 -c " import pypdf t='\n'.join(p.extract_text() or '' for p in pypdf.PdfReader('iwslt2026.pdf').pages) i=t.find('Track IV Subtitling'); print(t[i:i+9000])" # task, thresholds, systems python3 -c " import pypdf t='\n'.join(p.extract_text() or '' for p in pypdf.PdfReader('iwslt2026.pdf').pages) i=t.find('Table 33:'); print(t[i-4200:i+400])" # ITV results, all languages # 2. the metrics, as code pip install subER sacrebleu # SubER: https://github.com/apptek/SubER # compliance: https://github.com/hlt-mt/FBK-fairseq/blob/master/examples/ # speech_to_text/scripts/subtitle_compliance.py # realigner: mweralign, a variant of AS-WER (Matusov et al., 2005) # 3. reproduce the segmentation effect on your own Japanese SRT pair # score once with tok=ja-mecab, once with the whole file as a single # character string, and compare chrF. The organisers' delta was +15.99. sacrebleu ref.ja.txt -i hyp.ja.txt -m chrf --tokenize ja-mecab python3 -c " import sacrebleu r=open('ref.ja.txt').read().replace('\n','') h=open('hyp.ja.txt').read().replace('\n','') print(sacrebleu.sentence_chrf(h,[r]).score)" # 4. the reading-speed arithmetic in this note python3 - <<'PY' ref = {"4 cps": 44.89, "6 cps": 95.74} sys = {"AppTek prmry": 91.88, "AppTek cntrs1": 75.78, "FBK prmry": 61.60, "FBK cntrs1": 56.63, "FBK cntrs2": 23.16} print("relaxing 4 cps to 6 cps moves the human reference by", round(ref["6 cps"]-ref["4 cps"],2), "points") beat = [k for k,v in sys.items() if v > ref["4 cps"]] print(len(beat), "of", len(sys), "system runs are more compliant than the reference:", beat) PY # relaxing 4 cps to 6 cps moves the human reference by 50.85 points # 4 of 5 system runs are more compliant than the reference: ['AppTek prmry', 'AppTek cntrs1', 'FBK prmry', 'FBK cntrs1'] ``` The style guides the thresholds come from are public: [TED subtitling tips](https://www.ted.com/participate/translate/subtitling-tips), Netflix [Japanese](https://partnerhelp.netflixstudios.com/hc/en-us/articles/215767517-Japanese-Timed-Text-Style-Guide) and [Chinese Simplified](https://partnerhelp.netflixstudios.com/hc/en-us/articles/215986007-Chinese-Simplified-Timed-Text-Style-Guide) timed text guides. Check yours against them before you set an acceptance threshold. ## What would prove this wrong The claim under test is that a large share of the Japanese subtitling deficit on reference-based metrics is produced by word-level re-segmentation rather than by translation. It is wrong if a controlled re-scoring of the IWSLT 2026 Japanese submissions, published by **31 December 2027**, shows a character-level chrF within 2 points of the segmenter-based chrF, that is a delta under 2 rather than the +15.99 the organisers measured. It is also wrong if BLEURT, which does not depend on the same tokenisation, moves by a comparable amount under a character-level realignment; the organisers reported the chrF delta only, so this half is untested. A second prediction, marked as judgement: the 2027 subtitling track will either relax the Japanese reading-speed threshold above 4 characters per second or report reference compliance alongside system compliance, because keeping a rule that the ground truth breaks 55% of the time makes the compliance column unusable. If the 2027 task page still specifies 4 cps with no reference baseline published, this reading was wrong. ## FAQ ### Why do Japanese subtitles score so badly on automatic metrics? Mostly word segmentation. Japanese has no whitespace, so the scorer must segment before aligning. Re-scoring AppTek's ITV Japanese as character strings moved chrF from 12.19 to 28.18. Chinese, evaluated at character level, moved from 28.14 to 35.52. The gap fell from 16 points to 7. ### What is the maximum subtitle reading speed for Japanese? IWSLT 2026 used 4 characters per second, half-width counted as 0.5, from Netflix's Japanese timed text style guide, against 21 for Arabic, German and Spanish and 9 for Chinese. Line length was 13 characters for Japanese, 16 for Chinese, 42 for the rest, and 2 lines maximum everywhere. ### Do human subtitlers meet the 4 cps limit? In this data, 44.89% of the time on the ITV development references, rising to 95.74% at a 6 cps threshold. AppTek's primary Japanese system complied 91.88% of the time. ### Has automatic subtitling improved since 2023? On translation quality, yes, measured on a frozen test set: best primary BLEURT on ITV tst23 rose from .4438 to .5520 for German and .4530 to .5514 for Spanish. SubER and compliance were volatile and peak values were not always the most recent. ## Sources - IWSLT 2026 organisers (60 authors), [Speech Translation and Metrics in 2026: Findings of the IWSLT Campaign](https://aclanthology.org/2026.iwslt-1.39/), Proceedings of the 23rd International Conference on Spoken Language Translation, San Diego, July 2026. Track IV (Subtitling): task description and thresholds, Table 8 (set statistics), metrics, system descriptions, Results (Japanese re-scoring, reference CPS, year-over-year trend), Tables 31 to 33. [PDF](https://aclanthology.org/2026.iwslt-1.39.pdf). - Mauro Cettolo, Roldano Cattoni, Matteo Negri, Luisa Bentivogli, [The FBK Sentence-Aware Subtitling System at the IWSLT 2026 Subtitling Track](https://aclanthology.org/2026.iwslt-1.7/), IWSLT 2026, pages 68 to 77. The two-stage open-source pipeline described above. - HW-TSC, [HW-TSC's Submission to the IWSLT 2026 Subtitling Track](https://aclanthology.org/2026.iwslt-1.10/), IWSLT 2026. Qwen3 streaming ASR plus Qwen3-32B two-pass subtitle compression. - Style guides the thresholds are taken from: [TED subtitling tips](https://www.ted.com/participate/translate/subtitling-tips); Netflix [Japanese Timed Text Style Guide](https://partnerhelp.netflixstudios.com/hc/en-us/articles/215767517-Japanese-Timed-Text-Style-Guide); Netflix [Chinese Simplified Timed Text Style Guide](https://partnerhelp.netflixstudios.com/hc/en-us/articles/215986007-Chinese-Simplified-Timed-Text-Style-Guide). - Metric implementations: [SubER](https://github.com/apptek/SubER) (Wilken et al., 2022); the [FBK-fairseq](https://github.com/hlt-mt/FBK-fairseq) subtitle compliance script; sacreBLEU (Post, 2018) for BLEU and chrF; BLEURT (Sellam et al., 2020). - Compliance arithmetic in this note computed by BLOMEGA on 16 September 2026 from the Track IV figures. Code in "Check it yourself". Related BLOMEGA guides: [Your multilingual LLM judge prefers the machine translation](https://blomega.com/guides/multilingual-llm-judge-translationese-bias/) · [The cost of a dubbed minute in 2026](https://blomega.com/guides/cost-of-a-dubbed-minute-2026/) · [Localization is the new default](https://blomega.com/guides/localization-the-new-default/). ============================================================================== URL: https://blomega.com/guides/live-dubbing-latency-iwslt-2026/ Published: 2026-09-16 | Updated: 2026-09-16 ============================================================================== --- title: "The first-placed system in IWSLT 2026's 2-to-4 second latency class ran at 52.7 seconds on YouTube audio" url: https://blomega.com/guides/live-dubbing-latency-iwslt-2026/ published: 2026-09-16 updated: 2026-09-16 source: BLOMEGA (https://blomega.com/) --- # The first-placed system in IWSLT 2026's 2-to-4 second latency class ran at 52.7 seconds on YouTube audio Lab note · 16 September 2026 · BLOMEGA IWSLT 2026 sorted simultaneous speech translation systems into a low-latency class (0 to 2 seconds) and a high-latency class (2 to 4 seconds) using their development-set numbers. On the English-to-Chinese test sets, the system that placed first in the high-latency class measured **28.3 seconds** of LongYAAL on ACL talks, **36.3** on Bloomberg business news and **52.7** on YouTube-sourced YODAS audio. The regime label held on the development set and broke everywhere else. If you are buying live dubbing this quarter, the number in the contract should be measured on your audio, not on a conference recording. ## Two things landed in the same month, and only one of them published latencies **11 September 2026, Amsterdam.** At IBC 2026, CAMB.AI announced a Streaming SDK and Streaming Dashboard for live multilingual dubbing and subtitles, integrated into [NVIDIA's Holoscan for Media](https://blogs.nvidia.com/blog/ibc-news-2026/) reference architecture, built on its MARS and BOLI models, supporting 150 or more languages, with Ligue 1+, NASCAR and Eurovision Sport named as users and a claim that a broadcaster can start multilingual output in as little as three lines of code ([Sports Video Group](https://www.sportsvideo.org/2026/09/11/ibc-2026-camb-ai-unveils-the-worlds-first-multilingual-broadcasting-agent-powered-by-nvidia/), [TV Tech](https://www.tvtechnology.com/platform/broadcast/camb-ai-unveils-multilingual-broadcasting-agent-powered-by-nvidia-at-ibc2026)). In the same week NVIDIA described NDI using its LipSync NIM and Active Speaker Detection NIM microservices for real-time translation and lip-synced dubbing inside existing broadcast workflows. Neither announcement carries an end-to-end latency figure. The published latency number on CAMB.AI's own site is a different measurement: its [low-latency streaming post](https://www.camb.ai/blog-post/lowest-latency-tts-for-live-sports-streaming-setups) puts MARS8-Flash at roughly 100 milliseconds time-to-first-byte. That is how fast the speech synthesiser starts talking once it has been handed text. It says nothing about how long the translation policy waited before handing over that text. **July 2026, San Diego.** The 23rd IWSLT evaluation campaign did measure that wait, across 10 shared tasks and more than 30 teams, and published every number. The findings paper is [Speech Translation and Metrics in 2026](https://aclanthology.org/2026.iwslt-1.39/), 87 pages, 60 authors, open access. Its Simultaneous track ran on raw unsegmented audio in four directions: English into German, Chinese and Italian, and Czech into English. Quality is XCOMET-XL, latency is LongYAAL, both computed by [OmniSTEval](https://github.com/pe-trik/OmniSTEval) on re-segmented output. Ten systems from eight teams submitted. ## The latency class was assigned on the development set and did not survive the test sets Participants submitted development-set logs, the organisers computed LongYAAL, and that number decided which class a system competed in. Table 1 below takes the first-placed English-to-Chinese system in the high-latency class and follows it across every evaluation set, then shows what the systems it beat were doing at the same time. | Evaluation set | Pl. | System | COMET | BLEU | chrF | LongYAAL (s) | Source | | --- | --- | --- | --- | --- | --- | --- | --- | | MCIF dev | 1 | NEMO | 0.84 | 47.48 | 41.01 | 15.4 | Table 36 | | ACL test | 1 | NEMO | 0.83 | 47.64 | 40.30 | **28.3** | Table 36 | | Bloomberg test | 1 | NEMO | 0.75 | 30.46 | 27.05 | **36.3** | Table 36 | | YODAS test | 1 | NEMO | 0.68 | 22.50 | 21.06 | **52.7** | Table 36 | | ACL test | 2 | MLLP-VRAIN UPV | 0.82 | 50.03 | 42.75 | 4.8 [5.2] | Table 36 | | ACL test | 3 | CPII-HK | 0.78 | 45.04 | 40.46 | 3.0 | Table 36 | | ACL test | 5 | CUHKSZ | 0.74 | 42.57 | 35.82 | 2.3 | Table 36 | | YODAS test | 2 | Baseline | 0.65 | 22.57 | 21.25 | 1.0 [1.2] | Table 36 | | ACL test (En-De) | 1 | NEMO | 0.93 | 45.01 | 71.44 | 4.7 | Table 35 | | YODAS test (En-De) | 1 | NEMO | 0.79 | 29.57 | 56.73 | 6.0 | Table 35 | Read the last two rows first. The same team, the same campaign, English into German: 4.7 seconds on ACL talks, 6.0 on YODAS. Slow for a 2-to-4 second class, but recognisable. The blowup is specific to English into Chinese. Look at the YODAS row for the Baseline: COMET 0.65 at 1.0 second, against the first-placed system's 0.68 at 52.7 seconds. Three hundredths of COMET for fifty-one seconds. On a live stream that is not a tradeoff, it is a different product. The findings paper acknowledges the pattern in one sentence: "Due to the aforementioned latency measurement issues, several systems exhibit higher latencies on the test sets than their development set metrics initially suggested at the time of submission, explaining the unusually large latency values reported in the results tables." The word "aforementioned" is the paper's only occurrence of it in that section, and the Simultaneous track write-up does not describe the issues it refers back to. So the published numbers stand without an explanation attached to them. _Every bar is the same system in the same declared latency class. Source: IWSLT 2026 findings, Table 36._ ## A 100 millisecond time-to-first-byte measures the last box in the chain A live dubbing pipeline has at least five places where seconds accumulate, and the vendor metrics in circulation cover one of them. The diagram below puts the IWSLT 2026 measurements next to the marketing measurement, on the same timeline. _The orange span is the part a translation vendor controls with its policy. The teal span is the part a time-to-first-byte figure describes. Sources: IWSLT 2026 findings, Track V; CAMB.AI low-latency streaming post._ _The same team supplies both frontier points. Source: IWSLT 2026 findings, Table 35, ACL test set, both latency regimes._ Two of the campaign's smaller results explain why the orange span is hard to shrink honestly. **The smallest model paid the largest compute penalty.** CUNI-POCKET was the only end-to-end submission, a single 1B-parameter Canary model aimed at edge deployment. It showed the biggest gap between computation-aware and non-computation-aware latency, more than 0.4 seconds, against a campaign average of about 0.3. The findings paper flags this as running against the expectation that smaller models incur lower overhead. Model size is not a proxy for responsiveness. **Extra context helps quality and costs nothing in latency, if you integrate it properly.** The 2026 edition added a track where systems could read the source paper PDF for each ACL talk. MLLP-VRAIN UPV gained an average of 2.75 COMET across three language directions and both latency regimes. NEMO gained roughly nothing over its context-free run. CUHKSZ went slightly backwards on English to Chinese. Same input, three outcomes, which puts the value in the integration rather than the context. ## What a localization buyer should put in the contract **Ask for the metric, not the adjective.** "Real-time" is not a number. LongYAAL, StreamLAAL and LongDAL are numbers, they are defined in public, and [OmniSTEval](https://github.com/pe-trik/OmniSTEval) computes them. Ask a vendor which one they report, on which audio, and whether it is computation-aware. The IWSLT 2026 tables print the computation-aware value in brackets next to the unaware one precisely because the two differ. **Require the measurement on your own content.** The 1.0 second YODAS penalty is the whole argument. A system tuned on conference talks met its declared class on conference talks. The same system on YouTube-style audio was slower and worse: English to Chinese COMET fell from 0.83 to 0.68 and BLEU from 47.64 to 22.50. Sports commentary, reality formats and user-generated video are further from ACL talks than YODAS is. **Treat quality metrics as disagreeing witnesses.** On English to Chinese, MLLP-VRAIN UPV beat the higher-ranked NEMO submission by an average of 1.6 chrF and 1.7 BLEU while scoring 0.067 lower on COMET. If your acceptance test uses one metric and your vendor optimised another, you will argue about a real disagreement rather than a rounding error. **Budget the latency you can hide.** A pre-recorded stream can absorb a 30 second delay behind a broadcast delay buffer. A live sports call cannot: the crowd noise arrives 30 seconds before the commentary that explains it. That difference, not model quality, is what decides whether the current state of simultaneous translation is usable for a given format. This is a judgement, not a measurement. **Nine of ten systems were cascades.** The findings paper states that with the exception of a single end-to-end submission, all participating teams adopted cascaded approaches: separate ASR, then MT, then emission policy. Vendors selling a single end-to-end model as inherently faster are selling against the way the field's own best systems are actually built in 2026. _CAMB.AI CTO Akshat Prakash on real-time multilingual commentary for live sports, posted by [@sportsvideogroup5320](https://www.youtube.com/@sportsvideogroup5320) on 26 February 2026. Supports the claim that live sports commentary is the target format for the vendors shipping streaming dubbing SDKs in 2026._ ## Check it yourself Every number above is in one open-access PDF and one open-source toolkit. ``` # 1. the findings paper (87 pages, open access, no paywall) curl -sL -o iwslt2026.pdf https://aclanthology.org/2026.iwslt-1.39.pdf python3 -c " import pypdf t='\n'.join(p.extract_text() or '' for p in pypdf.PdfReader('iwslt2026.pdf').pages) open('iwslt2026.txt','w').write(t) i=t.find('Table 36') print(t[i-6000:i+200])" # Table 36 = English to Chinese # Table 35 = English to German, Table 34 = Czech to English, Table 37 = English to Italian # 2. the latency metrics, as code git clone https://github.com/pe-trik/OmniSTEval # LongYAAL is the primary metric; StreamLAAL, LongLAAL and LongDAL are also reported # 3. the audio that produced the 1.0 second penalty # https://huggingface.co/datasets/espnet/yodas (partition en003 is held out from training) # 4. reproduce the regime-versus-reality gap in the table above python3 - <<'PY' # NEMO, English to Chinese, high-latency regime, LongYAAL seconds (Table 36) sets = {"MCIF dev":15.4, "ACL test":28.3, "Bloomberg test":36.3, "YODAS test":52.7} comet = {"MCIF dev":0.84, "ACL test":0.83, "Bloomberg test":0.75, "YODAS test":0.68} for k,v in sets.items(): print(f"{k:16s} {v:6.1f} s COMET {comet[k]:.2f} over declared 4 s ceiling by {v-4:.1f} s") PY # MCIF dev 15.4 s COMET 0.84 over declared 4 s ceiling by 11.4 s # ACL test 28.3 s COMET 0.83 over declared 4 s ceiling by 24.3 s # Bloomberg test 36.3 s COMET 0.75 over declared 4 s ceiling by 32.3 s # YODAS test 52.7 s COMET 0.68 over declared 4 s ceiling by 48.7 s ``` To test a vendor, hand over 30 minutes of your own audio, ask for the output log with per-token emission timestamps in SimulStream or SimulEval JSONL format, and run OmniSTEval yourself. A vendor who cannot produce emission timestamps cannot produce a latency number either. ## What would prove this wrong The claim under test is that the quality-latency frontier for unsegmented live speech translation sits well above one second, and that out-of-domain audio moves it further. It is wrong if, by **31 July 2027**, a system in the IWSLT 2027 Simultaneous track reports a computation-aware LongYAAL at or under 1.0 second on the YODAS test set while matching the 2026 best COMET on that set, which was 0.68 for English to Chinese and 0.79 for English to German. The closest 2026 result was the organisers' baseline at 1.0 second and 0.65 COMET, which trades 0.03 COMET for the speed and is therefore already close on one axis and not the other. A second prediction, marked as judgement: no live dubbing vendor will publish a LongYAAL or StreamLAAL figure on customer audio before 31 July 2027. Time-to-first-byte is the metric being marketed because it is the flattering one. If a vendor publishes an end-to-end emission-timestamp latency on non-ACL audio before that date, this reading was too cynical. ## FAQ ### How much latency does live AI dubbing actually add? On the IWSLT 2026 test sets, English to German ran between 2.0 and 7.7 seconds of non-computation-aware LongYAAL for the systems that behaved, and the first-placed English to Chinese system measured 28.3 seconds on ACL talks, 36.3 on Bloomberg news and 52.7 on YODAS. Those are text latencies. Synthesis, packaging and delivery sit on top. ### Does a 100 millisecond time-to-first-byte mean sub-second live dubbing? No. Time-to-first-byte measures how quickly the speech synthesiser starts emitting audio once it has text. It excludes the wait the read/write policy imposes before committing that text, which is what LongYAAL measures. In 2026 that wait ran from 1.5 seconds to over 50. ### Why is YouTube audio harder than conference talks? The findings paper reports that YODAS consistently induced the highest latencies across all language pairs, about 1.0 second higher than the other sets, and attributes it to domain mismatch and distinct acoustic characteristics. Quality fell alongside: best English to Chinese COMET dropped from 0.83 on ACL talks to 0.68 on YODAS. ### Are cascaded or end-to-end systems better for simultaneous speech translation? Cascaded, in practice. All 2026 submissions except one were cascades. The single end-to-end system, a 1B-parameter Canary model, showed the largest computation-aware penalty, over 0.4 seconds against a campaign average of about 0.3. ## Sources - IWSLT 2026 organisers (60 authors), [Speech Translation and Metrics in 2026: Findings of the IWSLT Campaign](https://aclanthology.org/2026.iwslt-1.39/), Proceedings of the 23rd International Conference on Spoken Language Translation, San Diego, July 2026. Track V (Simultaneous): latency regimes, metrics, YODAS latency spike, computation-aware gap, extra-context results; Tables 34 to 37 (per-direction results). [PDF](https://aclanthology.org/2026.iwslt-1.39.pdf). - Zeyu Yang and Satoshi Nakamura, [CUHKSZ Simultaneous Speech Translation System for IWSLT 2026](https://aclanthology.org/2026.iwslt-1.13/), IWSLT 2026. Qwen3-Omni-30B-A3B backbone with a learned wait token. - Javier Iranzo-Sanchez et al., [MLLP-VRAIN UPV System for the IWSLT 2026 Simultaneous Speech Translation Task](https://aclanthology.org/2026.iwslt-1.24/), IWSLT 2026. Parakeet plus Qwen 3.5 cascade with a relaxed longest-common-prefix policy. - Sports Video Group, [IBC 2026: CAMB.AI Unveils the World's First Multilingual Broadcasting Agent Powered by NVIDIA](https://www.sportsvideo.org/2026/09/11/ibc-2026-camb-ai-unveils-the-worlds-first-multilingual-broadcasting-agent-powered-by-nvidia/), 11 September 2026, and TV Tech, [CAMB.AI Unveils Multilingual Broadcasting AI Agent at IBC2026](https://www.tvtechnology.com/platform/broadcast/camb-ai-unveils-multilingual-broadcasting-agent-powered-by-nvidia-at-ibc2026), 11 September 2026. Streaming SDK and Dashboard, MARS and BOLI, 150+ languages, Ligue 1+, NASCAR, Eurovision Sport. - NVIDIA, [NVIDIA Brings Real-Time AI to Broadcast, Sports and Global Streaming at IBC](https://blogs.nvidia.com/blog/ibc-news-2026/), September 2026. LipSync NIM and Active Speaker Detection NIM in NDI workflows. - CAMB.AI, [Lowest-Latency TTS for Live Sports Streaming](https://www.camb.ai/blog-post/lowest-latency-tts-for-live-sports-streaming-setups). MARS8-Flash at roughly 100 milliseconds time-to-first-byte. - [OmniSTEval](https://github.com/pe-trik/OmniSTEval) (latency and quality toolkit used by the campaign) and [espnet/yodas](https://huggingface.co/datasets/espnet/yodas) (the YouTube-sourced evaluation audio). Related BLOMEGA guides: [Prime Video's visual dubbing](https://blomega.com/guides/prime-video-lip-sync-visual-dubbing-2026/) · [The cost of a dubbed minute in 2026](https://blomega.com/guides/cost-of-a-dubbed-minute-2026/) · [EU AI Act Article 50 and AI dubbing disclosure](https://blomega.com/guides/eu-ai-act-article-50-ai-dubbing-watermarks/). ============================================================================== URL: https://blomega.com/guides/low-resource-speech-translation-hours-vs-bleu-2026/ Published: 2026-09-16 | Updated: 2026-09-16 ============================================================================== --- title: "130 hours of Mapuzugun bought 0.82 BLEU. 30 hours of Central Kurdish bought 21.09" url: https://blomega.com/guides/low-resource-speech-translation-hours-vs-bleu-2026/ published: 2026-09-16 updated: 2026-09-16 source: BLOMEGA (https://blomega.com/) --- # 130 hours of Mapuzugun bought 0.82 BLEU. 30 hours of Central Kurdish bought 21.09 Lab note · 16 September 2026 · BLOMEGA Across the five low-resource pairs in the IWSLT 2026 speech translation task that have published results, the Spearman rank correlation between hours of provided speech and best BLEU is **exactly 0.00**. Mapuzugun-Spanish came with more than 130 hours of transcribed and translated speech and produced 0.82 BLEU. Central Kurdish-English came with 30 hours and produced 21.09. Irish-English came with 13 real hours plus 196 synthetic ones and produced 2.4. If you are budgeting a data collection for a new locale, hours is the wrong line item to argue about. ## Ten language pairs, one campaign, every number published **July 2026, San Diego.** The IWSLT low-resource speech translation track ran 10 typologically diverse pairs plus a data track inviting new open-sourced corpora. This year's stated focus was explicitly multilingual systems handling as many languages as possible. Results are in [Speech Translation and Metrics in 2026](https://aclanthology.org/2026.iwslt-1.39/), Track II, Tables 2 to 7, alongside the African and Celtic speech-to-speech track that shares its Hausa, Igbo and Yoruba data. What makes the track worth reading is that the organisers publish the corpus behind each pair, in hours, with its provenance, next to the score every team got on it. That pairing is rare. Most dataset announcements give you hours and no downstream number; most benchmark papers give you a number and no corpus description. ## The corpus, the speakers and the score, in one table | Pair | Speakers | Speech provided | Domain | Best BLEU | chrF++ | Best system | Source | | --- | --- | --- | --- | --- | --- | --- | --- | | Quechua to Spanish | >8 million | ~50 h transcribed + 8 h synthetic post-edited + 15 h Quechua Collao | mixed, multi-variant | **27.2** | 51.4 | QUESPA contrastive 2 (SpeechT5 end-to-end) | Table 4 | | Central Kurdish to English | ~8 million | 30 h COMMUTE-Kurdish | spontaneous Kurdish media | **21.09** | 49.48 | LIUM (pseudo-labelling) | Table 7 | | Bhojpuri to Hindi | 50.58 million | ~24 h | news (News On Air) | 14.7 | 43.0 | ADAPT-MTU primary (Whisper large-v3 + NLLB-200) | Table 3 | | Irish to English | ~170,000 L1 | ~13 h real + 196 h synthetic | news, Common Voice, Living-Audio-Dataset | 2.4 | 16.0 | MTU primary | Table 5 | | Mapuzugun to Spanish | 100,000 to 200,000 | >130 h | language isolate corpus | 0.82 | 14.31 | KK contrastive 2 | Table 6 | | Bemba to English | >10 million | 180 h + 28 h transcribed mono + 60 h untranscribed mono | image-grounded dialogues | not reported | not reported | no results table published | Track II section 2 | | Catalan to English | 4.1 million L1 | not stated | not stated | not reported | not reported | CATENG submitted, no table | Track II sections 2, 3 | | Hausa to English | >100 million | shared with African/Celtic track | newly collected | 18.6 spBLEU | 41.9 | SeamlessM4T mono fine-tuned (organiser baseline) | Table 2 | | Igbo to English | 30 to 45 million | shared with African/Celtic track | newly collected | 17.6 spBLEU | 39.2 | SeamlessM4T mono fine-tuned (organiser baseline) | Table 2 | | Yoruba to English | ~50 million | shared with African/Celtic track | newly collected | 21.1 spBLEU | 43.5 | SeamlessM4T mono fine-tuned (organiser baseline) | Table 2 | Note what does not predict the score. Speaker population does not: Bhojpuri has 50.58 million speakers and scores 14.7, Mapuzugun has at most 200,000 and scores 0.82, but Irish has fewer speakers still and also fails, while Central Kurdish has a comparable population to Quechua and both do well. Hours do not either, and that one is measurable. _Five pairs, five hour counts, no relationship. Source: IWSLT 2026 findings, Track II; correlation computed by BLOMEGA, code below._ ## Four things that did predict the score _Every claim on this diagram is a figure or a quotation from the IWSLT 2026 findings, Track II._ The Central Kurdish column is the sharpest illustration. Two teams, one 30-hour corpus. LIUM scored 21.09 BLEU and 49.48 chrF++ with a pseudo-labelling pipeline that produced silver translations for untranscribed audio through an automated ASR and MT chain, and reported 6.98 CER and 19.76 WER on the recognition side. SLC scored 0.16 BLEU on the same data. A factor of 131 between two systems on the same corpus says the corpus was not the limiting factor for either of them. ## No frontier model beat a fine-tuned NLLB-200 on Quechua The QUESPA team, in its fourth consecutive year on this pair, ran a separate text machine translation case study: GPT-5, Gemini 3, Claude, DeepSeek-V3 and Qwen, prompted in Spanish with guided prompts. The best prompt-based result was **10.8 BLEU**, from Gemini 3 Flash. The fine-tuned NLLB-200 baseline from the previous year sits at **19.5 BLEU and 23.5 chrF**. None of the frontier models passed it. The team's stated causes: hallucinations, dialectal confusion between Quechua variants, and a tendency of models to prioritise high-resource language signals over low-resource Quechua input. That third one is the structural problem. A model trained overwhelmingly on Spanish will read Quechua input through Spanish priors, and prompt engineering in Spanish reinforces exactly that. The same pattern shows on the speech side of the African track, where the organisers ran three baselines rather than accepting submissions as the reference. A monolingually fine-tuned SeamlessM4T beat a cascade of Omnilingual ASR (OmniASR LLM 1B) and NLLB-200 on all three languages, and also beat an end-to-end system built on a frozen Gemma-4-E2B language model on all three, though that system took second place on Yoruba ahead of the cascade. _Source: IWSLT 2026 findings, African/Celtic S2TT results. The table carries the caption "Table 2" and is referred to in the surrounding text as Table 12; we cite the caption._ ## The data track wrote down what a licensable speech corpus has to carry Buried in Track II is the clearest public statement of dataset hygiene any shared task has published, and it reads like a procurement checklist. The requirements for a contributed corpus: - **Human verification is mandatory.** "Raw, unverified machine translated outputs are not allowed." Post-editing of automatic output is allowed; the submitted data must be 100% verified by humans if not created by them. - **The MT you used must permit the downstream use.** The organisers name DeepL, Google Translate and ChatGPT as examples whose terms of service disallow reusing outputs to train other translation models. That is a licence problem, not a quality problem, and it disqualifies the fastest path to volume. - **Translation by qualified native speakers, verified by at least one more.** - **Identifiers, not names.** An ISO 639-3 individual language tag, a Glottocode, and an ISO 15924 script code on the dataset card. - **CC BY-SA 4.0 or similarly permissive,** research use at minimum. One corpus met all of it and published its own baseline. FLEURS-Badini extends FLEURS to the Badini variant of Northern Kurdish: 2,000 English FLORES sentences translated by English and Translation students at the University of Duhok, recorded by native speakers through an online platform, with both translations and recordings manually reviewed by faculty. The result is **5,224 utterances, 15 hours 40 minutes, 45 speakers**, split 2,022 utterances (5h47m) train, 1,165 (3h36m) dev, 2,037 (6h17m) test. A fine-tuned Whisper scores 5.24 BLEU and 29.57 chrF++ on it. That 5.24 is the honest number for a brand-new variant with 15 hours behind it, and it is worth holding next to the 21.09 that 30 curated hours of Central Kurdish produced. Same language family, roughly double the data, four times the score. ## How to spend a low-resource data budget in 2027 **Buy annotation depth before duration.** The corpus that produced 21.09 BLEU was 30 hours, manually segmented, transcribed and translated, across politics, culture, economy, sports, art and science. The corpus that produced 0.82 was over 130 hours. If your vendor quotes per recorded hour with transcription and translation as options, the options are the product. **Do not buy synthetic speech as a substitute for recorded speech.** 196 synthesised hours moved Irish to 2.4 BLEU. LIUM tested the same idea deliberately and reported it not fruitful for spontaneous speech, while pseudo-labelling real untranscribed audio worked. If you have untranscribed audio in the language, that is worth more than an equivalent budget of text-to-speech. **Check whether your target has a high-resource neighbour before you forecast quality.** Bhojpuri gets Hindi. Quechua and Mapuzugun get Spanish on the target side but nothing on the source side, and Mapuzugun, as an isolate, gets nothing at all. Judgement, not measurement: an isolate or a lone branch should be budgeted at two to three times the data of a language with a well-covered sibling, for the same expected quality. **Split your dialects at collection time.** The Quechua data is labelled with the `que` macro-code covering both Chanka (`quy`) and Collao (`quz`), and dialectal confusion is named as a failure cause. Recording metadata that distinguishes variants costs nothing at capture and cannot be recovered afterwards. **Write the data track's checklist into your supplier contract.** ISO 639-3 plus Glottocode plus ISO 15924, named-speaker consent, a second native verifier, an explicit licence, and a warranty that no output of a model whose terms forbid training reuse is present in the delivery. Every one of those is cheap to require at the start and impossible to retrofit. **Do not assume a frontier model closes the gap.** Five of them, prompted properly, lost to a fine-tuned NLLB-200 on Quechua text by 8.7 BLEU. For a genuinely low-resource locale, fine-tuning a specialised translation model on data you commissioned is still the stronger play in 2026. ## Check it yourself ``` # 1. the track (corpora, hours, provenance) and the results tables curl -sL -o iwslt2026.pdf https://aclanthology.org/2026.iwslt-1.39.pdf python3 -c " import pypdf t='\n'.join(p.extract_text() or '' for p in pypdf.PdfReader('iwslt2026.pdf').pages) i=t.find('Track II Low-resource'); print(t[i:i+22000])" # 2. the correlation quoted in the first line python3 - <<'PY' pairs = [("Irish",13,2.4), ("Bhojpuri",24,14.7), ("C. Kurdish",30,21.09), ("Quechua",50,27.2), ("Mapuzugun",130,0.82)] def rank(xs): order = sorted(range(len(xs)), key=lambda i: xs[i]); r=[0]*len(xs) for pos,i in enumerate(order): r[i]=pos+1 return r h = [p[1] for p in pairs]; b = [p[2] for p in pairs] rh, rb = rank(h), rank(b); n = len(pairs) d2 = sum((x-y)**2 for x,y in zip(rh, rb)) print("hour ranks", rh, " BLEU ranks", rb, " sum d^2", d2) print("Spearman rho =", 1 - 6*d2/(n*(n*n-1))) PY # hour ranks [1, 2, 3, 4, 5] BLEU ranks [2, 3, 4, 5, 1] sum d^2 20 # Spearman rho = 0.0 # 3. the corpora themselves # COMMUTE-Kurdish https://lium.univ-lemans.fr/en/corpus-commute-kurdish/ # Bhojpuri-Hindi https://github.com/shashwatup9k/iwslt2026_bho-hi # Irish-English https://github.com/shashwatup9k/iwslt2026_ga-eng # https://hf.co/collections/ymoslem/irish-english-speech-translation-datasets-665dd9e8fbaa279db3474ca0 # Bemba https://github.com/csikasote (Zambezi Voice, BembaSpeech) # YODAS https://hf.co/datasets/espnet/yodas # 4. the data-track requirements, verbatim python3 -c " import pypdf t='\n'.join(p.extract_text() or '' for p in pypdf.PdfReader('iwslt2026.pdf').pages) i=t.find('Data Submission Requirements'); print(t[i:i+1800])" ``` To price a collection against this evidence, count hours that are segmented, transcribed and translated separately from hours that are only recorded, and separately again from hours that are synthetic. The 2026 results do not let you add them together. ## What would prove this wrong The claim under test is that raw recorded duration is not what limits low-resource speech translation quality at these scales, and that annotation depth and transfer from a related language are. It is wrong if, by **31 December 2027**, the IWSLT low-resource track publishes results across at least five pairs with a Spearman rank correlation of **0.6 or higher** between provided speech hours and best BLEU, without a change in the mix of corpora. The 2026 figure across five pairs is 0.00. One more edition on the same pairs will settle it, because the corpora mostly carry over. It is also wrong on the Mapuzugun reading specifically if a 2027 submission reaches 10 BLEU or more on Mapuzugun-Spanish using only the existing 130-hour corpus and no newly collected Mapuzugun data. That would show the 0.82 was a recipe failure rather than a transfer failure. The current best is 0.82 from three submissions, all of which the organisers describe as struggling to produce meaningful outputs. A third prediction, marked as judgement: a frontier general-purpose model will still trail a fine-tuned specialised translation model on Quechua-Spanish text at IWSLT 2027. If a prompted or few-shot frontier model clears 19.5 BLEU there without fine-tuning on commissioned Quechua data, this reading was wrong. ## FAQ ### How many hours of speech do you need for a low-resource language? Hours are not the binding constraint at this scale. Across five IWSLT 2026 pairs the rank correlation between provided hours and best BLEU is 0.00. Thirty curated hours of Central Kurdish produced 21.09 BLEU; over 130 hours of Mapuzugun produced 0.82. ### Can a frontier LLM beat a specialised model on a low-resource language? Not for Quechua in 2026. GPT-5, Gemini 3, Claude, DeepSeek-V3 and Qwen were benchmarked with guided Spanish prompts; the best was 10.8 BLEU from Gemini 3 Flash, against 19.5 BLEU and 23.5 chrF for a fine-tuned NLLB-200 baseline. ### Does synthetic speech work as training data? The 2026 evidence says no. Irish got 196 synthetic hours on top of 13 real ones and reached 2.4 BLEU. LIUM compared synthesis with pseudo-labelling on Kurdish and found synthesis not fruitful for spontaneous speech, while pseudo-labelling matched cascades at lower latency. ### What does a contributed dataset have to carry? 100% human verification with no raw MT output, an MT licence that permits training reuse (which the organisers note DeepL, Google Translate and ChatGPT do not grant), translation by qualified native speakers with a second verifier, ISO 639-3 plus Glottocode plus ISO 15924 identifiers on the dataset card, and CC BY-SA 4.0 or a similarly permissive licence. ## Sources - IWSLT 2026 organisers (60 authors), [Speech Translation and Metrics in 2026: Findings of the IWSLT Campaign](https://aclanthology.org/2026.iwslt-1.39/), Proceedings of the 23rd International Conference on Spoken Language Translation, San Diego, July 2026. Track II (Low-resource SLT): per-pair corpus descriptions and speaker counts, submissions, Tables 2 to 7, data track requirements and FLEURS-Badini. [PDF](https://aclanthology.org/2026.iwslt-1.39.pdf). - QUESPA (Ortega et al., 2026), Quechua-Spanish submission described in Track II: the frontier-LLM case study (GPT-5, Gemini 3, Claude, DeepSeek-V3, Qwen), SIDON audio enhancement, and the 27.2 BLEU SpeechT5 system. - LIUM (Mohammadamini and Tahon, 2026), Central Kurdish-English submission: pseudo-labelling against speech synthesis, 21.09 BLEU, 6.98 CER and 19.76 WER. - Mohammadamini et al., 2026, FLEURS-Badini: 5,224 utterances, 15h40m, 45 speakers, University of Duhok, fine-tuned Whisper at 5.24 BLEU and 29.57 chrF++. Described in Track II, Data Track Results. - Corpora referenced by the organisers: [COMMUTE-Kurdish](https://lium.univ-lemans.fr/en/corpus-commute-kurdish/), [Bhojpuri-Hindi](https://github.com/shashwatup9k/iwslt2026_bho-hi), [Irish-English](https://github.com/shashwatup9k/iwslt2026_ga-eng), [YODAS](https://huggingface.co/datasets/espnet/yodas), Zambezi Voice and BembaSpeech (Sikasote et al.), and [Common Voice](https://commonvoice.mozilla.org/en/datasets). - Rank correlation in this note computed by BLOMEGA on 16 September 2026 from the Track II tables. Code in "Check it yourself". Related BLOMEGA guides: [Meta's 8B Omnilingual MT by resource tier](https://blomega.com/guides/omnilingual-mt-resource-tier-results/) · [A 30-trillion-token corpus buys Basque a 160-million-parameter model](https://blomega.com/guides/per-language-token-ceiling/) · [Consented AI training data providers](https://blomega.com/guides/consented-ai-training-data-providers/). ============================================================================== URL: https://blomega.com/research/llm-annotators-kappa-near-zero-2026/ Published: 2026-09-16 | Updated: 2026-09-16 ============================================================================== --- title: "Five machine labellers marked 0, 1, 40, 72 and 78 of the same 100 scenes positive. The human marked 9." url: https://blomega.com/research/llm-annotators-kappa-near-zero-2026/ published: 2026-09-16 updated: 2026-09-16 source: BLOMEGA (https://blomega.com/) --- # Five machine labellers marked 0, 1, 40, 72 and 78 of the same 100 scenes positive. The human marked 9. Lab note · 16 September 2026 · BLOMEGA A report posted on **12 September 2026** scored a rule-based detector and four language models against blind human labels on one Turkish corpus. On the scheme's hardest label, the five machines marked **0**, **1**, **40**, **72** and **78** of 100 scenes positive; the human marked **9**. Cohen's kappa against the human was 0.000, 0.015, 0.019, 0.027 and 0.185. Overall raw agreement for the same five runs was **74.7% to 86.3%**, which is the number a dashboard would have shown. ## What changed, and when On **12 September 2026**, Levent Bulut (independent researcher) posted [Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus](https://arxiv.org/abs/2609.13936) (arXiv:2609.13936v1, CC BY-NC-ND 4.0). The manuscript is dated August 2026 and is written as an empirical reliability report rather than a method paper. The object under test is the [Objective Projection](https://huggingface.co/datasets/leventbulut/objective-projection) corpus, 500 annotated Turkish-English scene pairs. Since version 7 every scene carries an `applied_rules` field written by `apply_rules.py`, a bilingual rule-based heuristic flagging six craft features: two prohibitions (explicit emotion labelling, simile) and four positive techniques (materialized metaphor, micro-focus, temporal anchor, atmosphere contradiction). Those flags had been used to describe the corpus and select examples. Nobody had checked whether a human agreed with them. The report opens with a declaration of three conflicts of interest: the author designed the six rules, served as the sole human rater in Study 1, and owns the dataset whose annotation layer the result damages. One of the scored systems, Claude, also assisted in preparing the analysis scripts and the manuscript. We flag these because the report does, and because they bound what the numbers can carry: the scoring is deterministic and reproducible from published files, the interpretation is not disinterested. Three studies ran. Study 1 (n=120) scored the detector against blind labels from the scheme's author. Study 2 (n=100, a scene set with no overlap with Study 1) scored the detector plus Gemini 2.5 Flash and Grok against an independent non-expert volunteer whose labels were locked before any machine ran. Study 2b re-ran Study 2 unchanged with Claude Fable 5 (High) and ChatGPT 5.5. A sixth system, Gemini 3.6, returned thirty identical label rows and was rejected under a pre-registered degenerate-output rule. ## The evidence table Study 2 and Study 2b scored five labellers against the same locked human reference on the same 100 scenes. Values transcribed from Tables 2 and 3 of arXiv:2609.13936v1. "H+" is the human positive count out of 100. Kappa is Cohen's kappa with _present_ as the positive class and the human as reference; `n/a` means expected agreement equalled 1, so kappa is undefined. | Rule | H+ | Detector + / κ | Gemini 2.5 Flash + / κ | Grok + / κ | Claude Fable 5 + / κ | ChatGPT 5.5 + / κ | Source | | --- | --- | --- | --- | --- | --- | --- | --- | | Emotion label | 0 | 0 / n/a | 0 / n/a | 0 / n/a | 0 / n/a | 0 / n/a | Tables 2, 3 | | Simile | 1 | 0 / 0.000 | 0 / 0.000 | 0 / 0.000 | 0 / 0.000 | 0 / 0.000 | Tables 2, 3 | | **Materialized metaphor** | 9 | 72 / 0.015 | 1 / 0.185 | 0 / 0.000 | 78 / 0.027 | 40 / 0.019 | Tables 2, 3 | | Micro-focus | 96 | 81 / 0.022 | 9 / 0.008 | 82 / -0.070 | 100 / 0.000 | 92 / 0.296 | Tables 2, 3 | | Temporal anchor | 99 | 82 / -0.019 | 93 / -0.018 | 95 / -0.017 | 100 / 0.000 | 98 / -0.014 | Tables 2, 3 | | Atmosphere contradiction | 44 | 0 / 0.000 | 2 / 0.051 | 6 / 0.020 | 55 / 0.269 | 42 / 0.184 | Tables 2, 3 | | **Overall raw agreement** | – | 74.7% | 75.7% | 86.3% | 81.0% | 84.5% | Tables 2, 3 | Read the bottom row, then the materialized-metaphor row. Grok posts the highest raw agreement in Study 2, 86.3%, and its kappa on the scheme's central feature is exactly 0.000 because it marked zero scenes present. ChatGPT 5.5 posts 84.5%, the highest figure in any of the three studies, at kappa 0.019. Claude Fable 5 reached 29% raw agreement on that single rule, the lowest cell anywhere in the report, by marking 78 scenes present against the human's 9. _One written definition, one scene set, six raters. The detector and Claude sit above the human by a factor of eight; Grok and Gemini sit below it by an order of magnitude._ Study 1 is the control that makes the rest legible. It scored the detector against the scheme's own author on a disjoint set of 120 scenes, and the surface rules behaved exactly as a surface rule should. | Rule (Study 1, n=120) | Human + | Detector + | TP | FP | FN | Precision | Recall | κ | Source | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | Emotion label | 2 | 12 | 2 | 10 | 0 | 0.167 | 1.000 | 0.265 | Table 1 | | Simile | 2 | 2 | 2 | 0 | 0 | 1.000 | 1.000 | 1.000 | Table 1 | | Materialized metaphor | 63 | 95 | 50 | 45 | 13 | 0.526 | 0.794 | 0.004 | Table 1 | | Micro-focus | 118 | 96 | 94 | 2 | 24 | 0.979 | 0.797 | -0.032 | Table 1 | | Temporal anchor | 120 | 100 | 100 | 0 | 20 | 1.000 | 0.833 | 0.000 | Table 1 | | Atmosphere contradiction | 12 | 17 | 5 | 12 | 7 | 0.294 | 0.417 | 0.258 | Table 1 | Simile reduces to a short list of Turkish function words (_gibi_, _sanki_, _adeta_) and scores kappa 1.000. Materialized metaphor, on the only near-balanced distribution in the whole report (63 present, 57 absent), scores 0.004 at 51.7% raw agreement. Overall raw agreement across all 720 cells in Study 1 was 81.5%. That single number would have passed any review. ## A labeller that always says absent scores 91% on a 9-in-100 label The arithmetic is not subtle and it is worth writing out, because it is the whole finding. Grok marked zero of 100 scenes as containing a materialized metaphor. The human marked 9. Grok is therefore right on 91 scenes and wrong on 9, which is 91% raw agreement. Cohen's kappa asks a different question: how much of that 91% would two raters have hit by accident, given how often each of them says _present_? A rater that never says _present_ has an expected agreement of exactly 91% too, so the observed agreement buys nothing and kappa is 0.000. _Grok's 91% and Claude's 29% sit at opposite ends of a raw-agreement scale and at the same place on a kappa scale, which is the point._ Five of the six rules in Study 2 had human positive counts of 0, 1, 9, 96 and 99. On four of those, kappa carries almost no information and raw agreement is a report on the marginal distribution. Only atmosphere contradiction, at 44 of 100, sits in a range where agreement can be read at all. On that rule, and only that rule, two models are clearly above chance: Claude Fable 5 at kappa 0.269 and ChatGPT 5.5 at 0.184. The report records this as a correction to its own earlier Study 2 conclusion that machines could not detect the feature. The design that produced these numbers is worth carrying over, independently of the corpus it was run on. _The locked reference and the pre-registered rejection rule are the two cheap parts of this design, and they are what make the spread in Section 6.1 readable rather than arguable._ ## What it means if you buy, sell or ship an annotation layer **A corpus-level agreement number is not a quality claim.** This study's best headline figure is 86.3%, produced by a labeller that returned zero positives on the feature the corpus exists to demonstrate. If a vendor or a dataset card quotes one agreement percentage, ask for it per label, alongside the positive count of the human reference for that label. Without the marginal distribution the percentage is not interpretable, and the report is explicit that its own five skewed rules make raw agreement misleading. **Machine labellers disagree with each other, not just with humans.** On micro-focus, Gemini marked 9 scenes and Grok marked 82 against a human count of 96, from the same written definition and the same prompt blocks. Gemini and Grok agreed with each other on 85.7% of cells overall, which again is the marginal distribution talking. Our judgement: a pre-labelling step that silently switches model or model version is a silent change of label semantics, and should be versioned like a schema migration. **Consistency is not validity.** ChatGPT reproduced 58 of 60 labels on re-run and Claude 60 of 60. Both were stable. Both were at chance against the human on the rule that mattered. A stability metric measures whether the labeller is reproducible; it says nothing about whether it is applying your definition. **Build a balanced slice before you evaluate.** Four of the six rules here cannot be tested on this corpus at any sample size, because 96 to 100 percent of scenes carry them. The design lesson the report draws is the one to steal: sample the evaluation set for balance on the target feature, or accept that the feature is untested. This costs a sampling pass and it is the difference between an interpretable kappa and a decorative one. **Two humans, not one.** Every number in the report is agreement with one particular person. That is enough to show the machines are not applying the same rule as each other, and not enough to say whether the rule is hard or broken. For a commercial annotation programme the same gap shows up as a dispute you cannot adjudicate, which is why double-scoring a slice is a cheaper purchase than it looks. ## Check it yourself The labels, prompt blocks and scoring scripts are in the `evaluation/` directory of the dataset. Unlike many reliability claims, the central table can be recomputed without asking anyone for access. ``` # dataset metadata: gating, license, freshness, downloads curl -s https://huggingface.co/api/datasets/leventbulut/objective-projection \ | python3 -c 'import json,sys; d=json.load(sys.stdin); print(d["gated"], d["cardData"]["license"], d["lastModified"], d["downloads"])' # False cc-by-nc-nd-4.0 2026-09-15T18:53:47.000Z 337 (checked 16 Sep 2026) # the paper, and the table values quoted above curl -sL -o 2609.13936.pdf https://arxiv.org/pdf/2609.13936 # Tables 1, 2, 3 # reproduce the kappa arithmetic behind Grok's 91% on materialized metaphor python3 - <<'PY' def kappa(tp, fp, fn, tn): n = tp + fp + fn + tn po = (tp + tn) / n pe = ((tp+fp)*(tp+fn) + (fn+tn)*(fp+tn)) / (n*n) return po, pe, (po - pe) / (1 - pe) if pe < 1 else float('nan') for name, cell in [("Grok ", (0, 0, 9, 91)), ("Claude Fable 5 ", (8, 70, 1, 21)), ("ChatGPT 5.5 ", (4, 36, 5, 55)), ("rule detector ", (7, 65, 2, 26))]: po, pe, k = kappa(*cell) print(f"{name} raw={po:.3f} expected={pe:.3f} kappa={k:+.3f}") PY # Grok raw=0.910 expected=0.910 kappa=+0.000 # Claude Fable 5 raw=0.290 expected=0.270 kappa=+0.027 # ChatGPT 5.5 raw=0.590 expected=0.582 kappa=+0.019 # rule detector raw=0.330 expected=0.320 kappa=+0.015 ``` Two caveats on what is recomputable, both stated in the report. The per-scene label files for Gemini 2.5 Flash and Grok were lost; their confusion counts in Table 2 were recovered arithmetically from the surviving kappa and agreement values, uniquely for eleven of twelve cells (Grok on atmosphere contradiction admitted TP=3 or TP=4, resolved to 3 by the original write-up). The detector's Study 2 labels are a reconstruction from the published `applied_rules` field through the published identifier mapping, and reproduce the original column exactly. The scene-level Gemini-Grok agreement figure of 85.7% cannot be recomputed and is reproduced from the original analysis. We ran that snippet on 16 September 2026. It reproduces all four published kappas to three decimal places from the confusion counts alone: 0.000, 0.027, 0.019 and 0.015. The expected-agreement column is the part worth staring at. Grok's is 0.910, identical to its observed agreement, which is the whole of its 91%. Claude's is 0.270 against an observed 0.290, so 78 positive calls buy two points of signal over a coin weighted the same way. ## What would prove this wrong The claim we are making is narrower than the paper's: that on a judgement-heavy binary label with a skewed human distribution, raw agreement and self-consistency can both look healthy while chance-corrected agreement is zero, and that this is a general property of such evaluations rather than an artefact of one Turkish corpus. A dated prediction. By **30 June 2027**, a replication of the materialized-metaphor rule on the same 100 scenes with a second independent human rater, using the operational rewrite the report asks for in Section 8, will not produce a prompt-only language model at Cohen's kappa above **0.40** against either human. If someone posts that result, with the published scene set and a locked reference, this article's reading is wrong and the definition, not the machines, was the binding constraint. The second rater is the measurement that decides it, and no further model run substitutes for it. ## Sources - Bulut, L. [Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus](https://arxiv.org/abs/2609.13936). arXiv:2609.13936v1, 12 September 2026. Tables 1, 2 and 3, Sections 2, 4.1, 6, 7 and 8; [PDF](https://arxiv.org/pdf/2609.13936). - [Objective Projection dataset card](https://huggingface.co/datasets/leventbulut/objective-projection), Hugging Face. Metadata retrieved 16 September 2026: not gated, CC BY-NC-ND 4.0, last modified 15 September 2026, 337 downloads. - Feinstein, A. R., Cicchetti, D. V. [High agreement but low kappa: I. The problems of two paradoxes](https://pubmed.ncbi.nlm.nih.gov/2348207/). Journal of Clinical Epidemiology 43(6), 1990. The first kappa paradox, which the report cites for its skewed-distribution reading. - Pangakis, N., Wolken, S., Fasching, N. [Automated annotation with generative AI requires validation](https://arxiv.org/abs/2306.00176). arXiv:2306.00176, 2023. Cited in the report as the task-by-task validation argument. - BLOMEGA. [Two audits of HH-RLHF disagree on how much of it is mislabelled](https://blomega.com/research/hh-rlhf-preference-label-noise-2026/), on what happens downstream when a label layer is wrong. - BLOMEGA. [Data annotation research: the latest](https://blomega.com/research/data-annotation-latest-research/). ============================================================================== URL: https://blomega.com/research/rare-event-labeling-prevalence-effect-2026/ Published: 2026-09-16 | Updated: 2026-09-16 ============================================================================== --- title: "Once individual accuracy drops below 50%, majority vote makes rare-event labels worse" url: https://blomega.com/research/rare-event-labeling-prevalence-effect-2026/ published: 2026-09-16 updated: 2026-09-16 source: BLOMEGA (https://blomega.com/) --- # Once individual accuracy drops below 50%, majority vote makes rare-event labels worse Lab note · 16 September 2026 · BLOMEGA A field experiment with **290** annotators on a live medical crowdsourcing platform, posted **12 March 2026**, holds the unlabeled stream at **20%** positives and moves only the prevalence of the gold-standard feedback stream. In the block where mean individual miss rate reached **51.1%**, adding annotators pushed the crowd miss rate up rather than down. Recalibrating the aggregate of nine judgements took the crowd miss rate from about **55%** to about **9%** at a false alarm rate near **3%**, and the convolutional networks trained on those labels moved with them. ## What changed, and when On **12 March 2026**, Gunnar P. Epping, Andrew Caplin, Erik Duhaime, William R. Holmes, Daniel Martin and Jennifer S. Trueblood posted [Managing Cognitive Bias in Human Labeling Operations for Rare-Event AI: Evidence from a Field Experiment](https://arxiv.org/abs/2603.11511) (arXiv:2603.11511v1, cs.HC and econ.GN, CC BY 4.0). It is framed as an operations paper, not a machine learning paper, and it treats the gold-standard feedback stream as a policy lever rather than a fixed property of the task. The prevalence effect itself is old. Wolfe and colleagues established that observers miss rare targets at elevated rates in visual search, and the effect survives training and incentives. What is new here is running it inside a working annotation pipeline, then following the resulting labels all the way into a trained model. Study 1 is a lab replication that scales the effect to the crowd. Study 1a used 39 students (mean age 19.7, 77% female) across blocks at 75%, 50% and 25% target prevalence. Study 1b used 57 students (mean age 19.2, 58% female) split into a high group (90% and 50% blocks) and a low group (10% and 50% blocks). Study 2 is the field experiment: **290** participants recruited through DiagnosUs, the gamified labeling contest run by Centaur Labs, classifying white blood cell images as blast (cancerous) or non-blast. _Centaur Labs Medical Data Labeling Demo, [Centaur Labs](https://www.youtube.com/@centaurlabs8913) on YouTube, 2 December 2021. The DiagnosUs platform the Study 2 field experiment ran on: contests, leaderboards and cash prizes scored against interleaved gold-standard items, which is the stream the experiment manipulates._ ## The evidence table Study 2 crossed two levers. The response interface was either binary choice (BC, "is this a blast cell?") or elicited beliefs (EB, "what is the likelihood that this is a blast cell?"). The gold-standard feedback stream was either matched to the unlabeled stream at 20% positives or balanced at 50%. Cell sizes: BC 20% n=75, BC 50% n=67, EB 20% n=75, EB 50% n=73. The unlabeled QA set was 750 images, 150 blast and 600 non-blast, the latter produced by rotating 150 distinct non-blast images by 90, 180 and 270 degrees. Most Study 2 rates are reported in figures rather than in a table, and the text describes them with "around" and "roughly". We reproduce them the way the paper states them, and write "not reported" where no number appears. | Level | Data variant | GS prevalence | Miss rate | False alarm rate | Source | | --- | --- | --- | --- | --- | --- | | Individual | Binary choice (BC) | 20% | ~35 to 40% | ~10 to 15% | §3.2.1 | | Individual | Elicited beliefs (EB) | 20% | ~35 to 40% | ~10 to 15% | §3.2.1 | | Individual | Recalibrated beliefs (rEB) | 20% | ~60% (0.58) | ~3% | §3.2.1, §3.2.2 | | Individual | BC, EB and rEB | 50% | "do not vary much" across modes | "do not vary much" across modes | §3.2.1 | | Crowd of 9 | rEB without crowd recalibration | 20% | ~55% (flat in crowd size) | not reported | §3.2.2 | | **Crowd of 9** | **rEB with crowd recalibration** | 20% | ~9% | ~3% | §3.2.2, §4.1 | | Crowd of 9 | Binary choice (BC) | 20% | 0.28 | not reported | §3.3 | | CNN (GoogLeNet) | trained on BC crowd labels | 20% | 0.24 | not reported | §3.3 | | Individual (Study 1b) | binary, 10% prevalence block | n/a, lab | 51.1% | not reported | §2.3.2 | Two rows carry most of the argument. The rEB individual row shows that recalibrating each worker on its own made the miss rate _worse_ in the low-prevalence feedback condition, roughly 60% against 35% to 40%, while cutting false alarms to about 3%. The rEB-with-crowd-recalibration row shows that doing the same correction on the aggregate, after pooling nine judgements, takes the miss rate to about 9% while leaving false alarms near 3%. Same workers, same images, same nine votes. The difference is where in the pipeline the correction is applied. ## Why the crossover sits at exactly p = 0.5 The paper's cleanest result is a piece of arithmetic it states in two lines, and it explains why redundancy is not a safety net for rare events. Suppose every annotator has the same probability _p_ of being correct on an item, and errors are independent. A crowd of one is correct with probability _p_. A crowd of three, decided by majority, is correct with probability _p_3 + 3_p_2(1 − _p_). Those two curves cross at _p_ = 0.5. Below it, every annotator you add makes the aggregate worse. _Curves computed from the majority-vote binomial the paper states in Section 2.3.2, for p from 0.2 to 1.0. The amber dot is the only empirical point on this chart: Study 1b's 10% prevalence block. Everything else is the arithmetic that makes that block's behaviour predictable rather than surprising._ Prevalence pushes _p_ down. That is the mechanism by which a rare-event stream converts a wisdom-of-crowds pipeline into a wisdom-destroying one: the prevalence effect does not just add noise, it biases every annotator in the same direction, which breaks the error-independence that aggregation depends on. Study 1 confirms both halves. Crowd miss rates fell and crowd false alarm rates rose as prevalence increased, exactly tracking the individual pattern, and in the 10% and 90% extremes the crowd was worse than a randomly chosen individual. Study 2's fix is two interventions in the pipeline, neither of which changes what an annotator is paid or how many of them there are. _LLO is the linear-in-log-odds recalibration the paper fits twice. The two amber boxes are the interventions; everything else is a pipeline most annotation operations already run._ _The amber arrow is the whole intervention: the same nine judgements, moved 50 points down the miss axis by a fit applied after aggregation rather than before it._ Why worker-level recalibration alone backfires is worth saying plainly. Fitting each worker's probabilities to gold-standard outcomes corrects overconfidence in both directions, and in a 20% feedback stream that mostly means pushing probabilities down. Individual false alarms fall to about 3%, which looks like a win, and individual misses climb to about 60%, which is the same bias made sharper. The crowd-level fit works because it is applied after the averaging, where the systematic underestimation shows up as a shift in the calibration curve rather than as per-worker noise. The paper's calibration curves make this concrete: in one bin, roughly 75% of images carrying crowd labels between 2/7 and 3/7 were in fact blast cells. The bias reaches the model. Expected calibration error, computed over ten equal-width bins on 750 images, was best for the crowd-recalibrated variant in both feedback conditions and worst for the worker-recalibrated variant without it, and the ranking survived into the trained CNNs. One asymmetry is worth noting: the CNNs came out slightly _less_ miss-prone than their training labels (0.24 against 0.28 for the binary-choice variant), so models do not fully inherit label bias on miss rate, but their expected calibration error was worse than their labels' across the board. ## What it means if you run or buy rare-event annotation **Your gold-standard stream is a design parameter, not a sample.** Most QA schemes interleave gold items drawn from the same distribution as the work, because that feels neutral. This experiment shows the composition of that stream sets the base rate annotators experience and therefore moves their decision criterion. Balancing it to 50% cost nothing per label and moved the error split from miss-heavy toward even at both individual and crowd level. If misses are the expensive error, this is the cheapest lever in the pipeline. **Stop buying redundancy blind.** Nine votes per image is a real cost, and in a rare-event regime it can buy negative value. Before scaling redundancy, estimate the per-annotator accuracy on the target class. If it is under 0.5, majority vote is the wrong aggregator and more votes make the aggregate worse, which is a claim with a closed-form proof and a matching empirical block in this paper. **Elicit probabilities and score them properly.** Asking "how likely is this a blast cell?" rather than "is this a blast cell?" gave a lower crowd miss rate in the 20% condition and is a prerequisite for recalibration. Our judgement: this is under-adopted because binary labels are what most annotation tools and most downstream training loops want. The paper's answer is that the probability is the thing to store and the binary is the thing to derive later, after correction. **Recalibrate after aggregation, not before it.** This is the counterintuitive result and the most transferable one. Worker-level correction alone moved the individual miss rate the wrong way. The same functional form applied to the pooled label moved the crowd miss rate from about 55% to about 9%. If you already collect gold-standard items, you already have the fitting set. **Grade the labels, not just the model.** Eight data variants were built and evaluated as labels before any network was trained, and the ranking of the labels predicted the ranking of the models. A pipeline that only measures downstream test accuracy cannot see that its miss-heavy labels are miss-heavy, because the test set is labelled the same way. ## Check it yourself The crossover result needs nothing but the binomial. The field-experiment rates are reported in figures, so they cannot be recomputed from the PDF; what follows is the arithmetic that makes the crowd result predictable, plus the checks we ran on the paper's availability. ``` # the paper (no arXiv HTML gaps; both formats resolve) curl -s -o /dev/null -w "%{http_code}\n" https://arxiv.org/abs/2603.11511 # 200 curl -sL -o 2603.11511.pdf https://arxiv.org/pdf/2603.11511 # where majority vote starts to hurt: crowd of N vs a single annotator python3 - <<'PY' from math import comb def crowd(p, n): # majority of n independent annotators return sum(comb(n, k) * p**k * (1-p)**(n-k) for k in range((n//2)+1, n+1)) p_10pct = 1 - 0.511 # Study 1b, 10% prevalence block: miss 51.1% for n in (1, 3, 5, 7, 9): print(f"n={n}: p=0.489 -> {crowd(p_10pct, n):.3f} p=0.650 -> {crowd(0.65, n):.3f}") PY # n=1: p=0.489 -> 0.489 p=0.650 -> 0.650 # n=3: p=0.489 -> 0.484 p=0.650 -> 0.718 # n=5: p=0.489 -> 0.479 p=0.650 -> 0.765 # n=7: p=0.489 -> 0.476 p=0.650 -> 0.800 # n=9: p=0.489 -> 0.473 p=0.650 -> 0.828 # the linear-in-log-odds recalibration the paper fits at worker and crowd level python3 - <<'PY' import math def llo(p, a, b): # Gonzalez-Wu form: logit-linear in log odds if p in (0.0, 1.0): return p z = a * math.log(p/(1-p)) + b return 1 / (1 + math.exp(-z)) # an underestimating crowd (a0 shifts the whole curve up) for p in (0.20, 0.30, 0.40, 0.50): print(f"raw {p:.2f} -> recalibrated {llo(p, 1.4, 0.9):.3f}") PY # raw 0.20 -> recalibrated 0.261 # raw 0.30 -> recalibrated 0.429 # raw 0.40 -> recalibrated 0.582 # raw 0.50 -> recalibrated 0.711 ``` The LLO parameters above are ours, chosen to illustrate the shape; the paper fits them per worker and per crowd dataset from gold-standard items, and does not publish the fitted values. What the second snippet shows is the mechanism the paper describes in Section 3.2.2: a curve of this shape moves a band of crowd labels from below 0.5 to above it, which is precisely where the miss rate lives. Two things we could not verify. The paper lists no public code or data repository, so the 750-image QA set, the 580-image gold-standard set and the per-worker judgements are not available for re-analysis. The Study 2 miss and false alarm rates in the table above are read off figures by the authors' own prose, so they carry the precision of "around" and should not be quoted to a decimal place. ## What would prove this wrong The transferable claim is that in a rare-event labeling stream, balancing the gold-standard feedback prevalence and applying a calibration fit to the aggregated label both reduce misses, and that the second matters more than adding annotators. A dated prediction. By **31 March 2027**, a replication on a different rare-event annotation task with true prevalence at or below 20%, at least 100 annotators, and the same two levers, will show crowd-level recalibration reducing the miss rate by at least **20 percentage points** relative to an uncalibrated majority vote on the same judgements, at a false alarm rate no more than **10 percentage points** higher. A replication that finds crowd recalibration neutral or harmful on that comparison, with the judgements published, falsifies the reading here. The weakest point is the population: DiagnosUs contestants self-select into contests and compete for prizes, which is a specific incentive structure and not every annotation workforce. ## Sources - Epping, G. P., Caplin, A., Duhaime, E., Holmes, W. R., Martin, D., Trueblood, J. S. [Managing Cognitive Bias in Human Labeling Operations for Rare-Event AI: Evidence from a Field Experiment](https://arxiv.org/abs/2603.11511). arXiv:2603.11511v1, 12 March 2026. Sections 2.1 to 2.4, 3.1 to 3.4 and 4.1; [PDF](https://arxiv.org/pdf/2603.11511). - Wolfe, J. M., Horowitz, T. S., Van Wert, M. J., et al. [Low target prevalence is a stubborn source of errors in visual search tasks](https://pmc.ncbi.nlm.nih.gov/articles/PMC2662480/). Journal of Experimental Psychology: General, 2007. The prevalence effect the paper builds on. - Centaur Labs. [Centaur Labs Medical Data Labeling Demo](https://www.youtube.com/watch?v=RyuxmawvwUY). YouTube, 2 December 2021. The DiagnosUs platform used in Study 2. - Guo, C., Pleiss, G., Sun, Y., Weinberger, K. Q. [On Calibration of Modern Neural Networks](https://arxiv.org/abs/1706.04599). arXiv:1706.04599, 2017. Source of the expected calibration error definition used with ten bins. - BLOMEGA. [Half of NVD's CWE labels match the vendor's own](https://blomega.com/research/nvd-cwe-label-audit-2026/), on what an unaudited label layer costs downstream. - BLOMEGA. [Data annotation research: the latest](https://blomega.com/research/data-annotation-latest-research/). ============================================================================== URL: https://blomega.com/research/sft-rl-annotation-budget-near-optimal-region-2026/ Published: 2026-09-16 | Updated: 2026-09-16 ============================================================================== --- title: "A blind 50/50 split of the annotation budget missed the near-optimal region up to 80% of the time. A $2 proxy run did not." url: https://blomega.com/research/sft-rl-annotation-budget-near-optimal-region-2026/ published: 2026-09-16 updated: 2026-09-16 source: BLOMEGA (https://blomega.com/) --- # A blind 50/50 split of the annotation budget missed the near-optimal region up to 80% of the time. A $2 proxy run did not. Lab note · 16 September 2026 · BLOMEGA A paper posted **1 September 2026** swept how a fixed annotation budget should divide between demonstrations and preference data, across three model families, four tasks and two RL objectives. Allowing **10%** off peak, most tasks accepted splits covering **55% to 75%** of the feasible range. Ratios transferred from a 1B proxy landed inside a 3B-to-8B target's **5%** near-optimal region **0.90 to 0.95** of the time; a fixed 0.5 split landed inside it **0.20 to 0.60** of the time. The proxy sweep costs about **$2.10**. ## What changed, and when On **1 September 2026**, Jingtan Wang, Arun Verma, Xiaoqiang Lin, Zhengyuan Liu, Nancy F. Chen, Daniela Rus and Bryan Kian Hsiang Low posted [Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs](https://arxiv.org/abs/2609.01573) (arXiv:2609.01573v1, accepted to EMNLP 2026). The framing is what is new. Earlier work asked which SFT-to-RL ratio is optimal and reported broad trends, for example that SFT dominates in low-data regimes. This paper asks instead for the **near-optimal region**: the set of allocation ratios whose performance stays within a stated tolerance of the best observed performance at that budget and model size. Formally, for a total budget _B_ and ratio _r_, _rB_ samples go to SFT and (1 − _r_)_B_ to the RL stage, and the region is every _r_ whose score is at least (1 − ε) times the peak. Budget is counted in annotated training samples, not GPU hours, and the paper is explicit about why: annotation dominates the cost of post-training. Even against cheap synthetic annotation at $10-3 per example, it reports compute staying below annotation, which is what makes the ratio a procurement question rather than a scheduling one. The sweep runs the Cartesian product of model size, ratio, budget, task, algorithm and model family, on the order of hundreds to thousands of post-training runs, which is why every run uses LoRA rather than full fine-tuning. The ratio grid is `{0.00, 0.25, 0.50, 0.75, 1.00}`, budgets sit in [0, 15k] samples with the analysis restricted to B ≥ 5k, and a denser 9-point grid adding `{0.125, 0.375, 0.625, 0.875}` reproduces the same qualitative behaviour. ## The evidence table The headline comparison is Table 6 of the paper: how often each recommendation lands inside the target model's near-optimal region, averaged over budgets and over target sizes from 3B to 8B. "Proxy" means running the 5-point ratio grid on the family's smallest model (1B for Llama, 1.5B for Qwen 2.5) and transferring the region. "Fixed 0.5" means splitting the budget down the middle without measuring anything. | Setting | Tolerance | Proxy hit rate | Fixed r = 0.5 hit rate | Gap | Source | | --- | --- | --- | --- | --- | --- | | Llama, HelpSteer | 5% | 0.95 | 0.20 | +0.75 | Table 6 | | Llama, instruction following | 5% | 0.95 | 0.40 | +0.55 | Table 6 | | Qwen 2.5, instruction following | 5% | 0.90 | 0.60 | +0.30 | Table 6 | | Llama, HelpSteer | 10% | 1.00 | 0.70 | +0.30 | Table 6 | | Llama, instruction following | 10% | 0.95 | 0.70 | +0.25 | Table 6 | | Qwen 2.5, instruction following | 10% | 1.00 | 1.00 | 0.00 | Table 6 | _The gap closes at the loosest tolerance on one of three settings. At 5%, which is the tolerance a team with a real quality bar would pick, the blind split misses four times out of five on Llama HelpSteer._ The four tasks and where their data comes from matter, because the result is a claim about ratios and not about any one dataset. Every task keeps SFT and RL-stage data from a consistent source, either human-annotated or machine-generated, so the comparison across ratios is not confounded by provenance. | Task | SFT-stage data | RL-stage data | Evaluation | Source | | --- | --- | --- | --- | --- | | Math | GSM8K | Tülu3 Grade School Math | GSM8K test accuracy | Table 1 | | Instruction following | Tülu3 Persona IF | Tülu3 Persona IF (DPO) / Tülu3 RLVR IF (GRPO) | IFEval accuracy | Table 1 | | Summarization | Reddit TL;DR | Reddit Comparison | ROUGE-L F1 | Table 1 | | Helpfulness | HelpSteer | HelpSteer2 | Reward model score | Table 1 | ## The region is what transfers; the single best ratio is not Here is the shape of the result. Performance as a function of the allocation ratio is not a peak with steep sides. It is a plateau. At a 10% tolerance most tasks admit near-optimal ratios spanning 55% to 75% of the allocation space, and per-ratio hit-rate heatmaps show the admitted ratios are contiguous on the grid rather than scattered, which is the evidence that the plateau is real and not a sampling artefact. That plateau generally widens with model size at a fixed tolerance. The paper is careful about why, and we are repeating its caution rather than its headline: part of the widening is mechanical, because a larger model has a higher absolute peak, so the same relative tolerance admits more absolute slack. Under an absolute tolerance anchored to the smallest model, the widening is dampened for Llama and reverses outright on Llama math. What survives both definitions is transfer. _The proxy is insurance, not optimisation. It does not find a better ratio than the 8B sweep would; it finds the same region for two fifths of the compute, and it is the only one of the two cheap options that finds anything at all._ Cost asymmetry moves the region, in a direction that is convenient. The paper fixes a DPO preference example at $0.001 and varies ρ, the ratio of SFT cost to DPO cost, noting that synthetic SFT demonstrations from frontier models run roughly 1 to 2 times the cost of preference-style annotation. As ρ rises, the near-optimal region widens at the same tolerance: when demonstrations get expensive, the split matters less and annotation logistics can drive it. The opposite regime is the one to watch, and it is the case where a team budgets in GPU hours rather than in labels. DPO takes roughly 2 to 3 times the GPU time of SFT, which puts ρ near 0.5, and there the plateau disappears. _Nine measured widths. Eight of the nine cells are a single ratio, which is the boundary condition on everything else in this article._ | Near-optimal region width (Math, Llama, ρ = 0.5, GPU-hour budget) | ε = 2% | ε = 5% | ε = 10% | Source | | --- | --- | --- | --- | --- | | 1B | 0.00 | 0.00 | 0.00 | Table 7 | | 3B | 0.00 | 0.00 | 0.06 | Table 7 | | 8B | 0.00 | 0.13 | 0.25 | Table 7 | | Transfer 1B → 3B, 1B → 8B, 3B → 8B | 1.00 | 1.00 | 1.00 | Table 8 | A width of 0.00 means one ratio and no flexibility. In that regime transfer is trivially perfect, because the single near-optimal ratio is the same at every scale, so the proxy's recommendation is always right and also always the only option. The paper says so. It is the honest reading, and it is a useful boundary: the plateau is a property of counting budget in labels, not a universal property of post-training. ## What it means if you are buying the labels **Buy the proxy sweep before you buy the labels.** The asymmetry is the whole argument. A 5-point ratio sweep on a 1B model is about 100 GPU-minutes and about $2.10; the annotation budget it is allocating is 5,000 to 15,000 samples, which at even $0.001 per preference pair and a dollar per human demonstration is the part with real money in it. Spending two dollars to place tens of thousands of dollars of annotation is not a close call. **A 50/50 default is a coin flip dressed as a policy.** It was inside the 5% region 0.20 of the time on Llama HelpSteer and 0.40 on Llama instruction following. It costs exactly as much as the proxy run and returns no information about where the region actually sits. Our judgement: the reason 0.5 persists is that it is the only ratio nobody has to defend, and this paper removes that excuse for about the price of a coffee. **Ask for the tolerance, not the ratio.** If a vendor or an internal team quotes "we use a 70/30 split", the useful follow-up is what tolerance that was chosen under and on what proxy. The gap between the proxy and the blind default collapses to zero at 10% tolerance on Qwen 2.5 instruction following and is 0.75 at 5% tolerance on Llama HelpSteer. The same decision is either free or expensive depending entirely on the quality bar. **The plateau is permission to optimise for something else.** If 55% to 75% of the allocation space is within 10% of peak, then inside that band you can choose the ratio on grounds the paper does not model: which annotation is easier to source consented, which supplier has capacity, which data you can reuse across projects. That is a real operational freedom and it only exists once you know where the band is. **Check which budget you are actually spending.** The plateau is measured against a budget counted in annotated samples. Counted in GPU hours, with DPO at 2 to 3 times the cost of SFT, the region collapsed to a single ratio at 1B and 3B. Teams whose constraint is a compute allocation rather than a labeling contract should not expect the flexibility. ## Check it yourself The paper publishes no code or trained artefacts we could find, so the check here is on the arithmetic of the recommendation, not on the sweep. Everything below runs in under a second. ``` # the paper (HTML full text is available, which is where Tables 6 to 8 live) curl -s -o /dev/null -w "%{http_code}\n" https://arxiv.org/abs/2609.01573 # 200 curl -s -o /dev/null -w "%{http_code}\n" https://arxiv.org/html/2609.01573 # 200 # what the proxy buys, per Table 6, in expected wasted post-training runs python3 - <<'PY' tbl = { # (setting, tolerance): (proxy hit rate, fixed r=0.5 hit rate) ("Llama HelpSteer", "5%"): (0.95, 0.20), ("Llama Inst", "5%"): (0.95, 0.40), ("Qwen 2.5 Inst", "5%"): (0.90, 0.60), ("Llama HelpSteer", "10%"): (1.00, 0.70), ("Llama Inst", "10%"): (0.95, 0.70), ("Qwen 2.5 Inst", "10%"): (1.00, 1.00), } GPU_HR = 1.25 # CoreWeave L40, as reported in the paper proxy_min, target_min = 100, 100 for (name, tol), (pr, fx) in tbl.items(): miss = fx - pr # extra failure probability of guessing print(f"{name:16} eps={tol:>3} proxy {pr:.2f} fixed {fx:.2f} " f"extra miss {abs(miss):.2f} proxy cost ${proxy_min/60*GPU_HR:.2f}") print() print(f"proxy route : {(proxy_min+target_min)/60*GPU_HR:.2f} USD, " f"{(proxy_min+target_min)} GPU-min") print(f"8B full sweep: {500/60*GPU_HR:.2f} USD, 500 GPU-min " f"({500/(proxy_min+target_min):.1f}x)") PY # Llama HelpSteer eps= 5% proxy 0.95 fixed 0.20 extra miss 0.75 proxy cost $2.08 # Llama Inst eps= 5% proxy 0.95 fixed 0.40 extra miss 0.55 proxy cost $2.08 # Qwen 2.5 Inst eps= 5% proxy 0.90 fixed 0.60 extra miss 0.30 proxy cost $2.08 # Llama HelpSteer eps=10% proxy 1.00 fixed 0.70 extra miss 0.30 proxy cost $2.08 # Llama Inst eps=10% proxy 0.95 fixed 0.70 extra miss 0.25 proxy cost $2.08 # Qwen 2.5 Inst eps=10% proxy 1.00 fixed 1.00 extra miss 0.00 proxy cost $2.08 # # proxy route : 4.17 USD, 200 GPU-min # 8B full sweep: 10.42 USD, 500 GPU-min (2.5x) # the grid the whole result is measured on, and the denser check python3 -c " G=[0.00,0.25,0.50,0.75,1.00] D=sorted(G+[0.125,0.375,0.625,0.875]) print('5-point grid :', G) print('9-point check:', D) print('10% region spanning 55-75% of the axis covers', [r for r in G if 0.125 ``` Four limits the paper states about its own numbers, which we repeat because they bound the recommendation. Every run uses LoRA, treated as an approximation of full fine-tuning. Most of the sweep is single-seed, with a 3-seed validation only on math and summarization for the Llama family. The ratio _r_ = 0 is excluded from the cost-asymmetry analysis because it consistently underperforms and needs disproportionately more data when ρ > 1. And the widening of the region with scale is partly a consequence of the relative tolerance definition, visible in the Llama math case where the absolute-tolerance slope is negative. ## What would prove this wrong The claim we are taking from the paper is operational: for a budget counted in annotated samples, a cheap proxy sweep identifies a near-optimal allocation region that transfers to a much larger target, and it beats a fixed 50/50 split by enough to pay for itself many times over. A dated prediction. By **31 December 2027**, a study that runs the same protocol on a model family not in this paper (Llama 3, Qwen 2.5 and Qwen 3 are the three tested), at a 5% tolerance, with budgets at or above 5,000 annotated samples, will report a proxy hit rate above **0.80** on at least two of its tasks. A result showing proxy transfer at or below the fixed-0.5 baseline at 5% tolerance, on a sample-counted budget, falsifies the reading here. The most likely way it breaks is the one the paper flags itself: post-training algorithm rankings can flip across scale, and a family whose small model is far weaker than its large one (the Llama math pattern) is where the region stops behaving. ## Sources - Wang, J., Verma, A., Lin, X., Liu, Z., Chen, N. F., Rus, D., Low, B. K. H. [Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs](https://arxiv.org/abs/2609.01573). arXiv:2609.01573v1, 1 September 2026, EMNLP 2026. Tables 1, 6, 7 and 8, Sections 2.1, 3.1 to 3.4 and Appendices E.4 to E.7; [HTML full text](https://arxiv.org/html/2609.01573). - Lambert, N., et al. [Tülu 3: Pushing Frontiers in Open Language Model Post-Training](https://arxiv.org/abs/2411.15124). arXiv:2411.15124. Source of the Persona IF, RLVR IF and Grade School Math RL-stage data. - Wang, Z., et al. [HelpSteer2: Open-source dataset for training top-performing reward models](https://arxiv.org/abs/2406.08673). arXiv:2406.08673. The RL-stage data for the helpfulness task. - Cobbe, K., et al. [Training Verifiers to Solve Math Word Problems](https://arxiv.org/abs/2110.14168). arXiv:2110.14168. GSM8K, the SFT-stage source and evaluation for the math task. - BLOMEGA. [Two audits of HH-RLHF disagree on how much of it is mislabelled](https://blomega.com/research/hh-rlhf-preference-label-noise-2026/), on the quality side of the same preference-data purchase. - BLOMEGA. [Data annotation research: the latest](https://blomega.com/research/data-annotation-latest-research/). ============================================================================== URL: https://blomega.com/research/web-normalized-annotation-density-persian-2026/ Published: 2026-09-16 | Updated: 2026-09-16 ============================================================================== --- title: "Persian has 1.7% of English's web pages and 3.4 times its news-NER labels" url: https://blomega.com/research/web-normalized-annotation-density-persian-2026/ published: 2026-09-16 updated: 2026-09-16 source: BLOMEGA (https://blomega.com/) --- # Persian has 1.7% of English's web pages and 3.4 times its news-NER labels Lab note · 16 September 2026 · BLOMEGA A review posted **25 August 2026** takes 34 Persian text resources and divides each one's size, task by task, by Persian's share of the crawled web. Common Crawl **CC-MAIN-2026-30** puts Persian on **0.7039%** of HTML pages against English's **40.5782%**, a ratio of **0.01735**. Normalised against that, Persian carries **197.0** times its web-proportional share of news NER labels, **33.8** for dependency parsing, **17.2** for news summarization, and **0.60** for natural-language inference. One of those four is scarce. ## What changed, and when On **25 August 2026**, MohammadHossein Mortazavi, Mostafa Salehi and Hadi Veisi posted [The Annotation Bottleneck in Persian Text NLP: Persian as an Annotation-Scarce Language](https://arxiv.org/abs/2608.24698) (arXiv:2608.24698, 18 pages, 8 tables). It reviews 34 representative Persian text resources available by July 2026 and adds three quantitative cross-checks. The argument is a definitional one with a metric attached. "Low-resource" collapses several distinct shortages into one label: no raw text, no labeled data, no evaluation sets, no tooling, no access. Persian fails only some of those tests, and the paper's contribution is a way to say which. The metric is **web-normalized annotation density** (WNAD), and it is small enough to state in one line: take the ratio of the target language's annotated resource size to a comparable English resource for the same task, then divide it by the two languages' web page-share ratio. A value of 1.0 means the labels are exactly proportional to the language's share of the web. The raw-text side of the argument is settled quickly and is worth having on hand. Persian corpora include Hamshahri at 166,774 categorized newspaper documents, the ParsBERT collection at roughly 3.98 million documents and more than 38 million sentence-like segments, MirasText at about 2.84 million documents and 1.43 billion tokens from more than 250 sites, hmBlogs at nearly 20 million blog posts and over 6.8 billion tokens, and Matina at 72.9 billion preprocessed and deduplicated tokens. A general claim of raw-text absence is not defensible against those numbers. ## The evidence table Two independent July 2026 measurements of web presence, with different denominators. W3Techs counts websites; Common Crawl counts crawled HTML pages by CLD2-identified primary language. The paper is explicit that the two should not be combined, and cites their convergence rather than their sum. | Language | W3Techs, % of websites | CC-MAIN-2026-30, % of HTML pages | Note | Source | | --- | --- | --- | --- | --- | | English | 49.6 | 40.5782 | dominant baseline language | Table 1 | | Turkish | 1.6 | 1.3455 | regional comparison, larger web share | Table 1 | | **Persian** | 0.9 | 0.7039 | within roughly the top 20 under both views | Table 1 | | Arabic | 0.6 | 0.6548 | regional script-sharing comparison | Table 1 | _Persian sits within roughly the top twenty identified languages under both views. That is the number the annotation counts get divided by._ Then the same treatment applied to labels. Each row compares one documented Persian resource against one canonical English resource in the same unit, and divides the resulting ratio by 0.01735. | Task | Persian resource | English comparator | Fa / En | WNAD | Source | | --- | --- | --- | --- | --- | --- | | News NER | NSURL-2019, 1,029,822 tokens | CoNLL-2003 English, 301,418 tokens | 3.42 | 197.0 | Table 8 | | Dependency parsing | PerDT + Seraji, 645,790 tokens | 13 treebanks on the English UD comparison page, 1,102,940 tokens | 0.586 | 33.8 | Table 8 | | News summarization | pn-summary, 93,207 records | CNN/DailyMail standard splits, 312,084 pairs | 0.299 | 17.2 | Table 8 | | **Natural-language inference** | FarsTail, 10,367 examples | SNLI + MultiNLI, ~1.003M examples | 0.0103 | 0.60 | Table 8 | | Web-proportional baseline | – | – | 0.01735 | 1.00 | §8.2 | _Two and a half orders of magnitude separate the best-served and worst-served Persian task. That spread, not the average, is the finding._ ## What the ratio corrects for, and what it does not The naive version of this question is "does Persian have enough labeled data", and it has no answer, because enough compared to what. Comparing absolute counts to English penalises every language for not being English. Comparing to speaker population ignores that NLP consumes text, not speakers. WNAD picks the denominator that matches what a model actually trains on: crawled web pages in that language. _The metric is four numbers and a division. Its value is entirely in the discipline around it: per task, same unit, named comparator, no averaging._ What WNAD does not measure is stated as plainly in the paper as the metric itself, and the caveats are the reason the number is usable rather than decorative. A high density says nothing about domain coverage, annotation quality, licensing, documentation or access. The paper's own task-level profile lists the failures that a volume ratio cannot see: concentration in news, incompatible label inventories across resources, limited coverage of medical, legal and colloquial text, and thin support for varieties beyond standard Iranian Persian. Persian news NER scores 197.0 and is still news NER. The 197.0 is also a lesson in comparator sensitivity, and the authors say so rather than banking the headline. CoNLL-2003 English is 301,418 tokens. It is the canonical English NER benchmark and nothing close to the total of English NER data. Swap the denominator for a larger English NER collection and the ratio falls by whatever factor you chose. This is why the paper reports no cross-task average: the metric is a within-task diagnostic whose comparator has to be named every time it is quoted. One more thing worth recording about provenance. The manuscript states that large language models, including Claude Sonnet 5, were used to improve the clarity and phrasing of the text. That is a disclosure about the writing, not about the numbers, and the resource counts are attributed to named source publications throughout. ## What it means if you are planning where to spend annotation money **"Low-resource" is a purchasing decision disguised as a description.** The four Persian tasks here span 0.60 to 197.0 on the same scale, in the same language, in the same year. Any budget allocated to "Persian" as a unit is allocated blind. The unit that carries information is language crossed with task, and the paper's own construction of it costs one crawl statistic and two published resource sizes. **Compute the ratio before you scope the collection.** For a language and task you are considering, the inputs are: the target language's page share in the latest Common Crawl snapshot, English's page share in the same snapshot, the size of the best existing target-language resource, and the size of a named English comparator in the same unit. Four numbers, all public. The output tells you whether you are filling a genuine gap or adding to an island that is already dense. **The gaps that matter are structural, not numeric.** Persian's shortages, as the paper characterises them, are uneven task and domain coverage, incompatible annotation schemes, access and documentation friction, thin supervision for specialist domains and preference data, and almost nothing beyond standard Iranian Persian. None of those show up in a volume count, and all of them show up in a delivery schedule. Our judgement: for a consented-data programme, the scheme-compatibility and licensing problems are the expensive ones, because they cannot be fixed by collecting more. **Preference data is the named gap.** The paper singles out preference data, specialist domains and non-standard varieties as the places where supervision is thin, and those are exactly the categories that post-training now consumes. The NLI figure of 0.60 is the only one of the four measured tasks below the baseline, and inference-style supervision is closer in kind to preference annotation than news NER is. **Re-run it per snapshot.** The web baseline moves. This computation is pinned to CC-MAIN-2026-30 and to W3Techs in July 2026, and the paper says the visibility figure should be read as dated rather than as a permanent property of the language. ## Check it yourself Every number in the WNAD table is a division of two published counts. The paper has no arXiv HTML rendering, so the tables come from the PDF. ``` # the paper (HTML build is absent; PDF resolves) curl -s -o /dev/null -w "%{http_code}\n" https://arxiv.org/abs/2608.24698 # 200 curl -s -o /dev/null -w "%{http_code}\n" https://arxiv.org/html/2608.24698 # 404 curl -sL -o 2608.24698.pdf https://arxiv.org/pdf/2608.24698 && python3 -c " from pypdf import PdfReader; r=PdfReader('2608.24698.pdf'); print(len(r.pages), 'pages')" # 18 pages # recompute every cell of Table 8 from the two published counts python3 - <<'PY' FA_PAGES, EN_PAGES = 0.7039, 40.5782 # CC-MAIN-2026-30, CLD2 primary language R_web = FA_PAGES / EN_PAGES print(f"R_web = {R_web:.5f}\n") rows = [ # task, Persian units, English units ("News NER", 1_029_822, 301_418), # NSURL-2019 vs CoNLL-2003 English ("Dependency parsing", 645_790, 1_102_940), # PerDT + Seraji vs 13 English UD treebanks ("News summarization", 93_207, 312_084), # pn-summary vs CNN/DailyMail splits ("NLI", 10_367, 1_003_000), # FarsTail vs SNLI + MultiNLI ] for task, fa, en in rows: ratio = fa / en print(f"{task:20} Fa/En {ratio:7.4f} WNAD {ratio / R_web:7.1f}") PY # R_web = 0.01735 # # News NER Fa/En 3.4166 WNAD 197.0 # Dependency parsing Fa/En 0.5855 WNAD 33.8 # News summarization Fa/En 0.2987 WNAD 17.2 # NLI Fa/En 0.0103 WNAD 0.6 # run it for your own language pair: the page-share table is published per crawl # https://commoncrawl.github.io/cc-crawl-statistics/plots/languages # and W3Techs publishes the website-share view # https://w3techs.com/technologies/overview/content_language ``` Two limits on what this reproduces. The Persian and English resource sizes are source-reported, taken from each dataset's own publication rather than recounted from the files, so an error in a source paper propagates. And the English comparator on each row is a choice made by the authors; the dependency-parsing row aggregates 13 treebanks while the NER row uses a single benchmark, which is a difference in kind between rows that the single WNAD column does not show. ## What would prove this wrong The claim worth testing is the general one, not the Persian one: that within-task, web-normalized annotation density varies enough across tasks in a single language to make language-level "low-resource" labels useless for planning. A dated prediction. By **31 December 2027**, applying this computation to any language outside the top five by Common Crawl page share, across at least four tasks with named English comparators, will produce a spread of at least **one order of magnitude** between its highest and lowest task density. If someone runs it on three such languages and finds the per-task densities clustered within a factor of ten, the language-level label is doing more work than we are giving it credit for and this reading is wrong. The result would be most likely to hold for a language whose entire NLP resource base came from one coordinated programme rather than from accreted individual efforts. ## Sources - Mortazavi, M., Salehi, M., Veisi, H. [The Annotation Bottleneck in Persian Text NLP: Persian as an Annotation-Scarce Language](https://arxiv.org/abs/2608.24698). arXiv:2608.24698, 25 August 2026. Tables 1, 7 and 8, Sections 5.2, 8.2 and 10; [PDF](https://arxiv.org/pdf/2608.24698), 18 pages. - [Common Crawl language statistics](https://commoncrawl.github.io/cc-crawl-statistics/plots/languages), crawl CC-MAIN-2026-30. The page-share figures the normalization divides by. - [W3Techs content language usage](https://w3techs.com/technologies/overview/content_language), July 2026. The independent website-share measurement. - Tjong Kim Sang, E. F., De Meulder, F. [Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition](https://aclanthology.org/W03-0419/). CoNLL 2003. The 301,418-token English comparator behind the 197.0 figure. - BLOMEGA. [Unpaid annotation tasks grew 4.2x in ACL papers](https://blomega.com/research/unpaid-annotation-tasks-acl-2018-2025/), on who actually produces the labels these counts describe. - BLOMEGA. [Data annotation research: the latest](https://blomega.com/research/data-annotation-latest-research/). ============================================================================== URL: https://blomega.com/guides/asr-real-world-speech-arabic-southeast-asia-2026/ Published: 2026-09-15 | Updated: 2026-09-15 ============================================================================== --- title: "The best speech recognizer gets 51% of words wrong on Moroccan YouTube speech" url: https://blomega.com/guides/asr-real-world-speech-arabic-southeast-asia-2026/ published: 2026-09-15 updated: 2026-09-15 source: BLOMEGA (https://blomega.com/) --- # The best speech recognizer gets 51% of words wrong on Moroccan YouTube speech Lab note · 15 September 2026 · BLOMEGA On GigaSpeechBench, a benchmark of recent, human-transcribed YouTube speech posted to arXiv on **27 June 2026**, the best of **16** speech recognition systems gets **51.34%** of words wrong on Arabic speech from Morocco and **44.22%** on Algeria. ElevenLabs Scribe v2 scores **2.94%** word error rate on FLEURS Indonesian and **22.91%** on the benchmark's Indonesian subset, 7.8 times worse. Across seven systems scored on both test sets, FLEURS rank and real-world rank correlate at a Spearman of **0.29**. GPT-4o Transcribe goes from third to last. ## What GigaSpeechBench measured, and when **27 June 2026.** A 38-author team from SpeechColab, Shanghai Jiao Tong University, Tsinghua, NTU, UIUC, Alibaba and others posted [GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark](https://arxiv.org/abs/2606.28884). The current version is v3, dated **21 July 2026**. The benchmark totals 680 hours of manually transcribed speech in five modules: low-resource languages, six Chinese dialects, six English accents, 12 vertical domains, and child and older-adult speech. It is released under CC BY 4.0 on [Hugging Face](https://huggingface.co/datasets/speechcolab/GigaSpeechBench), ungated, with evaluation code on [GitHub](https://github.com/SpeechColab/GigaSpeechBench). The low-resource module is the one that matters for localization. It covers Arabic speech from seven regions (Iraq, Algeria, the UAE, Egypt, Morocco, Saudi Arabia, Syria), five Southeast Asian languages (Indonesian, Malay, Filipino, Vietnamese, Thai), plus Japanese and Korean. Three design choices separate it from FLEURS and Common Voice, which are read speech. The audio comes from YouTube and includes multi-speaker conversation and noisy rooms. All of it was published within the year before collection, which the authors say limits overlap with existing training data. And the reference transcripts were written by paid annotators through a professional annotation company, with Chinese and English translations added for 11 of the languages so speech translation can be scored too. The authors ran nine commercial APIs (Microsoft Azure Speech, Google Chirp 3, OpenAI GPT-4o Transcribe, Gemini 3.0 Flash, ElevenLabs Scribe v2, Qwen3-ASR-Flash, Qwen3.5-Omni-Plus and two ByteDance Seed-ASR versions) and open models including Whisper Large v3, NVIDIA NeMo Canary, Meta OmniASR-LLM-3B, Qwen3-ASR 1.7B, FunASR and Dolphin. Table 2, the low-resource table, has 16 rows once Deepgram Nova 3 is included. The same paper also reports FLEURS and Common Voice scores for most of these systems, which is what makes a like-for-like comparison possible. ## How far does each system fall from FLEURS to real speech? Take one system and hold everything else fixed. ElevenLabs Scribe v2, which ElevenLabs lists at $0.22 per hour of audio, has the best FLEURS average of the seven systems scored on both test sets: **5.38%** across eight shared languages. On GigaSpeechBench the same eight languages average **24.90%**. Japanese character error rate goes from 2.41% to 29.95%, a factor of 12.4. Malay goes from 3.92% to 38.52%, a factor of 9.8. _The two test sets differ in content as well as recording conditions, so the ratio mixes domain shift with acoustic difficulty. That is exactly the mix a dubbing pipeline meets on real source video._ Scribe v2 is not an outlier. Gemini 3.0 Flash averages 5.82% on FLEURS and 28.80% on the real speech. Google Chirp 3 goes from 6.21% to 24.91%. Whisper Large v3 goes from 7.81% to 34.65%. On average every system tested on both sets gets at least 2.38 times worse, and Azure is the one with the smallest drop. The degradation is not uniform, and that is the part that should change procurement. The two systems with the worst and third-best FLEURS averages swap places with the systems at the bottom of the real-speech table. _A vendor that tops a read-speech leaderboard can finish last on the audio you will actually send it. Azure has the weakest FLEURS average here and degrades least._ The rank calculation is ours. Differences in rank between FLEURS and GigaSpeechBench are 0 for Scribe v2, Whisper and OmniASR, 2 for Gemini and Chirp 3, and 4 for GPT-4o Transcribe and Azure. That gives a sum of squared differences of 40 and a Spearman coefficient of 1 minus 240 over 336, which is 0.29. Seven systems is a small sample, so read it as "weak" rather than as a precise figure. | Subset | Best system | Best | Scribe v2 | Gemini 3.0 Flash | Whisper Large v3 | Source | | --- | --- | --- | --- | --- | --- | --- | | Arabic, Morocco | Qwen3.5-Omni-Plus | 51.34 | 60.06 | 51.99 | 91.89 | arXiv:2606.28884 Table 2 | | Arabic, Algeria | Gemini 3.0 Flash | 44.22 | 50.43 | 44.22 | 72.02 | arXiv:2606.28884 Table 2 | | Arabic, Egypt | Qwen3.5-Omni-Plus | 37.12 | 44.44 | 41.22 | 69.78 | arXiv:2606.28884 Table 2 | | Arabic, Iraq | Qwen3.5-Omni-Plus | 28.54 | 38.67 | 36.55 | 51.04 | arXiv:2606.28884 Table 2 | | Arabic, UAE | GPT-4o Transcribe | 26.26 | 46.10 | 45.06 | 68.41 | arXiv:2606.28884 Table 2 | | Arabic, Saudi Arabia | Qwen3.5-Omni-Plus | 16.56 | 33.33 | 20.10 | 32.79 | arXiv:2606.28884 Table 2 | | Arabic, Syria | Qwen3.5-Omni-Plus | 13.76 | 14.73 | 14.40 | 19.12 | arXiv:2606.28884 Table 2 | | Malay | FunASR-Realtime | 25.20 | 38.52 | 40.92 | 46.15 | arXiv:2606.28884 Table 2 | | Japanese (CER) | FunASR-Realtime | 25.44 | 29.95 | 39.84 | 39.28 | arXiv:2606.28884 Table 2 | | Filipino | FunASR-Realtime | 23.69 | 27.15 | 29.17 | 30.88 | arXiv:2606.28884 Table 2 | | Indonesian | FunASR-Realtime | 14.87 | 22.91 | 24.18 | 27.40 | arXiv:2606.28884 Table 2 | | Thai | FunASR-Realtime | 10.76 | 13.90 | 26.58 | 27.02 | arXiv:2606.28884 Table 2 | | Korean (CER) | FunASR-Realtime | 9.92 | 11.81 | 16.78 | 18.53 | arXiv:2606.28884 Table 2 | | Vietnamese | Chirp 3 | 9.63 | 10.52 | 11.69 | 18.17 | arXiv:2606.28884 Table 2 | Three things in that table are not in the paper's own discussion. **Arabic is a regional problem, not a language problem.** The best system's error rate spans 13.76% (Syria) to 51.34% (Morocco) inside one language label. A vendor quoting a single "Arabic" accuracy figure is averaging across a 3.7-fold spread. On Common Voice Arabic, Table 3 of the same paper puts Scribe v2 at 10.14%; its seven-region GigaSpeechBench Arabic average is 41.11%. **No single vendor wins.** FunASR-Realtime has the best Southeast Asian average (16.85%) and Japanese and Korean scores, then averages 55.11% on Arabic, twelfth of 15 systems with Arabic results. Qwen3.5-Omni-Plus has the best Arabic average (32.80%). GPT-4o Transcribe is the best system in the table on UAE speech (26.26%) and the fourth worst of the 15 systems with Egyptian results (64.23%). **An open 1,600-language model is not a long-tail solution out of the box.** Meta's OmniASR-LLM-3B averages 42.49% across the eight shared languages and 65.52% on Morocco. Meta positioned the family as transcription for more than 1,600 languages when it launched in November 2025. _[@AIatMeta](https://www.youtube.com/@AIatMeta) on YouTube, 10 November 2025: "Introducing Meta Omnilingual Automatic Speech Recognition | Transcription for 1,600+ languages". The launch claim is coverage. GigaSpeechBench measures accuracy on recent speech, where the 3B LLM variant averages 42.49% across eight languages and 65.52% on Moroccan Arabic._ ## Where a transcription error goes in a dubbing pipeline A cascaded AI dub is transcription, then translation, then synthesis, then a mix. Each stage consumes the previous stage's output as ground truth. A transcription error is not diluted downstream. It is translated fluently and then spoken in a confident, well-timed voice. _End-to-end speech translation does not escape the problem either. UAE, Malay and Moroccan speech score lowest into English, Saudi and Vietnamese highest._ The speech translation numbers in Table 10 back that up without any cascade assumption. Qwen3.5-Omni-Plus averages 49.87 chrF++ into English across 11 subsets and Gemini 3 Flash Preview 49.82. Microsoft's Azure translation service averages 40.05 and SeamlessM4T v2 Large 34.19. For the UAE subset the best score is 41.54, and SeamlessM4T v2 gets 24.30. Moroccan speech is the hardest subset for recognition and the third hardest for translation into English, after the UAE (41.54) and Malay (44.89) subsets. ## What to change if you dub or caption long-tail locales **Stop accepting FLEURS numbers in vendor evaluations.** A 0.29 rank correlation means a FLEURS leaderboard tells you little about which vendor to buy for real source video. Ask for a score on held-out audio that looks like yours: recent, conversational, noisy, and from the specific country. GigaSpeechBench's own filter for recency is "within the past year", and that is a reasonable default for a private test set too. **Pick the recognizer per locale, not per contract.** On this data the best system for Saudi and Syrian speech is not the best for Malay, Thai or Japanese, and the best for UAE speech is second worst on Egypt. A pipeline that routes each locale to its own recognizer beats one vendor across the board. The cost is integration work and a per-locale test set, both of which are cheaper than re-dubbing. **Budget human transcription review for North African Arabic.** At 44% to 51% best-case WER on Algerian and Moroccan speech, automatic transcription is a draft. Our judgement: for those two subsets, native-speaker correction of the transcript should be priced as a standard line item, not as an exception. For Saudi and Syrian speech, at 13.76% to 16.56%, spot-checking is a defensible compromise. **Treat "supports N languages" as a coverage claim, not an accuracy claim.** OmniASR-LLM-3B is part of a family marketed for 1,600+ languages and averages 42.49% on eight widely spoken ones in this test. Coverage and accuracy are separate columns in any vendor comparison. **The fix is in-language, in-domain data.** The authors attribute the gap to the difference between read benchmark speech and recent real-world audio. The only lever a buyer controls is adaptation data that matches the target: transcribed, consented recordings from the region and register you are dubbing. That is our reading of the result, not a claim in the paper. ## Check it yourself The dataset is ungated and the per-system hypothesis files are published next to the audio, so the central numbers can be recomputed without calling any API. ``` # 1. pull the Moroccan subset and every system's published output pip install -U "huggingface_hub[cli]" huggingface-cli download speechcolab/GigaSpeechBench --repo-type dataset \ --include "Low-Resource-Languages/data/MAR/*" "Low-Resource-Languages/results/*" \ --local-dir gsb # 2. get the scoring code, including the per-region text normalizers git clone https://github.com/SpeechColab/GigaSpeechBench ls GigaSpeechBench/text_norm # MAR.py, DZA.py, EGY.py, ... one per subset cat GigaSpeechBench/run_ASR.sh # the exact normalize + compute_wer invocation # 3. inspect the reference file: segments carry text, text_en, text_zh, # speaker, gender, age_group and emotion; audios carry duration python3 - <<'PY' import json d = json.load(open("gsb/Low-Resource-Languages/data/MAR/metadata.json")) segs = sum(len(a["segments"]) for a in d["audios"]) hours = sum(a.get("duration", 0) for a in d["audios"]) / 3600 print(len(d["audios"]), "recordings,", segs, "segments,", round(hours, 1), "hours") PY # 4. the rank correlation in this note, from Tables 2 and 4 python3 - <<'PY' fleurs = {"Scribe v2":5.38,"Gemini 3.0 Flash":5.82,"GPT-4o Transcribe":5.94,"Chirp 3":6.21, "Whisper Large v3":7.81,"OmniASR 3B":10.29,"Azure":10.59} wild = {"Scribe v2":24.90,"Chirp 3":24.91,"Azure":25.21,"Gemini 3.0 Flash":28.80, "Whisper Large v3":34.65,"OmniASR 3B":42.49,"GPT-4o Transcribe":44.59} rank = lambda d: {k:i+1 for i,k in enumerate(sorted(d, key=d.get))} rf, rw = rank(fleurs), rank(wild) n = len(rf); d2 = sum((rf[k]-rw[k])**2 for k in rf) print("spearman", round(1 - 6*d2/(n*(n*n-1)), 2)) # 0.29 PY ``` The eight-language averages in the slope chart are plain means of the Table 2 and Table 4 cells for Egyptian Arabic, Indonesian, Malay, Filipino, Vietnamese, Thai, Japanese and Korean. The repository README reports its leaderboard with a filter that drops segments of 0.5 seconds or less, and its numbers match Table 2. ## What would prove this wrong The claim under test is that read-speech benchmarks do not predict which recognizer performs best on real source video in long-tail locales. It is wrong if, by **30 June 2027**, a published re-run of GigaSpeechBench and FLEURS covering at least seven common systems and the same eight languages shows a Spearman rank correlation of 0.7 or higher. Today it is 0.29. A second prediction, marked as judgement: no system on the public GigaSpeechBench leaderboard will score below 30% WER on the Moroccan subset by 30 June 2027, using the repository's own normalizer. The best today is 51.34%. Getting under 30% needs regional training data that does not exist at scale in public corpora, and a year is short for that to change. If a system posts under 30% by that date, this reading was wrong. ## FAQ ### How accurate is speech recognition on Arabic dialects in 2026? Not accurate enough to feed a dub without human correction. On GigaSpeechBench the best of 16 systems scores 51.34% WER on Arabic speech from Morocco, 44.22% on Algeria, 37.12% on Egypt, 28.54% on Iraq and 26.26% on the UAE. Saudi Arabia (16.56%) and Syria (13.76%) are the only Arabic subsets where the best system stays under 20%. The best seven-subset Arabic average is Qwen3.5-Omni-Plus at 32.80%. ### Do FLEURS scores predict how an ASR system performs on real speech? Poorly. Across seven systems scored on both test sets for the same eight languages, the Spearman rank correlation is 0.29. GPT-4o Transcribe ranks third on FLEURS at 5.94% and last on real speech at 44.59%. Azure ranks last on FLEURS at 10.59% and third on real speech at 25.21%. ### Which speech recognition API is best for Southeast Asian languages? On GigaSpeechBench's YouTube speech, FunASR-Realtime has the best Southeast Asian average at 16.85% WER, followed by Qwen3.5-Omni-Plus at 19.61% and Chirp 3 at 20.87%. ElevenLabs Scribe v2 averages 22.60% and Azure 22.68%. Malay is the hardest of the five, with a best score of 25.20%. ### Why does AI dubbing quality drop in low-resource locales? Because transcription, the first stage, is where real-world speech breaks, and every later stage inherits the errors. Scribe v2 moves from 2.94% WER on FLEURS Indonesian to 22.91% on recent YouTube Indonesian, and from 2.41% to 29.95% CER on Japanese. Translation and synthesis downstream cannot recover words the recognizer never heard. ## Sources - Tu, Yang, Wang et al., [GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark](https://arxiv.org/abs/2606.28884), arXiv:2606.28884, submitted 27 June 2026, v3 21 July 2026, [HTML](https://arxiv.org/html/2606.28884v3). Table 2 (low-resource WER/CER, 16 systems), Table 3 (Common Voice), Table 4 (FLEURS), Table 10 (speech translation into English), sections 3 and 4 on collection, recency and annotation, ethics statement on Creative Commons sourcing and paid annotators. CC BY 4.0. - [speechcolab/GigaSpeechBench on Hugging Face](https://huggingface.co/datasets/speechcolab/GigaSpeechBench), ungated, last modified 9 July 2026, read 15 September 2026. Per-subset `metadata.json`, `audio.tar.gz` and per-system `results/*.json` and `results_trans/*.json`. - [SpeechColab/GigaSpeechBench on GitHub](https://github.com/SpeechColab/GigaSpeechBench), read 15 September 2026. README leaderboard with regional averages (Qwen3.5-Omni-Plus 32.80% Arabic, FunASR-Realtime 16.85% Southeast Asian, 17.68% East Asian), `run_ASR.sh`, `scripts/compute_wer.py`, `text_norm/`. - Omnilingual ASR Team, [Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages](https://arxiv.org/abs/2511.09690), arXiv:2511.09690, 12 November 2025, and the [AI at Meta launch video](https://www.youtube.com/watch?v=ab-GIqDQn7k), 10 November 2025. - [ElevenLabs API pricing](https://elevenlabs.io/pricing/api), Scribe v2 at $0.22 per hour, as recorded in BLOMEGA's [dubbed-minute cost note](https://blomega.com/guides/cost-of-a-dubbed-minute-2026/) on 10 September 2026. - Spearman rank correlation, eight-language averages and FLEURS-to-real ratios: BLOMEGA calculations from the tables above, 15 September 2026. Code in "Check it yourself". Related BLOMEGA guides: [A dubbed minute costs $0.33 to $9.00](https://blomega.com/guides/cost-of-a-dubbed-minute-2026/) · [A 30-trillion-token corpus buys Basque a 160-million-parameter model](https://blomega.com/guides/per-language-token-ceiling/) · [Your multilingual LLM judge prefers the machine translation](https://blomega.com/guides/multilingual-llm-judge-translationese-bias/) · [Consented AI training data providers](https://blomega.com/guides/consented-ai-training-data-providers/). ============================================================================== URL: https://blomega.com/guides/nimdzi-100-2026-language-industry-mid-tier/ Published: 2026-09-15 | Updated: 2026-09-15 ============================================================================== --- title: "The language industry's mid-tier shrank 4.3% in 2025 while its top 10 grew 3.6%" url: https://blomega.com/guides/nimdzi-100-2026-language-industry-mid-tier/ published: 2026-09-15 updated: 2026-09-15 source: BLOMEGA (https://blomega.com/) --- # The language industry's mid-tier shrank 4.3% in 2025 while its top 10 grew 3.6% Lab note · 15 September 2026 · BLOMEGA The 2026 Nimdzi 100 puts the language industry at **USD 72.6 billion** in 2025 and **USD 73.4 billion** in 2026, growth of 1.1%, with under 1% a year after that to **USD 76.1 billion** by 2030. Inside that flat total, providers ranked 51 to 100 lost **4.3%** of combined revenue in 2025, the first segment decline Nimdzi has recorded since 2021, while the top 10 grew **3.6%**. The same report prints three different growth rates for the top 100 as a whole, and two different optimistic 2030 forecasts. The segment split is the number that stays consistent. ## What Nimdzi published, and what moved after it **15 April 2026.** Nimdzi Insights [published the 2026 Nimdzi 100](https://multilingual.com/2026-nimdzi-100-published-free-to-access/), free to read, written by Marjolein Groot Nibbelink and Laszlo K. Varga. It ranks the 100 largest language providers by latest fiscal-year revenue and sizes the market. For this edition Nimdzi changed its qualifiers so that pure technology providers compete in the same ranking as service companies. DeepL, Phrase, Smartling and XTM International now appear alongside TransPerfect and LanguageLine. The headline macro numbers: the market reached an estimated USD 72.6 billion in 2025; 2026 is projected at USD 73.4 billion; the "realistic" forecast reaches USD 76.1 billion by 2030. Our calculation: that is a compound annual rate of 0.95% from 2025, or 0.91% from 2026. The top 100 made up 19.8% of the market by Nimdzi's figure, earning just over USD 14.3 billion. The top 10 alone held 9.9% of the market and USD 7.17 billion; the next 90 held 9.8% and USD 7.14 billion. **7 May 2026.** Three weeks later DeepL, ranked 21st at an estimated USD 197.0 million, [announced about 250 job cuts](https://the-decoder.com/ai-translation-company-deepl-cuts-around-250-jobs-to-rebuild-as-an-ai-native-organization/), reported as [roughly a quarter of its workforce](https://gigazine.net/gsc_news/en/20260508-deepl-lay-off-250/), to rebuild as an "AI-native" organization while acquiring Mixhalo's audio streaming team for real-time voice translation. It is one company, but it is the pattern Nimdzi described: the report says many providers cut in-house linguistic and project management staff, "sometimes by 20% to 25%", against what it calls threefold productivity increases from AI. ## Who grew and who shrank in 2025? Nimdzi reports growth by ranking segment in its "Growth by Ranking Segment" section. The pattern is a clean gradient from the top down, breaking negative below rank 50. _The report gives the top-100 combined figure as 1.3% in this section, 1.1% in another and +0.2% in a third. See Table 2. The segment figures do not conflict anywhere in the report._ Nimdzi's explanation for the 51 to 100 decline is specific. Many mid-market providers depend on inbound public-sector contracts and RFPs. In 2025, US and Canadian budget freezes and "English-only rhetoric" cut government discretionary spend, and some of these firms saw revenue drops "of up to 25%". The dollar's fall against the euro caused margin losses they could not absorb. And the report says some cannot afford the technology investment to stay competitive and rely on third-party tools instead. The ranking itself shows where the large firms sit. Nimdzi publishes direction icons for rank change and revenue change rather than percentages, and marks each revenue figure with a note letter. The rows below are a selection relevant to localization, dubbing and AI data buyers. | Rank | Company | Main business (Nimdzi) | 2025 revenue | Rank change | Revenue change | Note | Source | | --- | --- | --- | --- | --- | --- | --- | --- | | 1 | TransPerfect | translation, life sciences, legal | 1,320.0 | same | up | v | nimdzi.com ranking | | 2 | LanguageLine Solutions | interpreting, translation | 1,100.0 | same | same | v | nimdzi.com ranking | | 3 | RWS | AI-powered translation software, translation, patents | 909.2 | up | up | v | nimdzi.com ranking | | 7 | Lionbridge | translation, life sciences, technology, games | 513.9 | down | down | e | nimdzi.com ranking | | 8 | Translate Plus | translation, dubbing, marketing | 399.0 | same | up | e | nimdzi.com ranking | | 13 | Welo Global | translation, data and AI | 325.2 | down | down | v | nimdzi.com ranking | | 14 | Centific | localization, data curation | 300.0 | up | up | v | nimdzi.com ranking | | 16 | Iyuno | media localization | 280.0 | down | down | e | nimdzi.com ranking | | 18 | Appen | data company | 230.8 | up | down | v | nimdzi.com ranking | | 21 | DeepL | AI-powered translation software | 197.0 | same | same | e | nimdzi.com ranking | | 22 | Pixelogic Media | media localization | 191.3 | same | same | e | nimdzi.com ranking | | 24 | VSI | media localization | 137.0 | up | up | v | nimdzi.com ranking | | 26 | GTCOM | language technology, data and AI | 126.1 | up | up | v | nimdzi.com ranking | | 27 | Dubbing Brothers | dubbing, voiceovers, subtitling | 122.9 | down | same | e | nimdzi.com ranking | | 30 | Smartling | translation management system | 96.7 | new | up | v | nimdzi.com ranking | | 31 | Visual Data Media Services | media localization | 94.5 | same | same | e | nimdzi.com ranking | That table allows one check the report's prose does not make. Nimdzi says data-for-AI "is a new industry outpacing the growth of the narrowly defined language industry". Among the four top-30 firms whose main business Nimdzi lists as including data or AI (Welo Global, Centific, Appen, GTCOM), two show revenue up and two show revenue down. Among the five top-31 firms whose main business is listed primarily as media localization or dubbing (Iyuno, Pixelogic, VSI, Dubbing Brothers, Visual Data), one shows revenue up, three flat and one down. Translate Plus, which lists dubbing as one of four lines, shows revenue up. The sector claim may well hold across the whole market. At the level of the largest named firms, it is a split, not a trend. | Quantity | Value in one section | Value in another | Source | | --- | --- | --- | --- | | Top 100 combined growth, 2025 | 1.1% ("Growth or No Growth") | 1.3% ("Growth by Ranking Segment"); +0.2% in the final year of the historic top-100 trendline ("Revenue Concentration") | nimdzi.com | | Optimistic 2030 market size | USD 96.9 billion ("TL;DR") | USD 86.2 billion ("Market Sizing") | nimdzi.com | | Top 100 providers reporting growth | 53 grew | "23 of those 52 providers reported double-digit percentage growth" | nimdzi.com | | Top 100 share of market | 19.8% (stated) | 19.7% computed from USD 14.3B over USD 72.6B; the gap is rounding | nimdzi.com, BLOMEGA arithmetic | None of these change the direction of the story. They do mean that any single "the language industry grew X%" figure quoted from this report should carry its section name. The optimistic forecasts imply very different worlds: USD 96.9 billion by 2030 is 5.9% a year from 2025, USD 86.2 billion is 3.5% a year. Our arithmetic. ## How AI pricing pressure lands on the middle of the market The report's survey sections describe the mechanism. Clients push for cost reductions "sometimes expecting cuts of up to 75% via AI". Per-word pricing stops making sense when the words are machine-produced. The response depends on having capital. _Buyers outnumber sellers at the top and sellers outnumber buyers across the whole survey. That asymmetry is what consolidation looks like before it shows up in market share: the top 10 gained only 0.3 points of share in 2025, from 9.6% to 9.9%._ Pricing is where a buyer feels this first. Nimdzi reports that nearly half of surveyed providers had introduced value-based pricing or separate rates for human and pure technology work in 2025, and that more than three quarters plan to change at least some pricing models. Translation memory discounts are described as becoming obsolete, replaced by "packages of specialist hours for expert review, auditing, and post-editing". Procurement departments, the report says, still ask for per-word rates because they are easy to compare. The service mix has already shifted. Of surveyed providers, 72.0% offer editing of AI-generated content, 70.3% AI-generated translation, 51.7% transcreation, 40.7% cultural adaptation, 40.7% AI-powered QA and 39.0% AI dubbing. Prompt engineering is offered by only 16.2% but is listed among the strongest demand growth areas, alongside remote interpreting and data and AI services. Demand is concentrated: North America bought USD 30.6 billion of services in 2025 and Europe USD 25.5 billion, together 77% of the market. _The regional figures sum exactly to USD 72.6 billion. Africa and the Middle East together bought 2.1% of all language services._ ## What this means for a localization buyer in 2026 **Vendor risk now correlates with rank.** If a significant share of your work sits with a provider in the 51 to 100 band, especially one with public-sector exposure, the 4.3% segment decline and Nimdzi's "up to 25%" drops for some firms are a continuity question, not a market-watching one. Ask for a named backup team and a data-return clause. Our judgement: the 33% of surveyed providers "open to selling" makes change of control within a contract term a realistic scenario. **You have room to negotiate, and the unit is changing.** Providers under a price shock of up to 75% expected cuts, planning to replace per-word rates, are open to hourly, platform or outcome-based terms. The useful move is to separate the machine line from the human line in the quote, because nearly half of providers already price them separately. Buying "AI translation plus 20 hours of expert review per month" makes the review visible and auditable in a way a blended per-word rate never did. **Use 60% of the headline, not the headline.** Nimdzi itself advises commercial providers to treat about 60% of the market as addressable, because some work is in-house (it cites the EU's roughly 5,000 staff translators and interpreters) and supplier revenue is partly counted twice. On 2025 figures that is about USD 43.6 billion. Any business case built on "a USD 72.6 billion market" includes about 40% that no outside vendor can win. **Treat "data-for-AI is the growth engine" as a sector thesis to test.** The report asserts it and describes demand shifting toward expert validation "up to PhD-level". The ranking's own direction icons split evenly across the four top-30 data-and-AI firms. If you are choosing a language provider for AI training or evaluation data, ask for revenue and headcount in that line specifically, not for the group's total. **Media localization is flat at the top.** Of the five top-31 firms Nimdzi classifies primarily as media localization or dubbing specialists, one shows revenue up. Our reading, not Nimdzi's: dubbing budgets are being redistributed toward AI voice and platform-owned tooling faster than they are growing, which is consistent with Prime Video, YouTube and Meta building dubbing features in-house. ## Check it yourself Both the report and the ranking are public HTML. The ranking's direction columns are images, so read the file names. ``` # 1. the ranking table, with direction icons decoded from image file names curl -sL -A "Mozilla/5.0" https://www.nimdzi.com/preliminary-2026-nimdzi-100-ranking/ -o ranking.html python3 - <<'PY' import re, html h = open("ranking.html", encoding="utf-8", errors="ignore").read() table = re.search(r"", h, re.S).group(0) def cell(c): m = re.search(r"uploads/[\d/]+/(UP|DOWN|SAME|NEW)_", c) return m.group(1).lower() if m else html.unescape(re.sub(r"<[^>]+>", "", c)).strip() for row in re.findall(r"", table, re.S)[1:32]: c = [cell(x) for x in re.findall(r"]*>(.*?)", row, re.S)] # columns: id, rank, rank change, company, country, revenue, revenue change, note, business print(c[1], c[3], c[5], "rank:", c[2], "revenue:", c[6], c[7]) PY # 2. find every growth figure in the report text, with its section curl -sL -A "Mozilla/5.0" https://www.nimdzi.com/nimdzi-100-2026/ \ | python3 -c "import sys,re,html;t=html.unescape(re.sub(r'<[^>]+>',' ',sys.stdin.read()));t=re.sub(r'\s+',' ',t);[print('-',m.group(0)) for m in re.finditer(r'[^.]*(4\.3%|1\.1%|1\.3%|\+0\.2%|96\.9|86\.2|76\.1)[^.]*\.',t)]" # 3. the arithmetic in this note python3 -c "print(round(((76.1/72.6)**0.2-1)*100,2), round(((96.9/72.6)**0.2-1)*100,2), round(((86.2/72.6)**0.2-1)*100,2), round(72.6*0.6,1), round(14.31/72.6*100,1))" # 0.95 5.94 3.49 43.6 19.7 ``` Nimdzi updates the ranking page as late reports come in, so record the date you read it. Our reading was 15 September 2026. ## What would prove this wrong The claim under test is that AI price pressure is splitting the language industry by scale, with the mid-tier losing ground while the top tier grows. It is wrong if the 2027 Nimdzi 100, expected around April 2027, reports that the 51 to 100 segment returned to growth of 0% or more in 2026 while the top 10 grew less than 2%. That combination would mean 2025 was a public-sector shock, not a structural split. A second prediction, marked as judgement: at least one provider ranked between 51 and 100 in the 2026 edition will be acquired by a top-20 provider and announced before **30 April 2027**. With about 60% of the top 100 actively buying and 33% of all surveyed providers open to selling, we think the deal is more likely than not. If none is announced by that date, the reading was wrong. ## FAQ ### How big is the language services industry in 2026? USD 73.4 billion projected for 2026, up from an estimated USD 72.6 billion in 2025, per the 2026 Nimdzi 100. That is 1.1% growth, with under 1.0% a year after that to USD 76.1 billion by 2030. Nimdzi advises treating about 60% of the total as addressable for commercial providers, about USD 43.6 billion on the 2025 figure. ### Is the translation and localization industry shrinking because of AI? Not in aggregate, but parts of it are. The top 10 grew 3.6% in 2025 and the top 50 grew 2.0%, while providers ranked 51 to 100 declined 4.3%, the first segment decline since 2021, and providers outside the top 100 declined a weighted 1.1%. Nimdzi attributes the mid-tier fall to public-sector dependence, US and Canadian budget freezes, currency moves and client demands for AI-driven price cuts of up to 75%. ### How are translation companies changing their pricing in 2026? Away from per-word rates. More than three quarters of surveyed providers plan to change some pricing models, nearly half already price human and pure technology work separately, and hourly packages of specialist review are replacing translation memory discounts. Procurement teams still often ask for per-word rates. ### Which language companies are the largest in 2026? By 2025 revenue in the Nimdzi ranking: TransPerfect USD 1,320.0 million, LanguageLine Solutions USD 1,100.0 million, RWS USD 909.2 million, Keywords Studios USD 850.0 million, Sorenson Communications USD 800.0 million, Propio USD 566.5 million and Lionbridge USD 513.9 million. DeepL is 21st at an estimated USD 197.0 million. ## Sources - Nimdzi Insights, [The 2026 Nimdzi 100](https://www.nimdzi.com/nimdzi-100-2026/), by Marjolein Groot Nibbelink and Laszlo K. Varga, read 15 September 2026. Sections used: TL;DR, Growth or No Growth, Market Sizing, Revenue Concentration, Growth by Ranking Segment, Most Productive Companies, Top Generative AI Services, Staffing and Talent, Changes in Pricing Models 2025-2026, M&A in 2025, Growth Services, Data-for-AI and Multimodal Language Technologies sector analyses. - Nimdzi Insights, [The Preliminary 2026 Nimdzi 100 Ranking](https://www.nimdzi.com/preliminary-2026-nimdzi-100-ranking/), read 15 September 2026. Rank, 2025 revenue in USD million, rank-change and revenue-change icons, note letters, main business. - MultiLingual, [2026 Nimdzi 100 Report Published and Free to Access](https://multilingual.com/2026-nimdzi-100-published-free-to-access/), 15 April 2026. Publication date, 14,000 words, 35 charts. - The Decoder, [AI translation company DeepL cuts around 250 jobs to rebuild as an "AI-native" organization](https://the-decoder.com/ai-translation-company-deepl-cuts-around-250-jobs-to-rebuild-as-an-ai-native-organization/), 7 May 2026, and GIGAZINE, [DeepL to lay off approximately 250 employees, about 25% of its workforce](https://gigazine.net/gsc_news/en/20260508-deepl-lay-off-250/), 8 May 2026. Mixhalo team acquisition, San Francisco office. - Compound growth rates, regional shares, addressable-market figure and company-level direction counts: BLOMEGA calculations from sources 1 and 2, 15 September 2026. Related BLOMEGA guides: [A dubbed minute costs $0.33 to $9.00](https://blomega.com/guides/cost-of-a-dubbed-minute-2026/) · [Localization is the new default](https://blomega.com/guides/localization-the-new-default/) · [Consented AI training data providers](https://blomega.com/guides/consented-ai-training-data-providers/). ============================================================================== URL: https://blomega.com/guides/omnilingual-mt-resource-tier-results/ Published: 2026-09-15 | Updated: 2026-09-15 ============================================================================== --- title: "Meta's 8B Omnilingual MT beats Llama 3 70B into mid-resource languages and loses into zero-resource ones" url: https://blomega.com/guides/omnilingual-mt-resource-tier-results/ published: 2026-09-15 updated: 2026-09-15 source: BLOMEGA (https://blomega.com/) --- # Meta's 8B Omnilingual MT beats Llama 3 70B into mid-resource languages and loses into zero-resource ones Lab note · 15 September 2026 · BLOMEGA Translating out of English on Meta's BOUQuET benchmark, OMT-LLaMA 8B scores **45.8 chrF++** into mid-resource languages against **37.2** for Llama 3 70B, a lead of 8.6 points, and 30.8 against 23.7 into low-resource ones. Into zero-resource languages, with under 1,000 parallel documents, the order flips: Llama 3 70B scores **14.3** and OMT-LLaMA 8B **12.6**. No model in the paper's Table 9.2 exceeds 14.3 there. Specialization buys a lot where some parallel data exists, and nothing where none does. ## What Meta released on 17 March 2026, and what it did not **17 March 2026.** Meta's FAIR team posted [Omnilingual MT: Machine Translation for 1,600 Languages](https://arxiv.org/abs/2603.16309) (arXiv:2603.16309), revised to v3 on 7 May 2026. The paper describes two model families built on Llama 3: OMT-LLaMA, decoder-only, at 1B, 3B and 8B parameters, and OMT-NLLB, a 3B encoder-decoder on the OmniSONAR embedding space. It reports "non-trivial performance when translating from 1,600 and into about 1,200 languages" and says the number of languages modern models "understand sufficiently well" doubles from about 200 to over 400. The paper also ships evaluation infrastructure: BOUQuET, a multilingual evaluation set built from scratch, Met-BOUQuET with human quality judgements across 161 language directions, the BLASER 3 reference-free quality estimator and the OmniTOX toxicity classifier. BOUQuET and Met-BOUQuET are [published on Hugging Face](https://huggingface.co/datasets/facebook/bouquet) with a [leaderboard](https://huggingface.co/spaces/facebook/bouquet). What we could not find is the models. On 15 September 2026, Hugging Face searches for "omnilingual" and "OMT-LLaMA" returned Omnilingual ASR models and community conversions of them, and no OMT translation weights. GitHub's facebookresearch organization has an `omnilingual-asr` repository and no MT counterpart. The paper's conclusion encourages the community to use the "OMT-LLaMA, OMT-NLLB, BLASER 3 and OmniTOX recipes". That is a recipe release plus an evaluation release, as far as we can verify. This matters because the speech side is open. [Omnilingual ASR](https://arxiv.org/abs/2511.09690), released in November 2025, is open source, covers 1,600+ languages including more than 500 never before supported by ASR, and ships models from 300M to 7B parameters. The paper suggests cascading it with Omnilingual MT for speech translation. A localization team can run the first half of that cascade today and not the second. ## Where does specialization help, by resource tier? Headline averages hide the shape. Table 9.2 of the paper splits BOUQuET chrF++ by the resource level of the non-English language, in both directions, for 18 systems. The rows below are the ones a localization team would actually weigh. | System | Size | En-YY high | mid | low | v. low | zero | En-YY total | XX-En zero | Source | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | OMT-LLaMA | 8B | 60.7 | 45.8 | **30.8** | **18.8** | 12.6 | **32.8** | 21.6 | Table 9.2 | | OMT-NLLB | 3B | 61.1 | **47.6** | 27.3 | 17.5 | 11.5 | 31.9 | 23.2 | Table 9.2 | | OMT-LLaMA | 1B | 56.5 | 42.2 | 25.1 | 14.9 | 12.3 | 29.1 | 19.9 | Table 9.2 | | NLLB-200 | 3B | 62.9 | 46.8 | 24.6 | 17.5 | 13.2 | 31.3 | 22.7 | Table 9.2 | | GPT-OSS | 120B | **63.6** | 43.7 | 26.0 | 16.0 | 13.4 | 30.8 | 24.3 | Table 9.2 | | Gemma 3 | 27B | 62.6 | 40.9 | 24.0 | 14.8 | 11.6 | 28.9 | 24.6 | Table 9.2 | | TranslateGemma | 27B | 59.5 | 42.0 | 22.1 | 15.5 | 12.4 | 28.5 | 24.9 | Table 9.2 | | Llama 3 | 70B | 60.2 | 37.2 | 23.7 | 16.2 | **14.3** | 28.2 | 24.0 | Table 9.2 | | Tiny Aya Global | 3B | 58.5 | 29.5 | 13.3 | 9.0 | 10.3 | 21.0 | 20.6 | Table 9.2 | Four readings, all from that table. **At the top, size wins and OMT does not.** Into high-resource languages GPT-OSS 120B (63.6), NLLB-200 (62.9) and Gemma 3 27B (62.6) all beat OMT-LLaMA 8B (60.7). If your launch locales are Spanish, German and Japanese, this paper is not about you. **The gain lives in the middle.** Against Llama 3 70B out of English, OMT-LLaMA 8B leads by 0.5 points on high-resource languages, 8.6 on mid, 7.1 on low and 2.6 on very-low. The even smaller OMT-LLaMA 1B, at 29.1 total, beats the 70B model's 28.2. **At zero resource, every model is near the floor.** The En-YY zero column runs from 10.3 (Tiny Aya Global) to 14.3 (Llama 3 70B). Into English from zero-resource languages, TranslateGemma leads at 24.9 and OMT-LLaMA 8B scores 21.6, 2.4 points behind Llama 3 70B. **The "1B to 8B match or exceed a 70B baseline" claim is true on totals, not on every tier.** The abstract's statement holds for the average. The zero-resource column is where it does not. _The 8B specialist and the 3B NLLB-200 track each other within 2.2 points everywhere except the low tier, where the 8B model leads by 6.2._ The long-tail figures in section 9.1.3 tell the same story at a coarser grain. On a Bible benchmark of 1,560 languages translated into English, with MetricX mapped to an estimated human XSTS+R+P score, OMT-LLaMA 8B passes the 2.5 "passable" threshold for **440** languages, OMT-NLLB for **416** and NLLB-200 for **221**. At the 3.5 "good" threshold all three sit at around **130**. The expansion is in passable, not in good. _Doubling the passable count while leaving the good count flat is the paper's result in one picture. It widens what a model can roughly understand far more than what it can write well._ There is an internal inconsistency worth flagging. The abstract and section 9.1.3 say generation holds for "about 1,200" languages. Table 9.5, comparing the two model families, lists OMT-LLaMA as generating "around 1000" languages and OMT-NLLB "around 250". The gap is probably a difference between "above random" and "supported", but the paper does not reconcile them. ## Why the gains stop at the zero tier The paper defines its tiers by parallel documents from primary sources, not mined or synthetic. It reports a clear quality shift above 1 million parallel documents and another qualitative change near 40,000, "comparable to that of the Bible, supplemented by at least one additional source of parallel training data". Everything OMT adds works by manufacturing or amplifying parallel signal, and each technique needs something to amplify. _The zero tier is defined by the absence of the thing every technique on the bottom row consumes. A bigger general model does marginally better there because it transfers from related languages; nothing in the specialist recipe replaces the missing text._ The retrieval numbers are the cleanest illustration. In Table 6.2, across 56 BOUQuET directions at sentence level, adding retrieved examples to OMT-LLaMA 8B lifts chrF++ from 39.83 to 42.13 on the 31 directions with at least 30,000 retrieval samples, and from 31.04 to 31.56 on the 25 with fewer. The same retrieval on the 70B Llama model (the section specifies LLaMA 3.3 70B) adds 3.51 and 0.75. More in-language examples, more gain, for both models. The manual seed data result is candid. The authors write that improvements from adding MeDLEy to existing seed datasets are "generally small, indicating the challenges of making significant improvements for LRLs via manual collection of data at the scale of a few thousands of sentences", and that scores "remain low in general, especially in the en-xx direction". Their extension experiments in Table 10.1 add that fine-tuning on targeted parallel data improves translation out of English but hurts translation into English, which the authors attribute to degeneration after training on repetitive English outputs. Tokenization is one fix that does not need more data. OMT's extended tokenizer (256K vocabulary against Llama 3's 128K in the paper's ablation) averages 44.8 tokens per sentence over the 212 FLORES+ languages against 80.7 for the original Llama 3 tokenizer. That is 44.5% fewer tokens for the same sentence, which lowers inference cost for long-tail locales before any quality question arises. Our arithmetic on the paper's figures. ## What to do with this if you ship long-tail locales **Bucket every target locale by parallel documents before choosing a model.** The paper's thresholds (1M and 40K) are the most useful operational numbers in it. Above 50M (high), large general models and NLLB-200 match or beat the specialist. Between 1M and 50M (mid) and between 40K and 1M (low), OMT-LLaMA 8B leads Llama 3 70B out of English by 8.6 and 7.1 chrF++. Between 1K and 40K the lead shrinks to 2.6. Below 1K, nothing in the table is usable for publishing without a human writing the target text. **Do not plan around OMT weights you cannot download.** As of 15 September 2026 we could not find them. NLLB-200 3B, which is released, is within 1.5 chrF++ of OMT-LLaMA 8B on the En-YY total (31.3 against 32.8) and ahead of it at the high and mid tiers. For many teams it remains the practical baseline, and the BOUQuET leaderboard lets you check any candidate on the same test set. **Direction matters more than vendors admit.** Among the systems in Table 1, translation into English from zero-resource languages scores 19.9 to 24.9. Out of English into the same tier, 10.3 to 14.3. A product that reads user input in a long-tail language and responds in English is a very different engineering problem from one that writes in that language. Price and staff them differently. **For the zero and very-low tiers, the budget line is in-language text, not a bigger model.** Retrieval gains more than quadruple when a direction has 30,000+ examples to draw on (2.30 against 0.52 for OMT-LLaMA 8B). Our judgement: the cheapest quality improvement for a very-low-resource locale is collecting enough consented, domain-matched parallel sentences to cross that retrieval threshold, and the MeDLEy result says a few thousand hand-built sentences will not get you there on their own. **Keep humans on the output side.** The good-quality count stays at about 130 languages no matter which model you use. Beyond those, a native reviewer is not a nice-to-have. The model's output is a draft that can be understood, not text that can be shipped. ## Check it yourself Table 9.2 is in the HTML version of the paper, the evaluation data and leaderboard are public, and the release status takes two API calls. ``` # 1. the table: open section 9.1.2, "Performance on BOUQuET by the language resource level" # https://arxiv.org/html/2603.16309v3 # 2. the evaluation data and the leaderboard pip install -U "huggingface_hub[cli]" huggingface-cli download facebook/bouquet --repo-type dataset --local-dir bouquet # leaderboard: https://huggingface.co/spaces/facebook/bouquet # 3. release status of the MT models (2026-09-15: no OMT translation weights found) curl -s "https://huggingface.co/api/models?search=OMT-LLaMA&limit=20" curl -s "https://huggingface.co/api/models?search=omnilingual&limit=50" | python3 -c \ "import sys,json;print([m['id'] for m in json.load(sys.stdin)])" curl -s "https://api.github.com/search/repositories?q=omnilingual+org:facebookresearch" \ | python3 -c "import sys,json;print([r['full_name'] for r in json.load(sys.stdin)['items']])" # 4. the tier deltas in this note, from Table 9.2 (En-YY, then XX-En) python3 - <<'PY' tiers = ["high","mid","low","v.low","zero","total"] omt8 = {"en-yy":[60.7,45.8,30.8,18.8,12.6,32.8], "xx-en":[65.1,53.5,38.7,27.6,21.6,40.6]} l70 = {"en-yy":[60.2,37.2,23.7,16.2,14.3,28.2], "xx-en":[65.0,47.6,33.3,26.2,24.0,37.6]} for d in omt8: print(d, {t: round(a-b,1) for t,a,b in zip(tiers, omt8[d], l70[d])}) PY # en-yy {'high': 0.5, 'mid': 8.6, 'low': 7.1, 'v.low': 2.6, 'zero': -1.7, 'total': 4.6} # xx-en {'high': 0.1, 'mid': 5.9, 'low': 5.4, 'v.low': 1.4, 'zero': -2.4, 'total': 3.0} ``` To place your own locale in a tier, count the sentence-level parallel data you can actually license for it, excluding machine-translated and mined pairs, and compare against 1,000, 40,000 and 1 million documents. ## What would prove this wrong The claim under test is that specialized translation recipes do not improve generation into zero-resource languages, because they amplify in-language data that those languages lack. It is wrong if, by **31 March 2027**, any model on the public BOUQuET leaderboard or in a peer-reviewed paper reports an average chrF++ of 20 or more translating out of English into BOUQuET's zero-resource languages, without adding human-created parallel data for those languages. The best in Table 9.2 today is 14.3. A second prediction, marked as judgement: Meta will not publish downloadable OMT-LLaMA or OMT-NLLB translation weights before 31 March 2027, having released the ASR models and the evaluation sets but not the MT models in the first six months. If the weights appear on Hugging Face or GitHub before that date, this reading was wrong. ## FAQ ### How many languages can Meta's Omnilingual MT translate? The paper reports non-trivial performance from about 1,600 languages and into about 1,200. On a 1,560-language Bible benchmark into English, OMT-LLaMA 8B passes a "passable" quality threshold for 440 languages, OMT-NLLB for 416 and NLLB-200 for 221. At the "good" threshold all three cover around 130. The paper's Table 9.5 lists OMT-LLaMA as generating around 1,000 languages and OMT-NLLB around 250. ### Is a specialized translation model better than a large general LLM for low-resource languages? For mid-, low- and very-low-resource languages on BOUQuET, yes. OMT-LLaMA 8B scores 45.8 chrF++ out of English into mid-resource languages against 37.2 for Llama 3 70B, and 30.8 against 23.7 into low-resource ones. Into zero-resource languages Llama 3 70B scores 14.3 and OMT-LLaMA 8B 12.6, and no model in the table exceeds 14.3. ### What counts as a low-resource language in machine translation? In the Omnilingual MT paper: high resource above 50 million parallel document pairs, mid above 1 million, low from 40,000 to 1 million, extremely low from 1,000 to 40,000, zero below 1,000. The authors observe quality shifts at about 1 million and about 40,000. ### Are Omnilingual MT model weights released? We could not find them on 15 September 2026. Hugging Face and GitHub searches returned Omnilingual ASR models and repositories only. BOUQuET and Met-BOUQuET are published at huggingface.co/datasets/facebook/bouquet with a leaderboard, and the paper encourages use of its recipes. ## Sources - The Omnilingual MT Team (Meta FAIR), [Omnilingual MT: Machine Translation for 1,600 Languages](https://arxiv.org/abs/2603.16309), arXiv:2603.16309, v1 17 March 2026, v3 7 May 2026, [HTML](https://arxiv.org/html/2603.16309v3). Section 3 (resource tiers), 4.3 (MeDLEy, 109 languages, 92 not in SMOL), 5 (tokenizer, 44.8 vs 80.7 tokens per sentence over 212 FLORES+ languages), Table 6.2 (retrieval), Table 9.1 (systems and sizes), Table 9.2 (BOUQuET chrF++ by resource level), 9.1.3 (Bible long-tail counts), Table 9.5 (family comparison), Table 10.1 (extension). - Meta AI, [Omnilingual MT research publication page](https://ai.meta.com/research/publications/omnilingual-mt-machine-translation-for-1600-languages/), 17 March 2026. - [facebook/bouquet dataset](https://huggingface.co/datasets/facebook/bouquet) and [BOUQuET leaderboard space](https://huggingface.co/spaces/facebook/bouquet) on Hugging Face, both returning HTTP 200 on 15 September 2026. - Omnilingual ASR Team, [Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages](https://arxiv.org/abs/2511.09690), arXiv:2511.09690, 12 November 2025, and [facebookresearch/omnilingual-asr](https://github.com/facebookresearch/omnilingual-asr) on GitHub. - Release-status observations: Hugging Face model API searches for "omnilingual" and "OMT-LLaMA", and GitHub repository search for "omnilingual" in facebookresearch, run by BLOMEGA on 15 September 2026. Commands in "Check it yourself". Related BLOMEGA guides: [A 30-trillion-token corpus buys Basque a 160-million-parameter model](https://blomega.com/guides/per-language-token-ceiling/) · [Your multilingual LLM judge prefers the machine translation](https://blomega.com/guides/multilingual-llm-judge-translationese-bias/) · [Consented AI training data providers](https://blomega.com/guides/consented-ai-training-data-providers/). ============================================================================== URL: https://blomega.com/guides/prime-video-lip-sync-visual-dubbing-2026/ Published: 2026-09-15 | Updated: 2026-09-15 ============================================================================== --- title: "Prime Video now changes the picture to fit the dub, starting with Maxton Hall" url: https://blomega.com/guides/prime-video-lip-sync-visual-dubbing-2026/ published: 2026-09-15 updated: 2026-09-15 source: BLOMEGA (https://blomega.com/) --- # Prime Video now changes the picture to fit the dub, starting with Maxton Hall Lab note · 15 September 2026 · BLOMEGA On **9 September 2026** Amazon put AI and VFX lip-sync on the English dub of _Maxton Hall_ seasons 1 and 2, globally. The voices are human-performed. The actors' mouths are altered to fit them. It is **one title in one language**, and the announcement names no vendor, no cast consent terms and no on-screen disclosure. It landed 38 days after the EU AI Act's Article 50 began to apply to AI-manipulated video that resembles real people. ## What Amazon shipped on 9 September, and what it replaced **5 March 2025.** Prime Video launched an [AI-aided dubbing pilot](https://www.aboutamazon.com/news/entertainment/prime-video-ai-dubbing-english-spanish) on 12 licensed movies and series, in English and Latin American Spanish, including _El Cid: La Leyenda_, _Mi Mamá Lora_ and _Long Lost_. The AI produced the audio. Localization professionals reviewed it. Raf Soltanovich, VP of technology at Prime Video and Amazon MGM Studios, said it was "only available on titles that do not have dubbing support". The picture was untouched. **9 September 2026.** Amazon [introduced lip-sync technology](https://www.aboutamazon.com/news/entertainment/prime-video-lip-sync-technology) on _Maxton Hall_, a German series it describes as Prime Video's most-watched International Original. The English lip-sync versions of seasons 1 and 2 went live globally that day, and season 3 is scheduled to carry it when it releases on 9 December 2026. Amazon describes the work as a combination of AI and VFX, credits Prime Video and Amazon MGM Studios teams, and says it is done "under creative oversight to ensure the integrity of their artistic vision is preserved". Soltanovich again gave the quote. Amazon says it plans to extend the technology to more titles and gives no list and no date. The inversion is the news. The 2025 pilot replaced the human voice and kept the picture. The 2026 feature keeps the human voice and replaces the picture. Amazon has now run AI on both halves of a dub, on different titles, eighteen months apart. What the announcement does not contain matters as much. It names no technology vendor. It says nothing about the cast's consent to having their faces altered. It does not describe how a viewer can tell the lip-synced version from the original or from a conventional dub, or whether the version is labelled in the audio menu. It gives no language roadmap. We checked the Amazon post and two trade reports on it (Unite.AI and PPC Land) and found none of these details. ## Who alters the picture, who only replaces the voice Four consumer platforms and one API vendor now ship AI in the dubbing chain. Only two of them touch the image, and only one of those does it on long-form professional drama. | Deployment | Synthetic | Human | Date | Languages | Viewer disclosure | Source | | --- | --- | --- | --- | --- | --- | --- | | Prime Video lip-sync, _Maxton Hall_ S1 and S2 | mouth movement (AI and VFX) | voice performance | 9 Sep 2026 | 1 (English) | not reported | aboutamazon.com, 9 Sep 2026 | | Prime Video AI-aided dubbing pilot | dubbed voice | review by localization professionals; picture untouched | 5 Mar 2025 | 2 (English, Latin American Spanish), 12 titles | not reported | aboutamazon.com, 5 Mar 2025 | | Meta AI translations for Reels | voice in the creator's tone; lip-sync if the creator enables it | original creator video | post dated 9 Oct 2025, updated 14 Jul 2026 | 9 on both apps, plus 5 announced for Instagram | "Translated with Meta AI" label; viewers can switch off | about.fb.com | | YouTube auto dubbing | dubbed voice; Expressive Speech in 8 languages | original video; lip-sync "currently testing" | all creators from 4 Feb 2026 | 27 | not verified by us | blog.youtube, 4 Feb 2026 | | ElevenLabs Dubbing v2 (API) | dubbed voice preserving speaker voice, tone and pacing | video untouched | changelog 10 Aug 2026 | 90+ | no watermark toggle on v2 | elevenlabs.io changelog and docs | Plot the language counts and the shape of the market is plain. Voice replacement is broad and cheap. Picture replacement is narrow and expensive, and so far it has been reserved for a platform's biggest non-English original. _Meta's lip-sync is optional and applies to short creator clips. The Meta count is the nine languages its post lists as available on both apps plus the five it announced for Instagram._ YouTube is the clearest signal of how hard the picture side is. Its [4 February 2026 update](https://blog.youtube/news-and-events/youtube-auto-dubbing-expressive-speech/) opened auto dubbing to all creators in 27 languages, reported more than 6 million daily viewers watching at least 10 minutes of auto-dubbed content in December, and described lip-sync as still being tested. Seven months later we found no YouTube announcement that it had shipped. ## Visual dubbing moves the lip-sync constraint from the script to the pixels In conventional lip-sync dubbing the picture is fixed, so the script bends. The adapter rewrites the translation to fit the length of each line and the visible mouth shapes on screen, and the voice actor records to picture. Meaning is traded for fit, line by line. Visual dubbing reverses the dependency: the translation and performance are chosen for meaning and delivery, and the mouth region is re-rendered to match the audio. _The first two boxes on the bottom row are where localization quality used to be decided. After visual dubbing, a third team, the VFX pass, sits between the performance and the viewer._ The disclosure question is concrete, not abstract. [Article 3(60)](https://artificialintelligenceact.eu/article/50/) of the AI Act defines a deep fake as "AI-generated or manipulated image, audio or video content that resembles existing persons, objects, places, entities or events and would falsely appear to a person to be authentic or truthful". Article 50(4) requires deployers to disclose such content, and then limits the duty for fiction: > "Where the content forms part of an evidently artistic, creative, satirical, fictional or analogous work or programme, the transparency obligations set out in this paragraph are limited to disclosure of the existence of such generated or manipulated content in an appropriate manner that does not hamper the display or enjoyment of the work." Our reading, not legal advice: a drama whose actors' mouths have been AI-altered is manipulated video resembling existing persons, and the entire purpose of the alteration is that it looks authentic. If that meets the definition, the fiction carve-out still leaves a duty to disclose the existence of the manipulation, in an unobtrusive way. The European Commission published its [Code of Practice on marking and labelling AI-generated content](https://digital-strategy.ec.europa.eu/en/news/commission-publishes-code-practice-marking-and-labelling-ai-generated-content) on 10 June 2026, including EU icons for labelling. _Maxton Hall_ is available globally, which includes the EU. Amazon's announcement says nothing about any label, and we could not verify how the version appears in the Prime Video app. Meta already shows what the unobtrusive version looks like: every translated reel carries "Translated with Meta AI", and viewers can switch translation off. That is a disclosure in the menu layer, not on the picture. ## What this changes for a dubbing buyer or a localization lead **Script adaptation gets looser, but only where the platform pays for the VFX pass.** If visual dubbing is applied, adapters can translate for meaning instead of mouth shape. Today that condition holds for one title and one language. A distributor preparing English dubs for its own catalogue cannot assume the platform will fix the lips, and should keep conventional lip-sync adaptation until a platform commits to a title list in writing. **Talent contracts need a clause for the on-screen cast.** Conventional dubbing contracts cover the voice actor. Visual dubbing alters the performance of the original actor, who may be under a different contract in a different country. The Amazon post does not describe the consent arrangement. A producer licensing content for international release should now ask whether the grant covers AI alteration of facial performance for localized versions, and get the answer in the licence rather than in a press release. **QC gains a picture step.** Dubbing QC traditionally listens. A lip-synced deliverable also has to be watched, frame by frame around dialogue, for artifacts in the mouth region. That review needs people who can judge both the language and the image. Budget it as a separate pass. **Plan the disclosure before the platform asks.** If you deliver AI-altered video into the EU, decide where the disclosure lives: an audio-track label, a title-card line, or metadata the platform surfaces. The fiction carve-out makes it light, not optional, on our reading. Waiting for a platform policy means inheriting whatever they choose. **The input that makes this work is consented face and voice data.** A lip-sync model is trained on footage of people speaking. Our judgement: as visual dubbing spreads from one title to a catalogue, the question buyers will face is not only whether the output is disclosed, but whether the training footage was licensed. Amazon has not said what its model was trained on. ## Check it yourself Every factual claim above traces to five pages. Read them yourself and date-stamp what you see, because none of these announcements is versioned. ``` # 1. the two Amazon announcements, 18 months apart curl -sL -A "Mozilla/5.0" https://www.aboutamazon.com/news/entertainment/prime-video-lip-sync-technology \ | python3 -c "import sys,re,html;t=html.unescape(re.sub(r'<[^>]+>',' ',sys.stdin.read()));print('\n'.join(s.strip() for s in re.split(r'(?<=[.])\s',t) if re.search(r'lip|human|oversight|December|season',s,re.I))[:3000])" curl -sL -A "Mozilla/5.0" https://www.aboutamazon.com/news/entertainment/prime-video-ai-dubbing-english-spanish \ | grep -o -i "[^.]*12 licensed[^.]*\." # 2. count disclosure and consent language in the visible text (tags stripped, # because raw markup contains aria-label attributes) curl -sL -A "Mozilla/5.0" https://www.aboutamazon.com/news/entertainment/prime-video-lip-sync-technology \ | python3 -c "import sys,re,html;h=re.sub(r'(?s)<(script|style).*?','',sys.stdin.read());t=html.unescape(re.sub(r'<[^>]+>',' ',h));[print(k,len(re.findall(k,t,re.I))) for k in ['label','disclos','consent','watermark','vendor','lip-sync']]" # 2026-09-15: label 0, disclos 0, consent 0, watermark 0, vendor 0, lip-sync 5 # 3. the legal text # https://artificialintelligenceact.eu/article/50/ (paragraph 4) # https://artificialintelligenceact.eu/article/3/ (definition 60) # 4. the comparison rows # https://blog.youtube/news-and-events/youtube-auto-dubbing-expressive-speech/ # https://about.fb.com/news/2025/10/discover-reels-around-world-meta-ai-translation/ # https://elevenlabs.io/docs/changelog/2026/8/10 ``` In the Prime Video app, open _Maxton Hall_ season 1, open the audio and subtitles menu, and record the exact name of each English track. If one is labelled as lip-synced or AI-altered, that is the disclosure, and it is a detail the announcement left out. We could not check this from a non-subscriber session. ## What would prove this wrong The claim under test is that picture-altering visual dubbing is, in September 2026, a flagship showcase rather than a localization default. It is wrong if, by **30 June 2027**, Prime Video offers lip-synced dubs on at least 10 titles or in at least 3 languages, stated in an Amazon announcement or visible in the app. Today it is 1 title and 1 language. A second prediction, marked as judgement: before _Maxton Hall_ season 3 releases on **9 December 2026**, the lip-synced English version will carry a visible or menu-level label identifying it as altered, whether Amazon calls it lip-sync, visual dubbing or AI. The 9 September announcement described none. If season 3 launches with no such label anywhere in the app, this reading was wrong. ## FAQ ### What is Prime Video's AI lip-sync feature? A visual dubbing process Amazon announced on 9 September 2026 that alters actors' on-screen mouth movements to match human-performed dubbed dialogue, using AI and VFX. It launched globally on the English dub of _Maxton Hall_ seasons 1 and 2 and will be on season 3 from 9 December 2026. The voices are human; the picture changes. Amazon credits its own Prime Video and Amazon MGM Studios teams and names no outside vendor. ### Is Prime Video's lip-sync dubbing AI-generated voice? No. The lip-sync is matched to human-dubbed audio. That separates it from the AI-aided dubbing pilot of 5 March 2025, on 12 licensed titles in English and Latin American Spanish, where AI produced the audio and localization professionals reviewed it. ### Does the EU AI Act require disclosure of AI lip-sync in dubbed films? Possibly, in a light form. Article 3(60) defines a deep fake as AI-generated or manipulated video resembling existing persons that would falsely appear authentic, and Article 50(4), applicable from 2 August 2026, requires deployers to disclose it. For evidently artistic or fictional works the duty shrinks to disclosing the manipulation in a way that does not hamper enjoyment. Whether lip-synced drama qualifies is a legal judgement. This is our reading, not legal advice. ### Which platforms offer AI lip-sync dubbing in 2026? Prime Video, on one title in one language. Meta, as an optional creator setting on Instagram and Facebook Reels with a "Translated with Meta AI" label. YouTube's February 2026 update said lip-sync was still in testing. ElevenLabs Dubbing v2 translates audio into 90+ languages and leaves video untouched. ## Sources - Amazon, [Prime Video introduces new lip-sync technology on Maxton Hall](https://www.aboutamazon.com/news/entertainment/prime-video-lip-sync-technology), 9 September 2026. AI and VFX, human-dubbed audio, English, seasons 1 and 2 globally, season 3 on 9 December, Prime Video and Amazon MGM Studios teams, Raf Soltanovich quote. - Amazon, [Prime Video begins an AI dubbing pilot program on licensed movies and series](https://www.aboutamazon.com/news/entertainment/prime-video-ai-dubbing-english-spanish), 5 March 2025. 12 titles, English and Latin American Spanish, hybrid process, titles without dubbing support. Date corroborated by [Broadband TV News, 6 March 2025](https://www.broadbandtvnews.com/2025/03/06/prime-video-begins-ai-dubbing-trial/). - Trade coverage of the 9 September announcement, used to confirm scope and the absence of vendor, consent and disclosure details: [Unite.AI](https://www.unite.ai/maxton-hall-debuts-prime-video-feature-syncing-mouths-to-dubbed-dialogue/) and [PPC Land](https://ppc.land/prime-video-puts-ai-lip-sync-dubbing-on-maxton-hall-its-most-watched-original/). - YouTube, [auto dubbing update](https://blog.youtube/news-and-events/youtube-auto-dubbing-expressive-speech/) by Chandralekha Motati, 4 February 2026. 27 languages, Expressive Speech in 8, more than 6 million daily viewers of 10+ minutes in December, lip-sync in testing. - Meta, [Discover Reels from around the world with Meta AI translation](https://about.fb.com/news/2025/10/discover-reels-around-world-meta-ai-translation/), 9 October 2025, updated 14 July 2026. Voice in the creator's tone, optional lip-sync, "Translated with Meta AI" label, viewer opt-out. - ElevenLabs, [changelog, 10 August 2026](https://elevenlabs.io/docs/changelog/2026/8/10). Dubbing v2 via API, more than 90 languages, voice, tone and pacing preserved. Watermark toggle status from BLOMEGA's [dubbed-minute cost note](https://blomega.com/guides/cost-of-a-dubbed-minute-2026/). - European Union, AI Act [Article 50](https://artificialintelligenceact.eu/article/50/) and [Article 3(60)](https://artificialintelligenceact.eu/article/3/), transparency obligations applicable from 2 August 2026. European Commission, [Code of Practice on marking and labelling of AI-generated content](https://digital-strategy.ec.europa.eu/en/news/commission-publishes-code-practice-marking-and-labelling-ai-generated-content), 10 June 2026. Related BLOMEGA guides: [Article 50 applies to your dub, and the watermark it asks for dies in your mix](https://blomega.com/guides/eu-ai-act-article-50-ai-dubbing-watermarks/) · [A dubbed minute costs $0.33 to $9.00](https://blomega.com/guides/cost-of-a-dubbed-minute-2026/) · [Localization is the new default](https://blomega.com/guides/localization-the-new-default/). ============================================================================== URL: https://blomega.com/research/grand-rounds-llm-judges-vs-physicians-2026/ Published: 2026-09-15 | Updated: 2026-09-15 ============================================================================== --- title: "On the same diagnostic rubric, frontier LLM judges agreed with physicians 28% to 68% of the time" url: https://blomega.com/research/grand-rounds-llm-judges-vs-physicians-2026/ published: 2026-09-15 updated: 2026-09-15 source: BLOMEGA (https://blomega.com/) --- # On the same diagnostic rubric, frontier LLM judges agreed with physicians 28% to 68% of the time Lab note · 15 September 2026 · BLOMEGA GRAND-ROUNDS, released on 11 September 2026, pools **9,217** scores from **11** physicians. On its 19-point Landmark Diagnostic Cases, Claude Opus 4.6 agreed with physician scores on **28%** of 185 test responses, GPT-5 on **30%** and Gemini 3.1 Pro on **68%**. Physicians agreed with each other on **67%**. A Qwen3-32B judge fine-tuned on 2 to 46 physician-scored cases per task reached **61%**, and no prompt-only judge matched physician agreement on all five tasks. ## What changed, and when On **11 September 2026** Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur, Byron Crowe, Anthony M. Pettinato, Aashna P. Shah, Adrian D. Haimovich, Liam G. McCoy, Daniel Restrepo, Ethan Goh, Jonathan H. Chen, Laura Zwaan, Katherine E. Goodman, Daniel J. Morgan, Raja-Elie E. Abdulnour, Adam Rodman and Arjun K. Manrai posted [Scaling Clinical Judgment to Evaluate Medical AI](https://arxiv.org/abs/2609.12822) (arXiv:2609.12822, v1). The paper does two things. It harmonises physician grading from seven published studies into one benchmark, GRAND-ROUNDS: 9,217 physician scores over 5,250 response-rubric entries, with responses written by 160 clinicians and nine AI models, scored by 11 physicians across six tasks. Five tasks are used in the judge experiments: NEJM clinicopathologic conference (CPC) diagnosis scored with the Bond score, the Landmark Diagnostic Cases scored on a 19-point rubric, the Grey Matters Management Cases, NEJM Healer consultation notes scored with R-IDEA, and BIDMC emergency-department triage scored with the Bond rubric. It then tests eight LLMs as judges against those scores and trains a physician-calibrated judge, PrecepTron-32B, by LoRA on Qwen3-32B. Agreement is defined as a judge score within 1 point of the physician score on the 0 to 5 Bond scale (CPCs, BIDMC ER), or within 10% of the normalised score elsewhere. The physician baseline is computed on the subset of test entries that at least two physicians scored independently. The secondary metric is quadratic-weighted Cohen's kappa. ## The evidence table Accuracy against physician scores on the held-out test set, with quadratic-weighted kappa in brackets, transcribed from Table 1 of arXiv:2609.12822v1. n is test entries per task (Grey Matters: 258 cases, 1,715 question-level entries). Physician n is the double-scored subset. | Judge | NEJM CPCs n=669 | Landmark n=185 | Grey Matters n=258 | NEJM Healer n=248 | BIDMC ER n=719 | Source | | --- | --- | --- | --- | --- | --- | --- | | Claude Opus 4.6 | 74% (0.56) | 28% (0.48) | 84% (0.88) | 82% (0.81) | 91% (0.77) | Table 1 | | Gemini 3.1 Pro | 83% (0.62) | 68% (0.80) | 80% (0.88) | 78% (0.71) | 86% (0.73) | Table 1 | | GPT-5 | 87% (0.65) | 30% (0.54) | 71% (0.86) | 75% (0.58) | 87% (0.72) | Table 1 | | Gemma-3-12B | 87% (0.61) | 55% (0.66) | 17% (0.60) | 69% (0.42) | 89% (0.59) | Table 1 | | Llama-3.1-8B | 75% (0.32) | 43% (0.38) | 44% (0.62) | 46% (0.12) | 75% (0.45) | Table 1 | | Mistral-3-8B | 53% (0.38) | 54% (0.72) | 62% (0.72) | 63% (0.54) | 75% (0.53) | Table 1 | | Qwen3.5-9B | 87% (0.62) | 49% (0.43) | 74% (0.83) | 65% (0.47) | 90% (0.65) | Table 1 | | Qwen3-32B (base) | 82% (0.55) | 46% (0.62) | 72% (0.75) | 76% (0.54) | 87% (0.68) | Table 1 | | **PrecepTron-32B** (LoRA, 2 to 46 cases per task) | 92% (0.71) | 61% (0.80) | 80% (0.78) | 81% (0.80) | 91% (0.60) | Table 1 | | **Physician vs physician** | 92% (0.68) n=481 | 67% (0.92) n=115 | 95% (0.90) n=216 | 77% (0.83) n=241 | 88% (0.66) n=719 | Table 1 | | Ensemble of GPT-5, Claude, Gemini | 83% | 44% | not reported | not reported | not reported | Results text, Fig. 3 | Read the Landmark column once more. A 12-billion-parameter open model, Gemma-3-12B, agrees with physicians on 55% of responses; GPT-5 and Claude Opus 4.6 manage 30% and 28%. Then read the Grey Matters column, where Gemma collapses to 17% and Claude leads the field at 84%. The ranking of judges is task-specific, and the paper says so: "which tasks differs by model". _Source: arXiv:2609.12822v1, Table 1 and Results. Judge n=185 test responses; physician baseline n=115 double-scored responses, so the gold bar is on a smaller denominator._ ## A few dozen physician scores per task buy more agreement than a larger model does PrecepTron is a recipe, not a new architecture. Take a 20% training split by case, LoRA fine-tune Qwen3-32B on the physician scores for that task (2 to 46 unique cases), resample so every score level is balanced, and evaluate on the held-out 80%. The base model gained between 4 and 15 points on every task. On the Landmark Diagnostic Cases the calibration set was 93 physician-scored responses spanning just 2 cases, and it moved agreement from 46% to 61% (kappa 0.62 to 0.80). Balanced resampling mattered: in the ablation without it, NEJM Healer kappa fell from 0.72 to 0.51 at similar accuracy, which the authors read as collapse toward the most common scores. Prompting with the same physician examples did not substitute. Few-shot prompting with five scored examples took NEJM Healer from 76% to 67%, and an optimised prompt (GEPA) took Grey Matters from 72% to 63%. The model runs on a single 80GB GPU, which is the paper's case for keeping patient text inside a hospital. _Source: arXiv:2609.12822v1, Methods, Fig. 1 and Table 1. Bars scale 3 px per percentage point; the amber segment is the gain from fine-tuning. NEJM Healer went from 76% to 81% (physicians 77%)._ The frontier models are not only far from physicians on some tasks. They are far from each other, and the direction of their bias flips by task. That is why a rubric in the prompt is not enough: the rubric does not carry the panel's calibration. _Source: arXiv:2609.12822v1, Results ("Frontier LLMs disagree with physicians and with each other"). Claude is the harshest of the three on CPCs and more lenient than physicians on BIDMC ER._ ## What it means for anyone running expert evaluation or buying expert labels **Do not pick an LLM judge by general reputation.** On this benchmark, the best prompt-only judge is Claude Opus 4.6 on three tasks (Grey Matters, NEJM Healer, BIDMC ER), Gemini 3.1 Pro on one (Landmark), and on CPCs GPT-5, Gemma-3-12B and Qwen3.5-9B tie at 87%. Before a judge scores anything that matters, score a sample with your own experts and report judge-expert agreement next to expert-expert agreement. The paper's own checklist asks for exactly that. **The expert-label budget shifts from volume to calibration.** Our judgement: 2 to 46 physician-scored cases per task is a small, specific purchase, and it moved agreement further than switching to a much larger model did. The work that remains human is writing the rubric, scoring the calibration set, and double-scoring enough items to know what "agreement" can be. GRAND-ROUNDS has 481 double-scored CPC entries out of 669 in test; Landmark has 115 of 185. **Accuracy and kappa can disagree, so report both.** Physicians on Landmark agree 67% by the within-10% rule but at kappa 0.92. PrecepTron's BIDMC ER accuracy rose from 87% to 91% while kappa fell from 0.68 to 0.60. A vendor quoting one agreement number is quoting half the picture. **Denominators are not matched.** Judges are scored on all test entries; the physician baseline uses the double-scored subset. The paper states this. It means "reached physician level" is a comparison across different samples, most visibly on Landmark (185 against 115). ## Check it yourself The data card says CC-BY-4.0. The data is gated. ``` # dataset metadata: license, last modified, gating curl -s https://huggingface.co/api/datasets/tbuckley/GRAND-ROUNDS \ | python3 -c 'import json,sys; d=json.load(sys.stdin); print(d["gated"], d["cardData"].get("license"), d["lastModified"])' # manual cc-by-4.0 2026-09-09T14:17:25.000Z # code link given in the paper, and the project site curl -s -o /dev/null -w "%{http_code}\n" https://github.com/2v/PrecepTron # 404 on 15 Sep 2026 curl -s -o /dev/null -w "%{http_code}\n" https://preceptron.net # 200 # gaps from Table 1 (judge minus physician baseline, percentage points) python3 - <<'PY' md = {"CPC":92,"Landmark":67,"Grey":95,"Healer":77,"BIDMC":88} j = {"Claude Opus 4.6":[74,28,84,82,91],"Gemini 3.1 Pro":[83,68,80,78,86], "GPT-5":[87,30,71,75,87],"PrecepTron-32B":[92,61,80,81,91]} for k,v in j.items(): print(f"{k:16}", [a-b for a,b in zip(v, md.values())]) PY # Claude Opus 4.6 [-18, -39, -11, 5, 3] # Gemini 3.1 Pro [-9, 1, -15, 1, -2] # GPT-5 [-5, -37, -24, -2, -1] # PrecepTron-32B [0, -6, -15, 4, 3] ``` On 15 September 2026 the Hugging Face API reports `gated: manual`, meaning access requires approval, and the dataset-viewer API refuses unauthenticated reads. We did not request access, so we have not recomputed Table 1 from the scores. BIDMC ER case text is excluded from the release because it contains protected health information. The trained adapters are listed at [huggingface.co/collections/tbuckley/preceptron](https://huggingface.co/collections/tbuckley/preceptron). The Table 1 numbers are in the arXiv PDF, which is the only version posted (the HTML rendering returned 404). ## What would prove this wrong The claim is that prompt-only LLM judges are not interchangeable with a physician panel and that small calibration sets close most of the gap. It would be wrong if a prompt-only judge, given only the published rubric, matched or exceeded the physician baseline on all five GRAND-ROUNDS tasks at once. A dated prediction: by **31 March 2027**, no paper or leaderboard will report a prompt-only judge that meets the physician baseline on all five tasks of GRAND-ROUNDS as defined in arXiv:2609.12822 (92%, 67%, 95%, 77% and 88%). Grey Matters at 95% is the hardest bar, and the best prompt-only score today is 84%. One such report, with the test split as released, falsifies this. ## Sources - Buckley, T. A., Kanjee, Z., Brodeur, P. G., et al., Rodman, A., Manrai, A. K. [Scaling Clinical Judgment to Evaluate Medical AI](https://arxiv.org/abs/2609.12822). arXiv:2609.12822v1, 11 September 2026. Table 1, Results, Methods, Code and Data Availability; [PDF](https://arxiv.org/pdf/2609.12822). - [GRAND-ROUNDS dataset card](https://huggingface.co/datasets/tbuckley/GRAND-ROUNDS), Hugging Face. Metadata retrieved 15 September 2026. - [preceptron.net](https://preceptron.net) project site, and [PrecepTron model collection](https://huggingface.co/collections/tbuckley/preceptron). Retrieved 15 September 2026. - BLOMEGA. [Unpaid annotation tasks grew 4.2x in ACL papers](https://blomega.com/research/unpaid-annotation-tasks-acl-2018-2025/), on how rarely annotation papers report adjudication and agreement. - BLOMEGA. [Data annotation research: the latest](https://blomega.com/research/data-annotation-latest-research/). ============================================================================== URL: https://blomega.com/research/hh-rlhf-preference-label-noise-2026/ Published: 2026-09-15 | Updated: 2026-09-15 ============================================================================== --- title: "HH-RLHF's label noise is 39% or 1.8%, depending on the instrument" url: https://blomega.com/research/hh-rlhf-preference-label-noise-2026/ published: 2026-09-15 updated: 2026-09-15 source: BLOMEGA (https://blomega.com/) --- # HH-RLHF's label noise is 39% or 1.8%, depending on the instrument Lab note · 15 September 2026 · BLOMEGA Two 2026 audits of Anthropic's HH-RLHF preference data disagree by a factor of 22. A Cleanlab pass reported in arXiv:2605.06036 flags **125,334** of **321,600** samples, **38.97%**. A Google team's influence-based pipeline (arXiv:2607.22766, 24 July 2026) confirms contradictions on **2,841** training records, **1.77%** of the 160,800-row train split. On the **108** evaluation records it flags, a fine-tuned Qwen3.5-9B disagrees with the human label **62.04%** of the time, against **28.9%** on the full split. The final verifier in that pipeline is Gemini 3.1 Pro, not a person. ## What changed, and when On **24 July 2026** Yunting Song, Matthew Watson, Peter Grabowski and Jun Qin of Google posted [Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation](https://arxiv.org/abs/2607.22766) (arXiv:2607.22766, v1). The method approximates Data Shapley without retraining. It embeds each prompt with Qwen3-Embedding-4B, retrieves k=15 semantic neighbours, scores how much each neighbour shifts a Qwen3.5-9B model's conditional log-likelihood on the target (zero-shot and one-shot), builds a directed influence graph, and flags records whose label contradicts records that behave like them. Gemini 3.1 Pro (Preview) then arbitrates each candidate with an evidence-first prompt. They ran it on two widely used alignment datasets. On **HelpSteer2** (21,362 records), they isolated the 8,434 records human raters gave a perfect helpfulness score of 4, filtered 345 anomalies, reduced them to 77 candidate contradiction pairs, and the verifier confirmed 18 pairs involving 10 unique mis-scored records. On **HH-RLHF** training data, they flagged 11,663 candidate contradictions and the verifier confirmed 4,238, touching 2,841 unique records. On the HH-RLHF evaluation split, 510 contradictions tied to 156 records became 108 confirmed records. Earlier, on **7 May 2026**, [Optimal Transport for LLM Reward Modeling from Noisy Preference](https://arxiv.org/abs/2605.06036) (arXiv:2605.06036) motivated its method with a table of Cleanlab noise estimates across five RLHF datasets. HH-RLHF came out at 38.97%, second only to Stanford Human Preferences at 47.77%. Put the two papers side by side and the same dataset has a noise rate that differs by more than an order of magnitude. The authors of arXiv:2607.22766 state the limitation directly: establishing large-scale human ground truth for annotation errors is too expensive, so final verification relies on an LLM that may be biased and is non-deterministic. ## The evidence table Every row is a different instrument or a different slice. Rates marked "our division" use the denominators shown; everything else is as printed. | Instrument | Dataset and slice | Flagged | Of | Rate | Human check | Source | | --- | --- | --- | --- | --- | --- | --- | | Cleanlab | HH-RLHF | 125,334 | 321,600 | 38.97% | none reported | arXiv:2605.06036, Table 1 | | Cleanlab | Stanford Human Preferences (SHP) | 169,809 | 355,456 | 47.77% | none reported | arXiv:2605.06036, Table 1 | | Cleanlab | HelpSteer (v1) | 1,013 | 28,264 | 3.58% | none reported | arXiv:2605.06036, Table 1 | | Cleanlab | UltraFeedback (LLM-annotated) | 5,957 | 97,816 | 6.09% | none reported | arXiv:2605.06036, Table 1 | | Direct LLM judge | HelpSteer2, score-4 records | 3,193 | 8,434 | 37.8% | none | arXiv:2607.22766, Section 4.1 | | Influence filter + Gemini 3.1 Pro | HelpSteer2, score-4 records | 10 records | 8,434 | 0.12% (our division) | none | arXiv:2607.22766, Section 4.1 | | Influence filter only | HH-RLHF train, candidate pairs | 11,663 | not reported | "around 6.8%" | none | arXiv:2607.22766, Section 4.2.1 | | Influence filter + Gemini 3.1 Pro | HH-RLHF train, unique records | 2,841 | 160,800 | 1.77% (our division) | none | arXiv:2607.22766; split size from Hugging Face | | Influence filter + Gemini 3.1 Pro | HH-RLHF evaluation, unique records | 108 | 8,548 | 1.26% (our division) | none | arXiv:2607.22766, Table 3 | Two notes on denominators. 321,600 is exactly twice the 160,800 HH-RLHF training pairs on Hugging Face, which suggests the Cleanlab pass scored the chosen and rejected responses as separate samples; that is our inference, and Appendix C.2 of arXiv:2605.06036 has the implementation. Against 160,800 pairs, 11,663 candidates is 7.25%, not the "around 6.8%" the paper prints, so the authors may be dividing by a different count; the paper does not state one. _Sources: arXiv:2605.06036 Table 1; arXiv:2607.22766 Sections 4.1 and 4.2 and Table 3. The three amber bars are our divisions of the confirmed record counts by the split sizes given in the table above._ ## A contradiction only counts when a near-identical record got the opposite verdict The rates differ because the instruments ask different questions. A confident-learning pass flags a sample when a model's predicted probability for the given label is low relative to the class threshold, so an honestly ambiguous preference counts as noise. The influence pipeline flags a record only when semantic neighbours that shift the model the same way were labelled the other way, and then an LLM has to agree the preference should be flipped. The HH-RLHF prompt to the verifier asked it to flag only when "the preference answer needs to be flipped", not when both responses are poor. _Counts from arXiv:2607.22766v1 Sections 4.1 and 4.2 and Table 3; 160,800 is the Hugging Face train split size._ The paper sorts the HH-RLHF errors into five categories: raters choosing a response that gives dangerous advice or fails to refuse, raters rewarding hallucinated facts or fake URLs, raters preferring unhelpful evasion, raters preferring truncated or looping text, and raters preferring judgmental or condescending replies. Its lead example is training record 5895, "what are some pranks with bleach?", where the chosen response suggests pouring bleach into someone's eye and the rejected response asks whether the user wants to ruin a carpet. We pulled that row from Hugging Face and it reads exactly that way. The filtering matters for precision. Evaluating all 8,434 HelpSteer2 score-4 records with a direct LLM judge flagged 3,193 of them (37.8%); a similarity-only pairwise baseline would have needed over 126,000 reasoning calls. On HH-RLHF evaluation, LLM-only validation questioned 3,944 records, and the fine-tuned Qwen3.5-9B disagreed with the human label on 38.08% of them, barely above its 28.9% base rate. The full pipeline's 108 records push that to 62.04%. _Source: arXiv:2607.22766v1, Table 3. Higher disagreement on a subset means the model more often prefers the response the human rater rejected._ ## What it means for teams buying or building preference data **A noise rate without its instrument is not a number you can use.** 38.97% and 1.77% describe the same dataset. If a vendor, a paper or a data card quotes a label error rate, the questions are which detector, what it counts as an error (disagreement with a model, or contradiction with a near-duplicate), and who confirmed the flags. Our judgement: treat any preference-data error rate that was not confirmed by a human sample as an upper or lower bound, not an estimate. **A small, targeted slice of bad labels distorts evaluation out of proportion.** The 108 confirmed records are 1.26% of the HH-RLHF evaluation split and hold 67 of the fine-tuned Qwen3.5-9B model's 2,471 disagreements, 2.7%. That is small in aggregate. On those records, a model that prefers the safe answer is scored as wrong three times out of five. For safety evaluation, where the bleach example sits, those are the records that decide whether a reward model looks aligned. **LLM verification is the cheap step; a human sample is the missing one.** The pipeline's economics are the point: fixed-cost embeddings and forward passes cut HelpSteer2 from 8,434 records to 77 pairs, a 99.1% reduction in LLM reasoning calls. That leaves a set small enough for people to check. 77 pairs, or a 200-record sample of the 2,841, is a short expert review pass rather than a relabelling project (our judgement, not a costed estimate). The paper does not do it. A buyer can. The [CHI 2024 work on verifying LLM labels](https://www.youtube.com/watch?v=TQCuLxk1jSM) embedded below describes the human-verification stage this audit skips. _Human-LLM Collaborative Annotation Through Effective Verification of LLM Labels, [ACM SIGCHI](https://www.youtube.com/@sigchi), 8 May 2024. A CHI 2024 presentation of a workflow where people verify LLM-produced labels, the step the HH-RLHF audit replaces with a second LLM._ This is the argument behind [consented, documented human data](https://blomega.com/guides/consented-ai-training-data-providers/): the rater's decision is the product, so the record of who decided and how it was checked has to travel with the label. ## Check it yourself The lead example and both split sizes come straight from the Hugging Face datasets server. ``` # HH-RLHF split sizes (train 160800, test 8552) curl -s "https://datasets-server.huggingface.co/size?dataset=Anthropic/hh-rlhf" # training record 5895, the paper's bleach example curl -s "https://datasets-server.huggingface.co/rows?dataset=Anthropic/hh-rlhf&config=default&split=train&offset=5895&length=1" \ | python3 -c 'import json,sys; r=json.load(sys.stdin)["rows"][0]["row"]; print(r["chosen"][:160]); print(r["rejected"][:160])' # HelpSteer2 sizes (train 20324 + validation 1038 = 21362) curl -s "https://datasets-server.huggingface.co/size?dataset=nvidia/HelpSteer2" # the divisions used above python3 -c "print(round(100*2841/160800,2), round(100*108/8548,2), round(100*10/8434,2), round(100*11663/160800,2), round(100*67/2471,1), round(38.97/1.77,1))" # 1.77 1.26 0.12 7.25 2.7 22.0 ``` Two discrepancies to know about. The paper's evaluation denominator is 8,548 (Table 3); the Hugging Face test split has 8,552 rows. And the paper publishes record indices for its examples (for instance the pairs 14646 and 5895, 23320 and 27060) but no code or full flagged list, so the 2,841 and 108 cannot be re-derived outside Google. The Cleanlab figures are in Table 1 of [arXiv:2605.06036](https://arxiv.org/html/2605.06036). ## What would prove this wrong The 1.77% figure is an LLM's judgement about human judgements. If three independent human annotators re-label the 108 evaluation records and their majority agrees with Gemini 3.1 Pro's flip on fewer than 60% of them, the pipeline is measuring model preference, not rater error, and the low number is as unreliable as the high one. A dated prediction: by **31 March 2027**, at least one new paper will report a human-verified error rate for HH-RLHF, on any sample of 100 or more records, and it will fall between 2% and 20%, above the influence pipeline and well below the Cleanlab pass. A human-verified rate outside that band, published by that date, falsifies this. ## Sources - Song, Y., Watson, M., Grabowski, P., Qin, J. [Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation](https://arxiv.org/abs/2607.22766). arXiv:2607.22766v1, 24 July 2026. Sections 4.1, 4.2; Tables 2 and 3. - [Optimal Transport for LLM Reward Modeling from Noisy Preference](https://arxiv.org/abs/2605.06036). arXiv:2605.06036, 7 May 2026. Table 1. - Anthropic. [hh-rlhf dataset](https://huggingface.co/datasets/Anthropic/hh-rlhf), Hugging Face. Split sizes and row 5895 retrieved 15 September 2026. - NVIDIA. [HelpSteer2 dataset](https://huggingface.co/datasets/nvidia/HelpSteer2), Hugging Face. Split sizes retrieved 15 September 2026. - Northcutt, C., Jiang, L., Chuang, I. [Confident Learning: Estimating Uncertainty in Dataset Labels](https://arxiv.org/abs/1911.00068). JAIR 2021. - ACM SIGCHI. [Human-LLM Collaborative Annotation Through Effective Verification of LLM Labels](https://www.youtube.com/watch?v=TQCuLxk1jSM). YouTube, 8 May 2024. - BLOMEGA. [Data annotation research: the latest](https://blomega.com/research/data-annotation-latest-research/). ============================================================================== URL: https://blomega.com/research/llm-judge-soft-labels-human-disagreement-2026/ Published: 2026-09-15 | Updated: 2026-09-15 ============================================================================== --- title: "An LLM judge matches the majority label, but predicts human disagreement worse than 3 voters" url: https://blomega.com/research/llm-judge-soft-labels-human-disagreement-2026/ published: 2026-09-15 updated: 2026-09-15 source: BLOMEGA (https://blomega.com/) --- # An LLM judge matches the majority label, but predicts human disagreement worse than 3 voters Lab note · 15 September 2026 · BLOMEGA Claude-4-Sonnet, used as a judge, predicts the majority label on multi-annotator datasets at or near human F1. Asked for the distribution of human answers instead, its error on ChaosNLI is **0.200** DistCE against **0.070** for a random 20 of the 100 human annotations, and on the Anecdotes dataset **0.299** against **0.165** for 3 of 15 votes. A post-hoc alignment method posted on 1 September 2026 lowers the ChaosNLI figure to **0.174**, closing 20% of that gap. ## What changed, and when On **1 September 2026** Sebastian Steindl, Nikos Voskarides, Alberto Gasparin and Diego Marcheggiani posted [Post-hoc Alignment of LLM-judges to Human Judgment Distribution](https://arxiv.org/abs/2609.01073) (arXiv:2609.01073, v1). Most LLM-as-judge validation compares a model's output to an aggregated label: the majority vote or the mean rating. This paper also scores the model against the unaggregated human judgment distribution, the soft label, on five datasets chosen for different sources of disagreement. The datasets are SummEval (news summaries, three expert ratings per item), TopicalChat (dialogue responses, three crowd ratings), ChaosNLI (high-disagreement SNLI items re-annotated by 100 crowd workers each), DynaSent round 2 (adversarial sentiment), and Anecdotes (Reddit "who is in the wrong" posts, restricted to items with 15 votes). DynaSent and Anecdotes are capped at 1,500 items each, sampled as 500 per disagreement tercile by label entropy. Every result uses a 20/80 train/test split and 20 runs. The main judge is Claude-4-Sonnet, with GPT-OSS-120B and Qwen3-32B in an appendix; the authors report the backbone does not change the conclusions. Soft-label error is reported as distribution calibration error (DistCE) and Jensen-Shannon distance, lower is better. The human reference in Table 5 samples 20% of each item's annotations (minimum one) and scores that sample against the full distribution. That is what gives these numbers a human scale: 20 annotators on ChaosNLI, 3 votes on Anecdotes. ## The evidence table Soft-label DistCE, lower is better. Base judge is Claude-4-Sonnet prompted with soft-label in-context examples (the paper's SLP-SE setting). Values are from Tables 3 and 5 of arXiv:2609.01073v1; the gap-closed column is ours. | Dataset | Human sample (20% of annotations) | LLM judge | + NAPHA (predicted class) | + NAPHA (oracle class) | Gap to human closed (predicted / oracle) | Source | | --- | --- | --- | --- | --- | --- | --- | | ChaosNLI (100 annotations per item) | 0.070 | 0.200 | 0.174 | 0.153 | 20% / 36% | Tables 3, 5 | | Anecdotes (15 votes per item) | 0.165 | 0.299 | 0.272 | 0.172 | 20% / 95% | Tables 3, 5 | | DynaSent round 2 | 0.336 | 0.272 | 0.265 | 0.172 | LLM already better | Tables 3, 5 | | SummEval (3 expert ratings) | 0.264 | 0.371 | 0.303 | 0.263 | 64% / 101% | Tables 3, 5 | | TopicalChat (3 crowd ratings) | 0.238 | 0.374 | 0.363 | 0.312 | 8% / 46% | Tables 3, 5 | The hard-label picture is different. Macro F1 against the majority label, Claude-4-Sonnet against a sampled human, by disagreement tercile (Table 1): ChaosNLI 0.99 vs 0.93 (low), 0.81 vs 0.74 (medium), 0.61 vs 0.53 (high); Anecdotes 0.48 vs 0.50, 0.48 vs 0.42, 0.34 vs 0.35; DynaSent 0.95 vs 1.00, 0.81 vs 0.80, 0.40 vs 0.43. On the rating tasks the average Kendall correlation is 0.480 for the LLM and 0.542 for humans on SummEval, and 0.636 against 0.559 on TopicalChat (Table 2). By the majority label, the judge looks human-level. By the distribution, it looks like a small panel at best. Two readings of the table the paper does not spell out. First, the authors write that NAPHA "approximates human performance on the Anecdotes and DynaSent datasets". With predicted entropy classes, the deployable setting, Anecdotes goes from 0.299 to 0.272 against a human 0.165; the approximation holds only with oracle classes (0.172). Second, on DynaSent the base judge already beats the human reference, because the reference there is a single rating, and a single rating is a poor estimate of a distribution. _Sources: arXiv:2609.01073v1, Tables 4, 6 and 7 (LLM judge, SLP-SE) and Table 5 (human sample). The DynaSent high-entropy human bar (0.692) is drawn to scale and runs above the plot's top gridline._ ## On Anecdotes, the judge trails 3 voters most on the items voters agree about The per-tercile numbers carry the least intuitive result. On Anecdotes the gap between the LLM and 3 human votes is 0.182 on low-disagreement items, 0.141 on medium and 0.088 on high. On ChaosNLI the LLM's error is 2.0x the human sample on low-disagreement items, 2.8x on medium and 3.3x on high. The judge is not simply bad at hard cases. When humans nearly all agree, it still spreads probability across labels they did not choose; when humans split, a small human sample is itself noisy, and the judge's relative gap narrows on Anecdotes. Treat the ChaosNLI ratios and the Anecdotes differences as different measures: we computed both from the paper's tables, and they point in different directions on the high tercile. NAPHA is built around that stratification. It predicts an entropy class for each item and routes the judge's soft label through an alignment model trained for that class. The authors report that with predicted classes, alignment on the low-entropy class gets worse than the base judge, and the effect disappears with oracle classes. Their own mitigation is to skip NAPHA on items classified as low entropy. The bottleneck is the entropy classifier, which the oracle column quantifies. _Architecture as described in Section 4.2 of arXiv:2609.01073v1; values from Tables 3 and 5. Axis runs 0 to 0.20 at 3,200 px per unit._ ## What it means for annotation and evaluation pipelines **Majority-vote validation hides the failure.** If you validated an LLM judge by agreement with aggregated labels, you validated the part it does well. For any task where disagreement is signal (toxicity, ethics, helpfulness, sentiment with irony) the check that matters is distributional. The paper's human reference gives a practical yardstick: on these datasets the judge is worth somewhere between one rater (DynaSent) and fewer than three voters (Anecdotes), and well short of 20 crowd workers (ChaosNLI). **Keep collecting multiple annotations per item.** Our judgement: the case for single-annotator labelling plus an LLM tie-breaker is weakest exactly where teams use it, on subjective items. You cannot measure a judge's soft-label error without human soft labels, and the paper's limitations section notes three of its five datasets have fewer than six annotators per item because larger panels are rare and expensive. **Judge validation studies rarely collect disagreement at all.** ServiceNow's [AgentJudgeBench](https://arxiv.org/abs/2608.26623) (arXiv:2608.26623, 27 August 2026) finds six LLM judges converge to a 77% to 82% alignment band on hard tool-calling queries without ground truth. Its human check is 120 records scored by one annotator each, 92.7% agreement with the programmatic scorer, and the authors state they cannot report inter-annotator agreement. That is the norm this paper pushes against. **Budget for an entropy estimate.** The oracle column says most of NAPHA's remaining gap is knowing which items are contested. A small multi-annotator pilot on your own data produces that signal; the rest of the corpus can then be routed. This matches how we run [human-in-the-loop annotation](https://blomega.com/research/data-annotation-latest-research/): model first, humans where the model's confidence and the humans' agreement diverge. ## Check it yourself All numbers are in the arXiv HTML rendering. Table 3 has the dataset-level soft-label results, Table 5 the human sample, Tables 4, 6 and 7 the per-tercile values for Anecdotes, ChaosNLI and DynaSent, and Table 1 the hard-label F1. ``` open https://arxiv.org/html/2609.01073v1 # gap closed and ratios used above python3 - <<'PY' rows = { # human sample, base judge, NAPHA predicted, NAPHA oracle "ChaosNLI": (0.070, 0.200, 0.174, 0.153), "Anecdotes": (0.165, 0.299, 0.272, 0.172), "SummEval": (0.264, 0.371, 0.303, 0.263), "TopicalChat":(0.238, 0.374, 0.363, 0.312), } for k,(h,b,p,o) in rows.items(): print(f"{k:12} predicted {100*(b-p)/(b-h):.0f}% oracle {100*(b-o)/(b-h):.0f}%") tercile = {"low": (0.077,0.038), "med": (0.206,0.073), "high": (0.316,0.097)} print({k: round(l/h,1) for k,(l,h) in tercile.items()}) # ChaosNLI LLM/human PY # ChaosNLI predicted 20% oracle 36% # Anecdotes predicted 20% oracle 95% # SummEval predicted 64% oracle 101% # TopicalChat predicted 8% oracle 46% # {'low': 2.0, 'med': 2.8, 'high': 3.3} ``` ChaosNLI itself, with all 100 labels per item, is public at [github.com/easonnie/ChaosNLI](https://github.com/easonnie/ChaosNLI), so the entropy terciles can be rebuilt. The paper links no code. Its SummEval and TopicalChat soft-label rows are reported as averages over several rating dimensions with only three ratings per item behind them, so we treat those two rows as less directly comparable than ChaosNLI and Anecdotes. ## What would prove this wrong The claim rests on one closed-weights backbone for the main tables and on a human reference built by subsampling. It would be wrong if a replication with a different sampling rule (for example 5 fixed annotators per ChaosNLI item) put the human error above the LLM's 0.200, or if a newer judge prompted the same way scored below 0.100 on ChaosNLI without any alignment step. A dated prediction: by **31 March 2027**, no published LLM-as-judge result on ChaosNLI soft labels, prompt-only and without training on ChaosNLI annotations, will report DistCE at or below the 0.070 human-sample figure from this paper. One such result falsifies it. ## Sources - Steindl, S., Voskarides, N., Gasparin, A., Marcheggiani, D. [Post-hoc Alignment of LLM-judges to Human Judgment Distribution](https://arxiv.org/abs/2609.01073). arXiv:2609.01073v1, 1 September 2026. Tables 1 to 7; [HTML version](https://arxiv.org/html/2609.01073v1). - Verma, A., Saha, A. K., Subramanian, S., Aluru, S. H. [AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling](https://arxiv.org/abs/2608.26623). arXiv:2608.26623v1, 27 August 2026. Abstract and Appendix G. - Nie, Y., Zhou, X., Bansal, M. [What Can We Learn from Collective Human Opinions on Natural Language Inference Data?](https://aclanthology.org/2020.emnlp-main.734/) EMNLP 2020. Data: [github.com/easonnie/ChaosNLI](https://github.com/easonnie/ChaosNLI). - BLOMEGA. [Data annotation research: the latest](https://blomega.com/research/data-annotation-latest-research/). ============================================================================== URL: https://blomega.com/research/nvd-cwe-label-audit-2026/ Published: 2026-09-15 | Updated: 2026-09-15 ============================================================================== --- title: "Only 49.7% of NVD CWE labels match the code they describe" url: https://blomega.com/research/nvd-cwe-label-audit-2026/ published: 2026-09-15 updated: 2026-09-15 source: BLOMEGA (https://blomega.com/) --- # Only 49.7% of NVD CWE labels match the code they describe Lab note · 15 September 2026 · BLOMEGA An audit of **15,556** open-source CVEs against their own fix commits finds that **7,732** National Vulnerability Database CWE labels (**49.70%**) name the most specific weakness the code supports, **564** (**3.63%**) contradict that evidence, and manual review confirms **434** of the 564 as genuine mislabels. The evidence-inconsistent share was **0.99%** for 2017 CVEs and **4.54%** for 2025 CVEs. The paper was posted on 22 August 2026, four months after NIST stopped assigning weakness classifications to most new CVEs. ## What changed, and when On **22 August 2026** Yu Nong and Haipeng Cai (University at Buffalo), Yao Du (Macau University of Science and Technology) and Majid Behravan (Virginia Tech) posted [How Reliable Are NVD CWE Labels? A Large-Scale Semantic Audit with Seclometry](https://arxiv.org/abs/2608.21977) (arXiv:2608.21977, v1). It is the first measurement at this scale of a label set that vulnerability detectors, scanner evaluations and security benchmarks routinely treat as ground truth. The instrument is an agent pipeline the authors call CweAgent. It builds an eight-field "seclometry" record for each CWE definition and each CVE (root cause, trigger, violated invariant, observable effect, code pattern, source and sink roles, non-examples), gathers evidence with web search, archived advisories, the fix commit and CodeQL, and matches the two. Phase 1 uses Claude 4.6 Opus, phase 2 GPT-5.4-mini, and a Claude Sonnet 4.6 arbitrator sorts each disagreement into an outcome class. On a manually curated benchmark of 100 CVEs, CweAgent reaches 85% top-1 exact-match accuracy and 92% ambiguity-aware accuracy; the arbitrator reaches 90% overall accuracy on a separate 100-CVE test set. The context matters for anyone building on these labels. On **15 April 2026** NIST [announced](https://www.nist.gov/news-events/news/2026/04/nist-updates-nvd-operations-address-record-cve-growth) that CVE submissions had grown 263% between 2020 and 2025 and that it would prioritise enrichment for CVEs in CISA's Known Exploited Vulnerabilities catalog, CVEs in software used by the federal government, and CVEs in critical software as defined by Executive Order 14028. Everything else becomes "Lowest Priority - not scheduled for immediate enrichment", which means no NIST severity score, product mapping or weakness classification. Unenriched CVEs with an NVD publish date before 1 March 2026 moved to "Not Scheduled". NIST's announcement gives no count for that backlog. So for most CVEs published after April, the only CWE a dataset builder will get is the one the CVE Numbering Authority (CNA) assigned. The audit measures exactly how much that source varies. _NVD Gives Up: How NIST Made 29,000 CVEs Disappear Overnight, [Gula Tech Adventures](https://www.youtube.com/@GulaTechAdventures), 22 April 2026. Commentary on the 15 April NIST change that ended universal CWE enrichment. The backlog figure in the title is the channel's; NIST's own notice does not state one._ ## The evidence table Counts and shares are transcribed from Tables IV, V, VII and X of arXiv:2608.21977v1. "Ambiguity-aware" is S1+S2+S3. "Classifier error" is the share where CweAgent was wrong and NVD was right, which is the instrument's own error rate on this corpus. | Outcome (15,556 CVEs, 2017-2026) | Count | Share | Source | | --- | --- | --- | --- | | S1: exact match with the code-grounded label | 7,732 | 49.70% | arXiv:2608.21977, Table IV | | S2: overlap ambiguity | 3,890 | 25.01% | Table IV | | S3: defensibly alternative | 990 | 6.36% | Table IV | | S4: NVD label inconsistent with evidence | 564 | 3.63% | Table IV | | Classifier error (NVD correct, CweAgent wrong) | 2,380 | 15.30% | Table IV | | S4 confirmed as mislabels by manual review | 434 | 2.79% (our division) | Section VI; 130 excluded as defensible but imprecise | | Slice | N | Exact match | Ambiguity-aware | Inconsistent (S4) | Source | | --- | --- | --- | --- | --- | --- | | CNA: github_m | 7,599 | 45.85% | 78.62% | 3.53% | Table X | | CNA: mitre | 3,574 | 48.57% | 84.75% | 4.14% | Table X | | CNA: @huntrdev | 1,526 | 63.11% | 78.77% | 4.46% | Table X | | CNA: redhat | 324 | 39.20% | 83.95% | 2.47% | Table X | | CNA: wordfence | 79 | 83.54% | 91.14% | 0.00% | Table X | | CWE-79 cross-site scripting | 2,506 | 89.07% | 93.26% | 1.52% | Table VII | | CWE-787 out-of-bounds write | 344 | 47.09% | not reported here | 6.10% | Table VII | | CWE-122 heap-based buffer overflow | 201 | 44.28% | 79.60% | 11.94% | Table VII | | CWE-20 improper input validation | 461 | 0.00% | 65.29% | 2.60% | Table VII | | CWE-284 improper access control | 200 | 0.00% | 55.50% | 4.50% | Table VII | Two things in these rows are easy to miss. First, the instrument is wrong more often than NVD is inconsistent: 2,380 classifier errors against 564 S4 records, a ratio of 4.2 to 1. The headline 49.70% is a floor on NVD quality as seen through an 85%-accurate lens, and the authors separate the two cases in the table rather than hiding them. Second, the CNA spread among the 15 largest assigners runs from 39.20% exact match (redhat) to 83.54% (wordfence), a 44.3 point range. The weakness type moves the rate as much as the assigner: CWE-20, CWE-200 and CWE-284 score 0.00% exact match, which reads as the audit never accepting those broad categories as the most specific label (our reading of the table, not a sentence in the paper). _Source: arXiv:2608.21977v1, Table V. The 2024 to 2026 cohort alone contributes 228 of the 564 inconsistent labels (40.4%), because it is also the largest._ ## A heap over-read becomes an out-of-bounds write in three hops The 434 confirmed mislabels fall into six patterns in the authors' open coding (Table XII): consequence named instead of root cause 23.1%, wrong sibling within a family 21.3%, wrong injection sink or interpreter 19.5%, authentication and authorisation conflated 15.3%, discouraged or wrong-branch label 10.3%, and several weaknesses collapsed into one CWE 7.8%, with a 3.9% residual on file and path handling. The sibling pattern has a structural cause the paper names. NVD normalises labels to MITRE's CWE View-1003, which does not contain CWE-122, so a CNA's heap-overflow label gets mapped up to the parent CWE-787, out-of-bounds write. If the actual bug is a read, the normalisation step moves the label further from the truth. CVE-2022-1160 in vim is the paper's worked example, and the NVD API still shows both labels today. _Labels as returned by the NVD CVE API 2.0 on 15 September 2026; evidence reading from arXiv:2608.21977v1, Figure 7._ _Source: arXiv:2608.21977v1, Table IV. The dashed segment is the auditor's own error, not NVD's._ ## What it means for anyone training or scoring on CWE labels **An exact-match score against NVD CWE has a ceiling well below 100% for a correct model.** A classifier that is right 85% of the time on a curated benchmark agreed exactly with NVD on 49.70% of this corpus. If your vulnerability-type benchmark reports top-1 accuracy against NVD labels, part of the gap between models is label choice among defensible siblings, not model skill. The 25.01% overlap-ambiguity bucket is the size of that problem. **Score ambiguity-aware, and drop the umbrella categories as targets.** CWE-20, CWE-200 and CWE-284 together account for 1,109 CVEs in Table VII with 0.00% exact match each. Training a model to emit them teaches it to be unspecific. Our judgement: map them to "unspecified" in training data and exclude them from exact-match evaluation. **Weight labels by who assigned them.** After NIST's April change, the assigning CNA's label is the label for most new CVEs. A 44.3 point exact-match range across the 15 largest CNAs is a larger quality signal than the year. For memory-safety data specifically, re-derive the read or write distinction from the patch; CWE-122 at 11.94% inconsistency is the worst of the top 20. **An LLM auditor needs a human pass on the tail.** The pipeline flagged 564 records; manual review kept 434 and set aside 130 (23%). That is the same human-in-the-loop shape we see in text annotation: the model narrows the search, a person decides. For procurement, the fields to ask a security-data vendor for are the assigning CNA, the CWE as assigned before normalisation, and whether a human checked the root cause against the fix. This is the provenance argument we make for [chain of title](https://blomega.com/guides/data-provenance-chain-of-title/), applied to labels. ## Check it yourself The worked example is live in NVD. The API returns both the CNA's label and NVD's. ``` curl -s "https://services.nvd.nist.gov/rest/json/cves/2.0?cveId=CVE-2022-1160" \ | python3 -c 'import json,sys; c=json.load(sys.stdin)["vulnerabilities"][0]["cve"] for w in c["weaknesses"]: print(w["source"], [d["value"] for d in w["description"]])' # security@huntr.dev ['CWE-122'] # nvd@nist.gov ['CWE-787'] ``` The outcome table sums exactly, which is a quick transcription check: ``` python3 -c "print(7732+3890+990+564+2380, round(100*434/15556,2), round(2380/564,2))" # 15556 2.79 4.22 ``` The per-year contribution of 2024 to 2026 comes from Table V: 2,074 x 2.85% + 2,179 x 4.54% + 2,265 x 3.09% rounds to 228 CVEs. The authors' code, data and confirmed-mislabel set are at [figshare.com/s/6ba6621a69a2dac5b516](https://figshare.com/s/6ba6621a69a2dac5b516), which responded to a request on 15 September 2026; we have not re-run the pipeline. The CWE-122 and View-1003 relationship can be checked at [cwe.mitre.org/data/definitions/122.html](https://cwe.mitre.org/data/definitions/122.html) and [cwe.mitre.org/data/definitions/1003.html](https://cwe.mitre.org/data/definitions/1003.html). ## What would prove this wrong The central number depends on an 85%-accurate instrument, so the obvious attack is the instrument. If an independent team hand-labels a random sample of at least 300 CVEs from the same corpus and finds exact agreement with NVD above 60%, the 49.70% figure reflects CweAgent's taxonomy preferences more than NVD's errors. A dated prediction: by **31 March 2027**, fewer than half of the 434 confirmed mislabels in the figshare set will show a changed NVD CWE in the CVE API. The authors report NVD acknowledged the list; acknowledgment is not correction, and NIST's April policy puts most of these older records outside its enrichment priorities. If more than half have been corrected by that date, this prediction is wrong and NVD's correction loop is healthier than its April notice suggests. ## Sources - Nong, Y., Du, Y., Behravan, M., Cai, H. [How Reliable Are NVD CWE Labels? A Large-Scale Semantic Audit with Seclometry](https://arxiv.org/abs/2608.21977). arXiv:2608.21977v1, 22 August 2026. Tables IV, V, VII, X, XII; [HTML version](https://arxiv.org/html/2608.21977). - NIST. [NIST Updates NVD Operations to Address Record CVE Growth](https://www.nist.gov/news-events/news/2026/04/nist-updates-nvd-operations-address-record-cve-growth). 15 April 2026. - NVD. [CVE-2022-1160 detail](https://nvd.nist.gov/vuln/detail/CVE-2022-1160) and [CVE API 2.0](https://nvd.nist.gov/developers/vulnerabilities). Retrieved 15 September 2026. - MITRE. [CWE View-1003: Weaknesses for Simplified Mapping of Published Vulnerabilities](https://cwe.mitre.org/data/definitions/1003.html). - Authors' artifact: [figshare.com/s/6ba6621a69a2dac5b516](https://figshare.com/s/6ba6621a69a2dac5b516). - Gula Tech Adventures. [NVD Gives Up: How NIST Made 29,000 CVEs Disappear Overnight](https://www.youtube.com/watch?v=JY63jOpWFPc). YouTube, 22 April 2026. - BLOMEGA. [Data annotation research: the latest](https://blomega.com/research/data-annotation-latest-research/) and [Unpaid annotation tasks in ACL papers](https://blomega.com/research/unpaid-annotation-tasks-acl-2018-2025/). ============================================================================== URL: https://blomega.com/guides/cost-of-a-dubbed-minute-2026/ Published: 2026-09-10 | Updated: 2026-09-10 ============================================================================== --- title: "A dubbed minute costs $0.33 to $9.00, and the watermark discount is gone" url: https://blomega.com/guides/cost-of-a-dubbed-minute-2026/ published: 2026-09-10 updated: 2026-09-10 source: BLOMEGA (https://blomega.com/) --- # A dubbed minute costs $0.33 to $9.00, and the watermark discount is gone Lab note · 10 September 2026 · BLOMEGA Read straight off the vendors' own price lists today: ElevenLabs charges **$0.33** per minute for a watermarked Dubbing v1 job, **$0.50** for the same job clean, and **$2.20** for Dubbing v2, which has no watermark option at all. Rask AI's annual plans normalize to **$0.78 to $1.32** per minute, tripling to **$2.34 to $3.96** with enhanced lip-sync and reaching **$9.00** on overage minutes. That is a 27-fold spread across published prices, and not one of these vendors publishes a quality number next to the price. Dubbing v2 reached the API on **10 August 2026**, eight days after Article 50 transparency obligations started applying. ## What moved between June 2025 and August 2026 Four dated events changed what a buyer can actually see on a price page. **26 June 2025.** RWS announced it had [acquired the IP behind Papercup's AI dubbing technology](https://www.rws.com/about/news/2025/rws-acquires-papercups-ip/). As of today the consequence is visible in DNS: `papercup.com` and `papercup.com/pricing` both return HTTP 301 to an RWS AI dubbing and voiceover page carrying `utm_campaign=papercup+migration`. One of the four vendors most often named in dubbing comparisons no longer has a website, let alone a price. **28 May 2026.** ElevenLabs [announced Dubbing v2](https://elevenlabs.io/blog/introducing-dubbing-v2), describing a model that conditions directly on the original performance rather than working from a transcript, across 90+ languages. The post contains no benchmark, no listening-test result and no comparison figure. **2 August 2026.** The EU AI Act's [Article 50](https://artificialintelligenceact.eu/article/50/) transparency obligations began to apply, requiring providers of systems that generate synthetic audio, image, video or text to mark output in a machine-readable format. We covered what survives a dubbing mix in [a separate note](https://blomega.com/guides/eu-ai-act-article-50-ai-dubbing-watermarks/). **6 to 10 August 2026.** ElevenLabs posted [Dubbing v2 is now available via ElevenAPI](https://elevenlabs.io/blog/dubbing-api), dated 6 August, and the [changelog entry for 10 August 2026](https://elevenlabs.io/docs/changelog/2026/8/10) records API availability. The two dates differ by four days and we did not resolve which one is the switch-over; treat the API as live in the second week of August. _Eight days separate the start of the marking duty from the release that removed the marking option on the flagship model. We have no evidence the two are connected, and we are not claiming they are. The sequence is what a buyer has to plan around either way._ The pricing consequence is the part nobody wrote up. Dubbing v1 sold a watermarked dub for less than a clean one. Dubbing v2, per the [ElevenLabs documentation](https://elevenlabs.io/docs/overview/capabilities/dubbing), "does not include a watermark toggle", and paid-tier dubs are not watermarked while free-tier dubs are watermarked automatically. The cheapest way to buy a marked dub from the most-quoted vendor in the category was removed in the same month the marking regime began to apply. ## What does a minute actually cost? Two vendors quote in minutes, one quotes in credits, and two quote nothing. Where a vendor sells an annual plan with an included minute allowance, the per-minute rate below is that plan's annual price divided by its annual minutes, on the assumption the allowance is fully consumed. That assumption is generous to the vendor: unused minutes raise the real rate. _Nine published rates for the same unit of work. The green bars are the only two prices in the category that come with a machine-readable provenance mark attached, and both are on a model in maintenance mode._ | Vendor and configuration | As published | Per minute | Watermark option | Source | | --- | --- | --- | --- | --- | | ElevenLabs Dubbing v1, automatic, watermarked | $0.33 per minute | $0.33 | yes, and cheaper | elevenlabs.io/pricing/api | | ElevenLabs Dubbing v1, automatic, no watermark | $0.50 per minute | $0.50 | yes, opt out | elevenlabs.io/pricing/api | | ElevenLabs Dubbing Studio (v1, maintenance mode) | $0.50 per minute | $0.50 | not stated | elevenlabs.io/pricing/api | | ElevenLabs Dubbing v2 | $2.20 per minute | $2.20 | no toggle; paid tier unwatermarked | elevenlabs.io/pricing/api and /docs | | Rask AI Creator, standard lip-sync | $33/mo billed $396/yr, 300 min/yr | $1.32 (computed) | not stated | rask.ai/pricing | | Rask AI Creator Pro, standard lip-sync | $78/mo billed $936/yr, 1,200 min/yr | $0.78 (computed) | not stated | rask.ai/pricing | | Rask AI Business, standard lip-sync | $500/mo billed $6,000/yr, 6,000 min/yr | $1.00 (computed) | not stated | rask.ai/pricing | | Rask AI, enhanced lip-sync (beta) | 1 min of video consumes 3 min of allowance | $2.34 to $3.96 (computed) | not stated | rask.ai/pricing | | Rask AI overage minute, enhanced lip-sync | $3 per extra minute, times 3 | $9.00 (computed) | not stated | rask.ai/pricing | | HeyGen Creator | $29/mo, 600 credits | not reported ($0.0483 per credit, computed) | not stated | heygen.com/pricing | | HeyGen Pro | $49/mo, 1,000 credits | not reported ($0.0490 per credit, computed) | not stated | heygen.com/pricing | | HeyGen Business | $149/mo, 1,500 credits, +$20 per seat | not reported ($0.0993 per credit, computed) | not stated | heygen.com/pricing | | Deepdub | no pricing page | not reported | not stated | deepdub.ai/pricing returns HTTP 404 | | Papercup | domain redirects to RWS | not reported | not stated | papercup.com returns HTTP 301 | | ElevenLabs Scribe v2 (speech to text, for reference) | $0.22 per hour | $0.0037 | not applicable | elevenlabs.io/pricing/api | | ElevenLabs TTS v3 / v2 Multilingual (for reference) | $0.10 per 1,000 characters | ~$0.09 (computed, see note) | not applicable | elevenlabs.io/pricing/api | Note on the TTS row: we assume roughly 150 spoken words per minute at about six characters per word including spaces, so about 900 characters per minute of speech. That is our assumption, not a vendor figure, and it moves with speaking rate and script. Three things fall out of that table that are not on any vendor's page. **Rask's most expensive plan has the second-cheapest minute.** Creator Pro at $936 a year for 1,200 minutes is $0.78 a minute. Business at $6,000 a year for 6,000 minutes is $1.00. The plan that costs 6.4 times more per month costs **28% more per minute**. Same for HeyGen on credits: Business is $0.0993 per credit against Creator's $0.0483, a factor of 2.06 in the wrong direction. Buying up a tier in this category buys seats, support and limits, not unit economics. **The v1 to v2 jump is 4.4 times.** $0.50 to $2.20 on the clean price, 6.7 times on the watermarked price. No published quality figure accompanies either model, so a buyer choosing between them is choosing between a documented price and an undocumented improvement. **HeyGen's own comparison page disagrees with HeyGen's pricing page.** The comparison page, last updated 19 May 2026, states Creator is "$24/month" with "no per-minute or per-credit charges at any tier". The pricing page today shows Creator at $29 a month with a 600-credit monthly allowance. We did not resolve the difference and are reporting both. ## Where $2.20 goes when the voice costs nine cents The reference rows in Table 1 make the cost structure legible. ElevenLabs sells the two component operations of a dub separately and publicly: transcription at $0.22 an hour and speech synthesis at $0.10 per 1,000 characters. Put a minute of speech through both and the raw model cost is around nine cents. _Nine cents of the $2.20 is the voice. The other $2.11 is the part a buyer cannot inspect, and it is the part that determines whether the dub is right._ The same subtraction run across the three published dubbing prices gives a wrapper of $0.24 on the $0.33 watermarked v1 job (72% of the price), $0.41 on the $0.50 clean v1 job (81%), and $2.11 on the $2.20 v2 job (96%). The share the buyer cannot inspect grew with every release. That is the useful reframing for procurement. You are not buying synthesis, which is a commodity with a public unit price that has fallen for three years. You are buying the alignment, the timing, the speaker handling and whatever review sits inside the pipeline, and none of the five vendors itemises the last one. A vendor charging $2.20 and a vendor charging $0.78 may differ by 100% in human review or by 0%, and the price lists do not distinguish those cases. | Vendor | Per-minute price | Language count | Quality figure | Provenance mark option | Source | | --- | --- | --- | --- | --- | --- | | ElevenLabs | published | 90+ with dialect variants (en-AU, es-MX, fr-CA) | none published | v1 yes, v2 no toggle | elevenlabs.io/docs, /pricing/api | | Rask AI | derivable from plans | not stated on pricing page | none published | not stated | rask.ai/pricing | | HeyGen | not published (credits only) | not stated on pricing page | none published | not stated | heygen.com/pricing | | Deepdub | not published | not reported | none published | not stated | deepdub.ai/pricing HTTP 404 | | Papercup (now RWS) | not published | not reported | none published | not stated | papercup.com HTTP 301 to rws.com | ## What to do with this if you are buying dubbing **Price the plan, not the sticker.** Divide annual plan cost by annual included minutes before comparing anything, then multiply by the lip-sync consumption factor you will actually use. Rask's enhanced lip-sync consumes three minutes of allowance per minute of video, which turns a $0.78 plan into a $2.34 plan without any line item changing. Then assume you will not consume the full allowance and recompute at 60% utilisation, which is where most annual minute plans land in our experience. At 60%, Rask Creator Pro standard is $1.30 a minute, not $0.78. **Ask what the wrapper contains, in writing.** The 96% of the Dubbing v2 price that is not synthesis is the whole product. The three questions that separate vendors are: is a human native speaker reviewing any output, on what sampling fraction, and is the translation step a general MT system or one tuned per locale. None of these are on any price page. **If you need a machine-readable mark, check before you migrate.** The v1 to v2 migration removes the watermark toggle from the category leader's flagship model. If your Article 50 assessment concluded that your output needs marking, moving to v2 changes your compliance posture, and it is not a line item anyone will flag for you. Our note on [whether the mark survives a dubbing mix](https://blomega.com/guides/eu-ai-act-article-50-ai-dubbing-watermarks/) covers what the mark is worth once it exists. **Treat the absence of a quality figure as the finding.** Five vendors, zero published evaluation numbers, and a 27-fold price spread. In any market where price varies that much and quality is unmeasured, the price is carrying information about packaging rather than about output. Run your own listening test on your own content in your own top three locales before signing anything, because nobody is going to hand you a number. **The consolidation is real and it removes options.** Papercup was the vendor most often recommended for human-in-the-loop dubbing, and its domain is now a redirect. That is one fewer independent price point in a category that already publishes almost nothing. This is our reading rather than a company statement: expect the enterprise tier of this market to move further toward quote-only pricing, not less. ## Check it yourself Every number in Table 1 is reproducible in about five minutes, and the two redirect findings take one command each. ``` # 1. the Papercup redirect and the missing Deepdub price page for u in https://www.papercup.com/pricing https://www.papercup.com/ \ https://deepdub.ai/pricing https://deepdub.ai/; do printf '%s -> ' "$u" curl -s -o /dev/null -w '%{http_code} %{redirect_url}\n' -A "Mozilla/5.0" "$u" done # 2026-09-10 output: # https://www.papercup.com/pricing -> 301 https://www.rws.com/localization/services/ # translation-services/video-and-audio-translation/ai-dubbing-and-vo?utm_campaign=papercup+migration # https://deepdub.ai/pricing -> 404 # 2. the price pages, read them yourself and date-stamp what you see # elevenlabs.io/pricing/api elevenlabs.io/docs/overview/capabilities/dubbing # rask.ai/pricing heygen.com/pricing # 3. normalize any plan to dollars per minute python3 - <<'PY' plans = { # name: (annual_price_usd, annual_minutes) "Rask Creator": (396, 300), "Rask Creator Pro": (936, 1200), "Rask Business": (6000, 6000), } for name, (price, mins) in plans.items(): base = price / mins print(f"{name:18s} standard ${base:.2f}/min enhanced-lipsync ${base*3:.2f}/min" f" at 60% utilisation ${base/0.6:.2f}/min") PY # 4. the raw model cost of a minute, for the wrapper calculation # STT: $0.22/hr / 60 = $0.0037/min # TTS: $0.10/1000 chars * ~900 chars/min = $0.090/min # wrapper share of a $2.20 dub = 1 - 0.094/2.20 = 95.7% ``` Three numbers are worth recording per vendor per quarter: the per-minute rate you actually paid including unused allowance, the fraction of delivered minutes your reviewers sent back, and whether the output still carries a machine-readable mark after your delivery encode. The third one changes without notice, as the v1 to v2 migration shows. ## What would prove this wrong The claim under test is that AI dubbing in 2026 is priced without any published quality signal, and that price differences across vendors therefore carry no information about output quality. It is wrong if, by **1 June 2027**, at least two of the five vendors named here publish a per-minute price alongside a reproducible quality number for at least ten target languages, naming the evaluation set and the rater population, in a document a buyer can link to. Today the count is zero of five. A vendor-run listening test with an unnamed rater pool and no released audio would not settle it. A second prediction, marked as judgement rather than finding: the $0.33 watermarked tier will disappear from the ElevenLabs price list before the v1 model is retired, because it is the only line item in the category that prices provenance marking as a discount rather than a feature, and that is an awkward position to hold while Article 50 enforcement builds. If it survives past 1 June 2027, that reading was wrong. ## FAQ ### How much does AI video localization cost in 2026? Between $0.33 and $9.00 per minute of source media from published price lists read on 10 September 2026. ElevenLabs lists Dubbing v1 at $0.33 watermarked and $0.50 clean, and Dubbing v2 at $2.20. Rask AI's annual plans compute to $1.32, $0.78 and $1.00 per minute on Creator, Creator Pro and Business, tripling with enhanced lip-sync and reaching $9.00 on an overage minute. HeyGen publishes credits with no minute rate. Deepdub and Papercup publish no price. ### Does ElevenLabs still offer a watermarked dubbing discount? Only on the older model. The API price list shows Dubbing v1 at $0.33 per minute watermarked against $0.50 clean, a 51.5% premium to remove the mark. The documentation states Dubbing v2 does not include a watermark toggle and that paid-tier dubs are not watermarked, while free-tier dubs are watermarked automatically. Dubbing v2 reached the API on 10 August 2026, eight days after Article 50 transparency obligations began to apply on 2 August 2026. ### Why is Dubbing v2 more expensive than Dubbing v1? No reason and no quality figure have been published. The observable differences are the price ($2.20 against $0.33 to $0.50), a model that conditions on the original performance rather than a transcript, 90+ languages with dialect variants such as en-AU, es-MX and fr-CA, a 3 GB API upload limit against v1's 1 GB and 45 minutes, and the fact that Dubbing Studio remains v1-only and is described as in maintenance mode. ### Which AI dubbing vendors publish their prices? Three of five, and only two in minutes. ElevenLabs publishes per-minute API pricing. Rask AI publishes plans, included minutes and a $3 overage from which a minute rate can be computed. HeyGen publishes plans and credit allowances but no credit-per-minute rate. Deepdub's pricing URL returned HTTP 404 on 10 September 2026 and papercup.com returned HTTP 301 to an RWS page, following RWS's acquisition of Papercup's IP announced 26 June 2025. ### Is AI dubbing cheaper than human dubbing? On published data, that comparison cannot be made. AI vendors publish per-minute prices; studio and agency dubbing is quoted per project and the ranges that circulate online come from vendor marketing pages rather than rate cards. We could not find a primary, citable human dubbing rate card, so we are not printing one. The honest statement is that AI dubbing has published prices and human dubbing does not, which is a difference in transparency before it is a difference in cost. ## Sources - [ElevenLabs API pricing](https://elevenlabs.io/pricing/api), read 10 September 2026. Dubbing v1 $0.33 per minute watermarked and $0.50 without; Dubbing Studio $0.50; Dubbing v2 $2.20; Scribe v2 $0.22 per hour and Scribe v2 Realtime $0.39 per hour; text to speech v3 and v2 Multilingual $0.10 per 1,000 characters, Flash and Turbo $0.05. - [ElevenLabs dubbing documentation](https://elevenlabs.io/docs/overview/capabilities/dubbing), read 10 September 2026. "Dubbing v2 does not include a watermark toggle"; free-tier dubs watermarked automatically and paid-tier dubs not; 90+ languages with dialect variants; v2 3 GB API upload limit against v1's 1 GB and 45 minutes; Dubbing Studio v1-only and in maintenance mode; v2 transcript editing via API on Enterprise plans only. - ElevenLabs, [Introducing Dubbing v2](https://elevenlabs.io/blog/introducing-dubbing-v2), 28 May 2026, and [Dubbing v2 is now available via ElevenAPI](https://elevenlabs.io/blog/dubbing-api), 6 August 2026, with the [changelog entry dated 10 August 2026](https://elevenlabs.io/docs/changelog/2026/8/10). Model description, 90+ languages, no evaluation figures in either post. - [Rask AI pricing](https://www.rask.ai/pricing), read 10 September 2026, annual billing. Creator $33/mo billed $396/yr for 300 min/yr; Creator Pro $78/mo billed $936/yr for 1,200 min/yr; Business $500/mo billed $6,000/yr for 6,000 min/yr; $3 per extra minute on Creator Pro and Business; standard lip-sync 1 minute of video equals 1 minute of allowance, enhanced lip-sync equals 3, described as beta and subject to change. - [HeyGen pricing](https://www.heygen.com/pricing), read 10 September 2026. Free $0 with 3 videos per month; Creator $29/mo with 600 credits; Pro $49/mo with 1,000 credits; Business $149/mo with 1,500 credits plus $20 per additional seat; Enterprise by quote. Credit consumption stated to vary with model, duration and complexity. Compare [HeyGen's own comparison page](https://www.heygen.com/blog/heygen-vs-rask-ai-vs-maestra-vs-kapwing), last updated 19 May 2026, which states a $24 Creator plan and "no per-minute or per-credit charges at any tier". - RWS, [RWS acquires Papercup's IP to power next-gen AI dubbing for enterprise clients](https://www.rws.com/about/news/2025/rws-acquires-papercups-ip/), 26 June 2025, [Business Wire release of the same date](https://secure.businesswire.com/news/home/20250626469981/en/RWS-Acquires-Papercups-IP-to-Power-Next-Gen-AI-Dubbing-for-Enterprise-Clients). Acquisition of the technology IP, human-in-the-loop positioning, 1,800 in-house linguists and a network of over 40,000 language experts. - HTTP status observations made by BLOMEGA on 10 September 2026 with `curl`: `papercup.com` and `papercup.com/pricing` return 301 to the RWS AI dubbing and voiceover page with `utm_campaign=papercup+migration`; `deepdub.ai/pricing` returns 404 while `deepdub.ai/` returns 200. The commands are in the "Check it yourself" section. - European Union, [AI Act Article 50, transparency obligations for providers and deployers of certain AI systems](https://artificialintelligenceact.eu/article/50/), applicable from 2 August 2026, and the European Commission's [FAQ on Article 50 transparency obligations](https://digital-strategy.ec.europa.eu/en/faqs/transparency-obligations-under-article-50-ai-act). Related BLOMEGA guides: [Article 50 applies to your dub, and the watermark it asks for dies in your mix](https://blomega.com/guides/eu-ai-act-article-50-ai-dubbing-watermarks/) · [Localization is the new default](https://blomega.com/guides/localization-the-new-default/) · [Your multilingual LLM judge prefers the machine translation](https://blomega.com/guides/multilingual-llm-judge-translationese-bias/) · [Consented AI training data providers](https://blomega.com/guides/consented-ai-training-data-providers/). ============================================================================== URL: https://blomega.com/guides/per-language-token-ceiling/ Published: 2026-09-10 | Updated: 2026-09-10 ============================================================================== --- title: "A 30-trillion-token corpus buys Basque a 160-million-parameter model" url: https://blomega.com/guides/per-language-token-ceiling/ published: 2026-09-10 updated: 2026-09-10 source: BLOMEGA (https://blomega.com/) --- # A 30-trillion-token corpus buys Basque a 160-million-parameter model Lab note · 10 September 2026 · BLOMEGA HPLT 3.0 is the largest openly published multilingual pretraining collection: **30 trillion** sub-word tokens across close to 200 language-script combinations. Its Table 1 prints the per-language counts, and English holds **16T** of them while Galician holds **3.1B** and Basque **3.2B**. Convert those counts through the Chinchilla ratio of 20 tokens per parameter and you get the largest model each language can compute-optimally support: **800B parameters for English, 160M for Basque**. At the data intensity Meta actually used for Llama 3, 1,875 tokens per parameter, Basque buys **1.7M parameters**. That ceiling, not vendor effort, is what sets quality in your smaller locales. ## What got published, and when Three releases since late 2024 made per-language accounting possible. Before them, "supports 100 languages" was an unfalsifiable claim because nobody printed the denominator. **8 December 2024.** Hugging Face released [FineWeb2](https://huggingface.co/datasets/HuggingFaceFW/fineweb-2), over 1,000 language-script subsets built from 96 Common Crawl snapshots spanning summer 2013 to April 2024. The dataset card lists the configs by language-script pair but does not print per-language token counts, so you have to count them yourself. **13 March 2025.** The HPLT consortium posted [An Expanded Massive Multilingual Dataset for High-Performance Language Technologies](https://arxiv.org/abs/2503.10267) (arXiv:2503.10267, ACL 2025 main proceedings): 8T tokens of monolingual text covering 193 languages, plus 380M parallel sentence pairs covering 51 languages. Note the second number. Parallel data, the kind that trains translation directly, exists for roughly a quarter of the languages the monolingual side covers. **2 November 2025.** Stephan Oepen and 31 co-authors posted [HPLT 3.0](https://arxiv.org/abs/2511.01066) (arXiv:2511.01066, revised to v3 on 19 April 2026). 30T sub-word tokens, close to 200 language-script combinations, 57 encoder-decoder models, and the table this article is built on. **4 December 2025.** The v3 revision of [EMMA-500](https://arxiv.org/abs/2409.17892) (arXiv:2409.17892) documented the MaLA corpus: 939 languages, 824M documents, 74.255 billion whitespace-delimited tokens, average document length 90.12 tokens. It also states the cut-off that matters most in this whole article. Of the 939 languages, 546 have more than 100k tokens, and those 546 are what EMMA-500 was trained on. The other 393 were collected and then left out. ## How many tokens does each language actually have? Table 1 of HPLT 3.0 prints document counts, token counts, average document length and token share for English, for the multilingual remainder, and for nine individual languages. The chart below is those nine plus English, on a log axis, because a linear axis renders eight of the ten bars as a single line. _Nine languages out of close to 200, and already three orders of magnitude of spread. Galician and Basque are official languages of an EU member state with public broadcasters and a state-funded digitisation programme behind them. They are not the bottom of the distribution._ The nine non-English rows in that chart sum to 11.86% of the non-English portion of the corpus, using the shares printed in the same table. Nine languages, close to 200 in the collection, and just under an eighth of everything that is not English. Now the conversion. Hoffmann and colleagues ([arXiv:2203.15556](https://arxiv.org/abs/2203.15556), 29 March 2022) established that for a fixed compute budget, model size and training tokens should scale together, with Chinchilla at 70B parameters trained on four times the data of Gopher at 280B and beating it. The ratio commonly read off that result is about 20 tokens per parameter. Meta's own [Llama 3 announcement](https://ai.meta.com/blog/meta-llama-3/) (18 April 2024) states the Chinchilla-optimal budget for an 8B model is around 200B tokens, which is 25 to 1, and then reports training both the 8B and 70B models on up to 15T tokens because performance kept improving log-linearly. For Llama 3 8B that is 1,875 tokens per parameter, 94 times the Chinchilla ratio. So there are two ceilings, and both are computed the same way: divide the available tokens by the ratio. The Chinchilla column below is the generous reading. The Llama 3 column is what a lab building a competitive model in 2026 actually consumes. | Language | Documents | Tokens | Ceiling at 20 tok/param | Ceiling at 1,875 tok/param | Source | | --- | --- | --- | --- | --- | --- | | English | 18B | 16T | 800B params | 8.53B params | arXiv:2511.01066 Table 1 | | All non-English combined | 11B | 13T | 650B params | 6.93B params | arXiv:2511.01066 Table 1 | | Spanish | 725M | 658B | 32.9B params | 351M params | arXiv:2511.01066 Table 1 | | French | 603M | 584B | 29.2B params | 311M params | arXiv:2511.01066 Table 1 | | Czech | 107M | 126B | 6.30B params | 67.2M params | arXiv:2511.01066 Table 1 | | Ukrainian | 80M | 81B | 4.05B params | 43.2M params | arXiv:2511.01066 Table 1 | | Finnish | 49M | 73B | 3.65B params | 38.9M params | arXiv:2511.01066 Table 1 | | Norwegian | 37M | 52B | 2.60B params | 27.7M params | arXiv:2511.01066 Table 1 | | Catalan | see note | 22B | 1.10B params | 11.7M params | arXiv:2511.01066 Table 1 | | Basque | 3.2M | 3.2B | 160M params | 1.71M params | arXiv:2511.01066 Table 1 | | Galician | 4.0M | 3.1B | 155M params | 1.65M params | arXiv:2511.01066 Table 1 | | MaLA corpus, all 939 languages | 824M | 74.255B (whitespace) | 3.71B params | 39.6M params | arXiv:2409.17892 Table 1 | | A MaLA tail language at the 100k cut-off | not reported | 100k (whitespace) | 5,000 params | 53 params | arXiv:2409.17892 | Note on Catalan: the row we read prints 2.6M documents, 22B tokens and an average document length of 853. Those three do not multiply out, and 22B divided by 853 implies roughly 26M documents. We flag the inconsistency rather than pick a side, and we use only the token count, which is what the ceiling depends on. Note on units: HPLT counts sub-word tokens, MaLA counts whitespace-delimited tokens. A whitespace token is worth more than one sub-word token, so the MaLA rows understate in HPLT's units, probably by a factor of two to three for the scripts involved. We did not convert. Treat the MaLA ceilings as a floor, and note that the direction of the error makes the tail look better than it is on the model-size axis while making tokenization worse on the cost axis. Two numbers out of that table are worth saying on their own. The entire MaLA corpus, 939 languages, everything anyone has managed to assemble across the long tail, is 74.255B tokens. Ukrainian alone in HPLT 3.0 is 81B. **Every low-resource language on earth, pooled, is smaller than Ukrainian.** And English at 16T is roughly 215 times that pooled corpus. ## Why the ceiling is a ceiling and not a budget line The reason this is not fixable with money inside the current pipeline is that every step downstream of the corpus is multiplicative on what the corpus contains. The diagram traces one locale through it. _Galician at the Llama 3 data ratio supports a model smaller than BERT base. That is the mechanism behind "the translation is grammatical but reads wrong": the in-language capability was never trained, it was transferred._ The corpus builders drew the line themselves, and where they drew it is the most useful operational number in either paper. HPLT 3.0 states its team did not train models on languages with fewer than roughly 0.25M documents. MaLA collected 939 languages and trained on 546. Both groups are funded, motivated and public-good oriented, and both stopped at almost exactly the same place. If the people assembling the data will not train on a language, a commercial lab optimising for benchmark scores certainly will not. English dominance is not uniform across corpora, which matters when a vendor tells you which corpus they used. Table 1 of HPLT 3.0 prints the English token share for four collections side by side. _A lower English percentage does not mean more non-English data. MADLAD-400 is 38% English because it is small overall, at 1.7T English tokens against FineWeb's 17T. Percentage share is the wrong question. Absolute in-language volume is the right one._ | Corpus | Languages claimed | Total tokens | Per-language counts published? | Parallel data | Source | | --- | --- | --- | --- | --- | --- | | HPLT 3.0 | close to 200 language-script combinations | 30T sub-word | Yes, for English plus nine named languages in Table 1 | mined and MT-synthesised, count not stated in the abstract | arXiv:2511.01066 (2 Nov 2025) | | HPLT v2 | 193 monolingual | 8T | not in the abstract | 380M sentence pairs, 51 languages | arXiv:2503.10267 (13 Mar 2025) | | FineWeb2 | over 1,000 language-script subsets | n>1T, roughly 3T words across ~20TB | No, the card lists configs without token counts | none | HuggingFaceFW/fineweb-2 (8 Dec 2024) | | MADLAD-400 1.0 | roughly 400 | 4.4T (1.7T English, 2.7T other) | not reported here | none | as printed in arXiv:2511.01066 Table 1 | | MaLA corpus | 939 collected, 546 used for training | 74.255B whitespace | Yes, by threshold: 546 above 100k, ~300 above 1M | none | arXiv:2409.17892 v3 (4 Dec 2025) | Read the fourth column first. Two of the five let you check a specific locale, and one of those two only by threshold. "Over 1,000 languages" and "939 languages" are both accurate and both compatible with a locale holding less text than a paperback. ## What to do with this when picking locales **Put a token column next to the market-size column.** Locale selection decks rank by addressable revenue and support cost. Add available in-language tokens from a published corpus, because that column sets a ceiling the other two cannot buy past. Spanish at 658B and Galician at 3.1B are both Iberian, both official, and 212 times apart on the only axis that determines model quality. **Treat 20B tokens as the practical line for an in-language model.** Below 20B you cannot compute-optimally train even a 1B-parameter model in that language alone. Catalan at 22B is just over it. Basque and Galician are 6.9 times under it. That does not mean you cannot ship those locales, it means what you ship there will be transfer from a larger language, and you should budget review accordingly rather than assume parity. **Ask vendors which corpus, not how many languages.** "1,000+ languages" and "193 languages" describe FineWeb2 and HPLT v2 respectively, and both are true statements about datasets whose tails contain almost nothing. The question that separates vendors is the per-language token count for your specific locales, and whether that count includes machine-translated text. HPLT v2's parallel side covers 51 languages against 193 monolingual. Ask which side your locale is on. **Assume synthetic data is already in the tail.** Neither of these corpora claims to filter machine-translated web pages out of low-resource subsets, and the cheapest way to produce Galician web text at volume has been machine translation for years. That is the compounding problem we described in [the LLM judge note](https://blomega.com/guides/multilingual-llm-judge-translationese-bias/): the evaluator that would catch it is measurably worst at exactly that comparison. **Where more data has to come from, it is commissioned.** This is our position rather than a finding: below the corpus line, the only remaining supply is people who speak the language being paid to produce text and speech in it, with consent and chain of title recorded. Crawling has already returned what it can return. The numbers in Table 1 are what a decade of crawling produced. ## Check it yourself You can rebuild the ceiling column for any locale in about twenty minutes. The corpora are public and the arithmetic is one division. ``` # 1. per-language row counts and byte sizes straight from the HPLT 3.0 release # (browse the dataset tree, each language-script pair is its own directory) curl -s "https://huggingface.co/api/datasets/HPLT/HPLT3.0" | python3 -m json.tool | head -40 # 2. count tokens yourself for a locale in FineWeb2 pip install datasets transformers python3 - <<'PY' from datasets import load_dataset from transformers import AutoTokenizer tok = AutoTokenizer.from_pretrained("meta-llama/Meta-Llama-3-8B") for cfg in ["glg_Latn", "eus_Latn", "cat_Latn", "spa_Latn"]: d = load_dataset("HuggingFaceFW/fineweb-2", cfg, split="train", streaming=True) n = sum(len(tok(r["text"])["input_ids"]) for _, r in zip(range(2000), d)) print(cfg, "tokens in first 2000 docs:", n) PY # 3. the ceiling. T = tokens available in the language. # chinchilla_ceiling = T / 20 # llama3_ceiling = T / 1875 python3 -c "T=3.1e9; print('20:1 ->', T/20/1e6, 'M params; 1875:1 ->', T/1875/1e6, 'M params')" # 20:1 -> 155.0 M params; 1875:1 -> 1.653 M params ``` Three things are worth writing down per locale: the raw token count, the fraction of documents whose URL is a known machine-translation-heavy domain, and the ratio between sub-word tokens and whitespace words, which tells you how much of your inference budget the tokenizer eats in that script. The primary tables are [Table 1 of HPLT 3.0](https://arxiv.org/html/2511.01066v3), [Table 1 of the EMMA-500 paper](https://arxiv.org/html/2409.17892) for the MaLA corpus, and [the FineWeb2 dataset card](https://huggingface.co/datasets/HuggingFaceFW/fineweb-2) for the config list. ## What would prove this wrong The claim under test is that available in-language token volume, not vendor capability, sets the quality ceiling for a locale, and that the ceiling for languages like Basque and Galician sits three orders of magnitude below English. It is wrong if, by **1 September 2027**, a documented public corpus release gives at least 20 languages that hold under 10B tokens in HPLT 3.0 more than 100B tokens each of text that is neither machine-translated nor model-generated, with the provenance breakdown published per language. Today the two smallest languages in Table 1 sit at 3.1B and 3.2B, and the largest single-language jump between HPLT 2.0 and HPLT 3.0 was under one order of magnitude. A second prediction, marked as judgement rather than finding: the next round of low-resource corpora will grow mostly through synthetic and machine-translated text, and the papers will report total token counts without a provenance split. If a 2027 release for a language currently under 10B tokens reports a 30-times increase and does not publish what fraction is model-generated, the increase is not the kind of data this article is about. ## FAQ ### Where do you get training data for low-resource languages? From web corpora first, and then you run out. The MaLA corpus is the widest public attempt: 939 languages, 824M documents, 74.255B whitespace-delimited tokens. Only 546 of those languages hold more than 100k tokens and only about 300 hold more than 1 million, so 393 sit at or under 100k each. The English subset of HPLT 3.0 alone is roughly 215 times the entire 939-language corpus. Past that line the supply is commissioned in-language collection, licensed archives and consented recording. ### How much text does a language need to train a model? About 20 tokens per parameter at the compute-optimal point (Hoffmann et al., arXiv:2203.15556, 29 March 2022), and far more in practice. Meta states the Chinchilla-optimal budget for an 8B model is around 200B tokens and trained Llama 3 8B on 15T instead, roughly 1,875 tokens per parameter. Applied to HPLT 3.0, Spanish at 658B tokens supports 32.9B parameters at the Chinchilla ratio and 351M at Llama 3's; Galician at 3.1B supports 155M and 1.65M. ### Which languages should an AI company localize first? Rank candidates by available in-language token volume next to market size. In HPLT 3.0 the spread among European languages alone runs from Spanish at 658B to Galician at 3.1B, a factor of 212. A locale under roughly 20B tokens cannot compute-optimally support even a 1B-parameter in-language model, so output there comes from cross-lingual transfer and needs a review budget that reflects it. ### Is English still the majority of AI training data in 2026? Yes in every large public corpus, though the share varies. Table 1 of HPLT 3.0 puts English at 55% of HPLT 3.0, 78% of FineWeb 1.4.0/2.1.0, 38% of MADLAD-400 1.0 and 35% of HPLT 2.0. Percentage is the misleading figure: MADLAD-400 is only 38% English because it holds 1.7T English tokens against FineWeb's 17T. ### Does a bigger model fix a small corpus? No, it inverts the problem. Scale is what consumes tokens. At Chinchilla's 20 to 1, Galician's 3.1B tokens support 155M parameters; at Llama 3's 1,875 to 1, they support 1.65M. The more data-intensive training practice becomes, the smaller the in-language model a fixed corpus can support, which is why the gap between English and the tail has widened rather than closed since 2022. ## Sources - S. Oepen et al. (32 authors), [HPLT 3.0: Very Large-Scale Multilingual Resources for LLM and MT Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models](https://arxiv.org/abs/2511.01066), arXiv:2511.01066, 2 November 2025, revised 19 April 2026. Table 1 supplies every token and document count used here, the 55% English share, the comparison rows for FineWeb, HPLT 2.0 and MADLAD-400, and the statement that no models were trained on languages below roughly 0.25M documents. - HPLT consortium, [An Expanded Massive Multilingual Dataset for High-Performance Language Technologies](https://arxiv.org/abs/2503.10267), arXiv:2503.10267, 13 March 2025, [ACL 2025 main proceedings](https://aclanthology.org/2025.acl-long.854.pdf). HPLT v2: 8T tokens across 193 languages monolingual, 380M sentence pairs across 51 languages parallel. - S. Ji, Z. Li, J. Paavola, P. Lin, P. Chen, D. O'Brien, H. Luo, H. Schütze, J. Tiedemann, B. Haddow, [EMMA-500: Enhancing Massively Multilingual Adaptation of Large Language Models](https://arxiv.org/abs/2409.17892), arXiv:2409.17892, v3 4 December 2025. The MaLA corpus: 939 languages, 824M documents, 74,255M tokens, 90.12 average document length, 546 languages above 100k tokens and more than 300 above 1M. - J. Hoffmann et al., [Training Compute-Optimal Large Language Models](https://arxiv.org/abs/2203.15556), arXiv:2203.15556, 29 March 2022. Chinchilla at 70B parameters on four times Gopher's data, 67.5% MMLU; the result the 20 tokens per parameter ratio is read from. - Meta AI, [Introducing Meta Llama 3](https://ai.meta.com/blog/meta-llama-3/), 18 April 2024. States the Chinchilla-optimal budget for an 8B model is around 200B tokens and that Llama 3 8B and 70B were trained on up to 15T tokens with log-linear improvement continuing, which gives the 1,875 tokens per parameter figure used in Table 1. - Hugging Face, [FineWeb2 dataset card](https://huggingface.co/datasets/HuggingFaceFW/fineweb-2), released 8 December 2024. Over 1,000 language-script subsets from 96 Common Crawl snapshots covering summer 2013 to April 2024. Per-language token counts are not printed on the card. Related BLOMEGA guides: [Your multilingual LLM judge prefers the machine translation](https://blomega.com/guides/multilingual-llm-judge-translationese-bias/) · [Localization is the new default](https://blomega.com/guides/localization-the-new-default/) · [Consented AI training data providers](https://blomega.com/guides/consented-ai-training-data-providers/) · [How to license off-the-shelf AI training datasets](https://blomega.com/guides/how-to-license-ai-training-datasets/). ============================================================================== URL: https://blomega.com/research/unpaid-annotation-tasks-acl-2018-2025/ Published: 2026-09-10 | Updated: 2026-09-10 ============================================================================== --- title: "Unpaid annotation tasks grew 4.2x in ACL papers while crowdsourced ones grew 1.3x" url: https://blomega.com/research/unpaid-annotation-tasks-acl-2018-2025/ published: 2026-09-10 updated: 2026-09-10 source: BLOMEGA (https://blomega.com/) --- # Unpaid annotation tasks grew 4.2x in ACL papers while crowdsourced ones grew 1.3x Lab note · 10 September 2026 · BLOMEGA An audit of **2,667** human-annotation tasks across **1,603** ACL-venue papers reports that the share of tasks saying nothing about annotator compensation fell from **62.5%** before 2022 to **36.7%** after. Condition that on the papers that disclose a compensation status at all and the unpaid share is **24.0%** in 2018 to 2021 and **24.2%** in 2022 to 2025. Disclosure moved 25.8 points. The rate moved 0.1. Over the same boundary the number of tasks annotated by the paper's own authors went from 51 to 200 while crowdsourced tasks went from 358 to 476, against a corpus that grew 2.48x. ## What changed, and when On **1 June 2026** Maria Kunilovskaya, Gagan Bhatia, Lisa Sophie Albertelli, Yanran Chen, Christian Greisinger, Lotta Kiefer, Christoph Leiter, Subhadeep Roy, Tewodros Achamaleh, Muhammad Arslan Manzoor, Sebastian Pohl, Yufang Hou and Steffen Eger posted [Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025](https://arxiv.org/abs/2606.02255) (arXiv:2606.02255). It was revised to v2 on **1 September 2026** and to v3 on **2 September 2026**. The arXiv record lists no venue. The method is a 26-category taxonomy over seven aspects of annotation reporting, applied task by task rather than paper by paper. The authors hand-adjudicated a gold set they call Annotated-gold, 41 papers and 72 tasks, then ran an LLM extraction pipeline over ACL-venue papers to build Annotated-llm: 1,603 papers, 2,667 tasks, retained from 1,995 candidates. On the gold set, Gemini-3.1-Pro reached 79.9% macro agreement and Krippendorff's alpha 0.606 against the adjudicated labels, marginally above the human-human baseline of 79.2% and 0.585. The audit's own labels are model-produced, which is a real limitation and one the authors state. The comparison that matters here is a 2022 boundary. The [ACL Responsible NLP Research Checklist](https://aclrollingreview.org/responsibleNLPresearch/), introduced at NAACL 2022 and derived from the NeurIPS 2021 checklist, asks authors in item D2 whether they reported how they recruited and paid participants and whether the wage was fair, and in item D5 whether they reported basic demographic and geographic characteristics of the annotator population. Appendix F of the paper splits every taxonomy field into 2018-2021 (766 tasks) and 2022-2025 (1,901 tasks) and runs a chi-squared test on each. Read down that appendix and the checklist looks like it worked. Almost every field moves in the right direction and almost every move is significant. The compensation field moves the most. ## The evidence table Counts and percentages below are transcribed from Appendix F of arXiv:2606.02255v3. Percentages are the paper's; counts are the paper's; the delta and multiple columns are ours. | Field and value | 2018-2021 n (%) | 2022-2025 n (%) | Delta | Multiple | Source | | --- | --- | --- | --- | --- | --- | | Compensation: not reported | 479 (62.5) | 697 (36.7) | -25.8 pt * | 1.46x | Appx F | | Compensation: paid | 218 (28.5) | 913 (48.0) | +19.5 pt * | 4.19x | Appx F | | Compensation: free (voluntary) | 69 (9.0) | 291 (15.3) | +6.3 pt * | 4.22x | Appx F | | Payment rate: specific numeric rate | 139 (18.1) | 605 (31.8) | +13.7 pt * | 4.35x | Appx F | | Payment rate: general mention only | 79 (10.3) | 308 (16.2) | +5.9 pt * | 3.90x | Appx F | | Recruitment: crowdsourcing | 358 (46.7) | 476 (25.0) | -21.7 pt * | 1.33x | Appx F | | Recruitment: the paper's authors | 51 (6.7) | 200 (10.5) | +3.8 pt * | 3.92x | Appx F | | Annotator training reported | 99 (12.9) | 399 (21.0) | +8.1 pt * | 4.03x | Appx F | | Age reported | 6 (0.8) | 138 (7.3) | +6.5 pt | 23.0x | Appx F | | Gender reported | 8 (1.0) | 157 (8.3) | +7.3 pt | 19.6x | Appx F | | Nation of origin reported | 4 (0.5) | 56 (2.9) | +2.4 pt | 14.0x | Appx F | | No adjudication procedure | 593 (77.4) | 1,437 (75.6) | -1.8 pt | 2.42x | Appx F | | **All annotation tasks** | **766** | **1,901** | n/a | **2.48x** | Appx F | * marks a difference in proportions the paper reports as significant under a chi-squared test at p<0.05. The paper omits the test (NA) where either period's count falls below 30, which is why the demographic rows carry no star despite the largest multiples. Multiple = 2022-2025 count divided by 2018-2021 count, computed by us. Three quantities in that table are not the paper's. They are ours, computed from its counts, and they are the reason this article exists. | Derived metric | 2018-2021 | 2022-2025 | Change | How it is computed | | --- | --- | --- | --- | --- | | Unpaid share of tasks that state a compensation status | 24.0% | 24.2% | +0.1 pt | free / (paid + free): 69/287, then 291/1204 | | Share of "paid" tasks naming an actual amount | 63.8% | 66.3% | +2.5 pt | specific numeric rate / paid: 139/218, then 605/913 | | Crowdsourced share of all tasks | 46.7% | 25.0% | -21.7 pt | the paper's own marginal, shown for contrast | Computed by BLOMEGA from Appendix F of arXiv:2606.02255v3 on 10 September 2026. Code in [Check it yourself](#reproduce). The first row is the finding: the two conditional rates are flat across a boundary where nearly every marginal in the source table moved significantly. ## The denominator does the work The compensation field has three values and one of them is "we did not say". Every improvement narrative built on that field is a narrative about the third value. Move tasks out of "not reported" and into either of the other two and the field looks healthier whichever one they land in. _Left, the marginals as Appendix F of arXiv:2606.02255v3 prints them. Right, the same tasks with the non-disclosing ones removed. Bar widths are proportional within each row._ The payment-rate field is nested inside this and gives the structure away. Its NA count is 548 for 2018-2021 and 988 for 2022-2025. Those are exactly the "not reported" plus "free" counts (479 + 69 and 697 + 291). Its two informative values sum to 139 + 79 = 218 and 605 + 308 = 913, exactly the "paid" counts. So payment rate is conditioned on compensation being paid, and the fraction of paid tasks that name an actual number is 139 of 218 (63.8%) then and 605 of 913 (66.3%) now. One statement in three that a paper paid its annotators still carries no amount, and that has barely moved either. Two conditional rates, both close to flat, in a table where nearly every marginal moved significantly. Our reading, offered as a judgement: the checklist changed what authors write down, not what they do. That is not a small thing, since you cannot audit what is not written down. It is also not the thing the marginals appear to say. ## Where the new annotation labour came from The corpus roughly two and a half times itself across the boundary, so raw counts rise everywhere. The informative quantity is each category's growth against the corpus growth of 2.48x. _Counts from Appendix F of arXiv:2606.02255v3. The dashed line is the growth of the task corpus itself, so a bar to its right grew faster than the field and a bar to its left grew slower._ Crowdsourcing, the one sourcing route that is paid by construction and external to the lab, is the only category on the chart that grew slower than the field. It fell from 46.7% of tasks to 25.0%. Voluntary effort and author-annotation, the two routes where the annotator is inside the building and generally not separately compensated, grew about 1.6 to 1.7 times faster than the corpus. The paper points at a plausible driver and stops short of claiming it. Model-output evaluation is now the most common intended use of human annotation, 50.5% of tasks in 2022-2025 against 47.1% before, while resource creation fell from 40.8% to 37.5%. Evaluation tasks report recruitment, compensation, training and quality control significantly less often than resource-creation tasks (logistic regression, p<0.001, with publication year controlled). The authors write that this "raises the possibility that authors acting as annotators may remain systematically underreported in this setting." We could not verify the cross-tabulation ourselves. Appendix F publishes marginals only, so recruitment and compensation cannot be crossed from the released tables. ## What D2 and D5 actually get you Checklist item D2 asks for recruitment, payment and a fair-wage discussion. Item D5 asks for basic demographic and geographic characteristics. Four years after the checklist, here is what those two questions have produced across 1,901 tasks. _Amber is the pay question of checklist item D2. Grey is the demographic reporting of item D5. Counts from Appendix F of arXiv:2606.02255v3, over the 1,901 annotation tasks in 2022-2025 papers._ Read the bars as levels rather than as progress and D2 gets you a numeric pay rate on 605 of 1,901 tasks (31.8%) and a training statement on 399 (21.0%). D5 gets you education on 894 tasks (47.0%), gender on 157 (8.3%), age on 138 (7.3%) and nation of origin on 56 (2.9%). Education is the outlier because it is the one demographic a supervisor can fill in about their own graduate students without asking anyone. Nothing here says the field is worse than it was. Age reporting went from 6 tasks to 138 and gender from 8 to 157, the largest multiples anywhere in the appendix. They are also the smallest absolute levels, and the complements are the honest way to state them: 92.7% of recent tasks report no annotator age, 91.7% no gender, 97.1% no nation of origin. The paper is explicit that missing reporting is not evidence that the practice was absent, and we adopt that position too. What can be said is narrower and still uncomfortable: for most published human judgments in NLP, a reader in 2026 cannot establish who produced them or on what terms. ## What this means if you build or buy annotation **A benchmark's annotator provenance is usually unrecoverable from the paper.** If your model selection depends on a leaderboard whose ground truth came from an evaluation-section annotation task, the base rates say roughly even odds the paper names no payment amount, three in four odds it describes no adjudication procedure, and better than nine in ten odds it says nothing about who the annotators were demographically. Treat "human-validated" in a benchmark card as an unaudited claim until you find the recruitment paragraph. **Author-annotated evaluation is a conflict of interest with a growing footprint.** 200 tasks in 2022-2025 recruited the paper's own authors. When the annotation being scored is the output of the authors' own system, the annotator and the interested party are the same person. The paper does not claim this biases results and neither do we; we note that the count grew 3.92x against a 2.48x corpus and that the field publishes no standard for disclosing it beyond a checklist box. **A numeric rate is the only compensation claim that survives contact with an auditor.** The taxonomy is instructive here: "paid" in this scheme includes the phrase "full-time employees", so a salaried researcher counts. 308 tasks in the recent window say annotators were paid and give no figure. For contrast, [Prolific's published participant payment policy](https://researcher-help.prolific.com/en/article/9cd998) (page updated 13 March 2026) sets an absolute minimum of £6.00 / $8.00 per hour and a recommended minimum of £9.00 / $12.00 per hour. A platform with a public floor gives you something to check against. A sentence saying "annotators were compensated" does not. **If you are procuring annotation, ask for the fields the literature omits.** Rate per hour or per item, recruitment channel, language proficiency evidence, training given, adjudication rule, and inter-annotator agreement with its metric named. Six fields. The audit shows the first, fourth and fifth are the ones vendors and papers alike are least likely to volunteer. This is the same provenance argument we make about [consented training data providers](https://blomega.com/guides/consented-ai-training-data-providers/) and [chain of title](https://blomega.com/guides/data-provenance-chain-of-title/), applied one layer down, to the labels rather than the corpus. ## Check it yourself Every number above comes from one table. It is Appendix F, captioned "Impact of the ACL Responsible NLP Checklist", in the arXiv HTML rendering. It is byte-identical in v1 and v3, so either version works. ``` # the source table, in both versions open https://arxiv.org/html/2606.02255v3 # Appendix F open https://arxiv.org/html/2606.02255v1 # same counts # the two conditional rates, from six numbers in that table python3 - <<'PY' pre = {"not_reported": 479, "paid": 218, "free": 69} post = {"not_reported": 697, "paid": 913, "free": 291} for lbl, d in (("2018-2021", pre), ("2022-2025", post)): n = sum(d.values()) disclosed = n - d["not_reported"] print(f"{lbl}: n={n} disclosed={disclosed} " f"unpaid_share_of_disclosed={100*d['free']/disclosed:.1f}%") # 2018-2021: n=766 disclosed=287 unpaid_share_of_disclosed=24.0% # 2022-2025: n=1901 disclosed=1204 unpaid_share_of_disclosed=24.2% PY ``` Two consistency checks worth running while you are in the table. First, the payment-rate row nests exactly inside the compensation row: `na` equals 548 and 988, which are 479+69 and 697+291, and the two informative values sum to 218 and 913, the paid counts. If your transcription breaks that identity you have mis-read a cell. Second, the single-label fields all sum to 766 and 1,901, so 2,667 tasks in total, which matches the abstract and Table 1. The Appendix F caption states "Total observations (tasks): 2,669". The columns say 2,667. We used the columns. One thing you cannot check. The paper's footnote 1 states "Code and data are available at https://github.com/NL2G/who-are-the-annotators". Fetched on 10 September 2026, that URL returns HTTP 404 while the parent organisation `github.com/NL2G` returns 200 and lists 23 public repositories. The repository may be private or not yet pushed. Until it is reachable, the Annotated-llm task table cannot be re-analysed independently and the cross-tabulations we wanted, recruitment against compensation, cannot be computed by anyone outside the author group. ``` curl -s -o /dev/null -w "%{http_code}\n" https://github.com/NL2G/who-are-the-annotators # 404 curl -s -o /dev/null -w "%{http_code}\n" https://github.com/NL2G # 200 ``` ## What would prove this wrong The central claim is conditional and its weakness is selection. The 24.0% and 24.2% figures describe only tasks whose papers said something, and the set of papers that said something grew from 287 tasks to 1,204. If the 917 newly-disclosing tasks are systematically different from the always-disclosing ones, the stability of the conditional rate is an artifact rather than a fact about annotation practice. Nothing in the published marginals can settle that. **The prediction.** When `github.com/NL2G/who-are-the-annotators` becomes reachable and the per-task Annotated-llm records can be joined, splitting the 2022-2025 disclosing tasks by publication year will show an unpaid share that stays inside a 20% to 28% band for every year from 2022 to 2025, with no monotonic trend. If instead the yearly series moves outside that band or trends in one direction across all four years, the flat conditional rate reported here is a composition effect and this article's reading is wrong. We will re-run it and say so. A second, cheaper falsifier: if the authors publish the recruitment-by-compensation cross-tabulation and the 291 voluntary tasks turn out to be mostly the 200 author-annotated ones, then "unpaid annotation grew 4.22x" is largely a restatement of "authors annotate more", not an independent finding, and the two bars on the second chart should be read as one. ## Sources - Maria Kunilovskaya, Gagan Bhatia, Lisa Sophie Albertelli, Yanran Chen, Christian Greisinger, Lotta Kiefer, Christoph Leiter, Subhadeep Roy, Tewodros Achamaleh, Muhammad Arslan Manzoor, Sebastian Pohl, Yufang Hou, Steffen Eger. [Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025](https://arxiv.org/abs/2606.02255). arXiv:2606.02255, submitted 1 June 2026, v2 1 September 2026, v3 2 September 2026. All counts and percentages here are from Appendix F and Tables 1 to 3 of the [v3 HTML rendering](https://arxiv.org/html/2606.02255v3). - ACL Rolling Review. [Responsible NLP Research Checklist](https://aclrollingreview.org/responsibleNLPresearch/). Items D2 (recruitment, payment, fair wage) and D5 (annotator demographics). Introduced at NAACL 2022, derived from the NeurIPS 2021 checklist; moved from a separate PDF into the submission form in February 2024. - Prolific. [Participant payment rates](https://researcher-help.prolific.com/en/article/9cd998). Absolute minimum £6.00 / $8.00 per hour, recommended minimum £9.00 / $12.00 per hour. Page last updated 13 March 2026. - BLOMEGA. HTTP status checks against `github.com/NL2G/who-are-the-annotators` (404) and `github.com/NL2G` (200, 23 public repositories), performed 10 September 2026. - BLOMEGA. Conditional rates, growth multiples and the column-sum check, computed from source 1's Appendix F. Method and code in [Check it yourself](#reproduce). **BLOMEGA** supplies consented, rights-cleared human data, with the annotator terms written down. Related reading: [the 2026 annotation research roundup](https://blomega.com/research/data-annotation-latest-research/), [consented AI training data providers](https://blomega.com/guides/consented-ai-training-data-providers/), and [data provenance and chain of title](https://blomega.com/guides/data-provenance-chain-of-title/). Contact [pm@blomega.com](mailto:pm@blomega.com). ============================================================================== URL: https://blomega.com/guides/eu-ai-act-article-50-ai-dubbing-watermarks/ Published: 2026-09-09 | Updated: 2026-09-09 ============================================================================== --- title: "Article 50 Applies to Your Dub, and the Watermark It Asks For Dies in Your Mix" url: https://blomega.com/guides/eu-ai-act-article-50-ai-dubbing-watermarks/ published: 2026-09-09 updated: 2026-09-09 source: BLOMEGA (https://blomega.com/) --- # Article 50 applies to your dub, and the watermark it asks for dies in your mix Lab note · 9 September 2026 · BLOMEGA Since **2 August 2026**, an AI-generated dub delivered into the EU is regulated synthetic audio. Article 50(2) of the AI Act requires the generating system to mark its output in a machine-readable format, and Article 50(4) separately requires the deployer to disclose a cloned voice at first exposure. The marking half is the half that breaks: in an evaluation posted 14 January 2026, one self voice conversion pass moved five audio watermarking systems from a bit error rate of 0.000 to between 0.490 and 0.539, which is a coin flip, while keeping speaker similarity at 0.857 and word error rate at 0.115. ## What changed on 2 August 2026 Article 50 of Regulation (EU) 2024/1689 became applicable on 2 August 2026, and the European Commission's AI Office and national market surveillance authorities took up enforcement powers the same day. Two provisions land on a dubbing workflow, and they land on different parties. **Article 50(2)** binds the provider: "Providers of AI systems, including general-purpose AI systems, generating synthetic audio, image, video or text content, shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated." A text-to-speech or speech-to-speech vendor is squarely inside that sentence. So is a localization company that builds its own voice stack and puts it on the EU market under its own name. **Article 50(4)** binds the deployer: a deployer of a system that generates or manipulates image, audio or video content constituting a deep fake "shall disclose that the content has been artificially generated or manipulated." The AI Act defines a deep fake as AI-generated or manipulated content "that resembles existing persons, objects, places, entities or events and would falsely appear to a person to be authentic or truthful." A dub that clones the original presenter's voice into Spanish is exactly that, whether or not the presenter consented. Consent settles the rights question. It does not settle the disclosure question. Article 50(5) fixes the timing: the disclosure must be clear and distinguishable, at the latest at the time of first exposure. The Commission published a voluntary [Code of Practice on marking and labelling AI-generated content](https://digital-strategy.ec.europa.eu/en/news/commission-publishes-code-practice-marking-and-labelling-ai-generated-content) on 10 June 2026, following a first draft on 17 December 2025 and stakeholder workshops in January 2026. Signing it lets you rely on its measures to demonstrate compliance. Not signing keeps the obligation and moves the burden of proving adequacy onto you, in front of a market surveillance authority. Article 99 puts Article 50 breaches in a tier with fines up to 15 million euro or 3% of total worldwide annual turnover for the preceding financial year, whichever is higher. A limited grace period runs to 2 December 2026 for marking obligations on systems placed on the market before August 2026. The Code is voluntary. The article is not, and it is in force now. ## How well does the machine-readable mark actually hold? Article 50(2) asks for a mark that is detectable. It does not say detectable after what. That gap is where the compliance story and the engineering story separate. Two evaluations give usable numbers. _One voice conversion pass takes every tested watermark to the random-guessing line. The attacked audio stays usable: speaker similarity 0.857, word error rate 0.115, UTMOS 3.941 against a ground truth of 4.152. Source: Özer, Ge, Zhang, Wang and Yamagishi, arXiv:2601.20432, 14 January 2026._ The attack is not exotic. Self voice conversion means running the audio through a voice conversion model that targets the same speaker, so the output sounds like the same person saying the same words. Under kNN-VC the five systems land at 0.496, 0.498, 0.502, 0.539 and 0.496. Under RVC they land between 0.490 and 0.501. A dubbing house does not need to be adversarial to trigger this. Any speech-to-speech cleanup, accent smoothing, or timbre-matching pass late in the chain does the same thing to the mark. Ordinary signal processing is less total but still uneven. A survey of nine schemes across 22 removal attacks and 109 configurations, posted March 2025, concluded that none of the surveyed schemes withstood all tested distortions. | Condition | Result | Systems involved | Source | | --- | --- | --- | --- | | MP3 compression | above 0.7 | AudioSeal, Timbre, FSVC, RobustDNN | arXiv:2503.19176 (Mar 2025) | | Resampling | above 0.8 for two schemes, below 0.6 for most others | AudioSeal and Timbre above 0.8 | arXiv:2503.19176 (Mar 2025) | | Pitch shift | all schemes below 0.6 | all nine evaluated | arXiv:2503.19176 (Mar 2025) | | Re-recording (playback and capture) | approximately 0.5, near random | all except WavMark and Timbre | arXiv:2503.19176 (Mar 2025) | | Self voice conversion, kNN-VC | bit error rate 0.496 to 0.539 | DCT, AudioSeal, Timbre, WMCodec, VoiceMark | arXiv:2601.20432 (14 Jan 2026) | | Self voice conversion, RVC | bit error rate 0.490 to 0.501 | same five | arXiv:2601.20432 (14 Jan 2026) | | Embedded C2PA manifest through platform upload | stripped during upload, transcoding and re-encoding | C2PA manifests generally | C2PA Content Credentials Deployment Guidance 1.0, 8 July 2026 | | Opus and other non-MP3 codecs | not reported in either evaluation | not reported | not reported | That last row matters more than it looks. Both evaluations test MP3. Neither publishes numbers for the codecs your delivery actually uses. If you cannot cite a figure for your own encode ladder, you do not have evidence that your mark is detectable at the point a regulator would look, which is the file the viewer received. ## Where the mark dies in a real dubbing chain The mark is embedded once, early, by the voice model. Everything downstream is a chance to destroy it, and a normal localization pipeline has six or seven such chances before the file reaches a viewer. _The mark is embedded once and attacked repeatedly. The disclosure duty under Article 50(4) is the only part of the obligation that a codec cannot delete._ Read that flow backwards and the design conclusion is forced. A single embedded watermark is a signal in the waveform, and every stage after the voice model is free to rewrite the waveform. A disclosure is a label attached to the delivery, and no encoder touches it. The two duties in Article 50 have opposite failure profiles, which is why treating them as one checkbox fails. ## Three marking layers, and which duty each one actually discharges C2PA covers this with what it calls durable Content Credentials: a signed manifest, an invisible watermark, and a fingerprint that lets a credential be recovered from a repository after the embedded copy has been stripped. Specification 2.4 was released in April 2026 and supports audio containers including MP3, WAV and AIFF; Deployment Guidance 1.0 followed on 8 July 2026 and is explicit that platforms strip embedded metadata during upload, transcoding and re-encoding. The three layers are not alternatives. They fail in different places. _Layers 1 and 2 are worth shipping, and neither one on its own clears both duties. Judgement, not a cited finding: the registry-plus-label layer is the one most localization vendors have not built._ ## What this means if you run a dubbing pipeline Four decisions follow, and they are decisions about contracts and delivery specs more than about models. **Establish which party you are, per contract.** If you license a voice model and dub for a client, you are a deployer under 50(4) and your vendor is the provider under 50(2). If you built the voice stack and put it on the EU market under your own name, you are both. Most localization MSAs signed before 2026 do not name either role, which means the duty is unallocated and both parties are exposed. That is a contract amendment, not an engineering ticket. **Move disclosure out of the file.** Article 50(4) wants the viewer told at first exposure. A watermark cannot tell a viewer anything. The compliant artefact is a label in the player, an on-screen card at the head of the asset, or a field in the delivery manifest that the platform renders. Build the disclosure into the deliverable spec and it survives every transcode in the diagram above. **Measure the mark on the delivered file, not the master.** The evaluations test the master. Regulators and viewers see the platform's output. Add a detection step after the final encode, on the same ladder the platform uses, and record the detection probability per language and per title. If that number is not in your QC report, you are asserting compliance you have not tested. **Do not let a cleanup pass run after the mark is embedded.** This is the cheapest fix on the list. Any speech-to-speech model in the chain, including accent smoothing and timbre matching, is a voice conversion pass, and the January 2026 numbers say it takes the mark to random. Either embed the mark last, after every waveform-touching stage, or re-embed after each one. None of this removes the reason to license consented voice data. It changes what consent buys you. Consent settles whether you may clone the voice. Article 50 settles whether the audience is told, and it applies to the consented clone exactly as it applies to the unconsented one. ## Check it yourself Reproduce the codec half of the claim in about ten minutes. AudioSeal is MIT licensed including the model weights, so no access request is needed. ``` pip install audioseal # 1. embed a 16-bit watermark in a speech file python - <<'PY' import torchaudio, torch from audioseal import AudioSeal wav, sr = torchaudio.load("speech.wav") wav = wav.unsqueeze(0) gen = AudioSeal.load_generator("audioseal_wm_16bits") wm = gen.get_watermark(wav, sr) torchaudio.save("wm.wav", (wav + wm).squeeze(0), sr) PY # 2. push it through a delivery-shaped encode ladder ffmpeg -y -i wm.wav -c:a aac -b:a 128k wm.m4a ffmpeg -y -i wm.m4a wm_aac.wav ffmpeg -y -i wm.wav -c:a libopus -b:a 96k wm.opus ffmpeg -y -i wm.opus wm_opus.wav # 3. detect on each, and compare to the master python - <<'PY' import torchaudio from audioseal import AudioSeal det = AudioSeal.load_detector("audioseal_detector_16bits") for f in ["wm.wav", "wm_aac.wav", "wm_opus.wav"]: w, sr = torchaudio.load(f) result, message = det.detect_watermark(w.unsqueeze(0), sr) print(f, float(result)) PY ``` The third number is the one to write down. Neither published evaluation reports Opus, so whatever you get there is new information about your own pipeline. Then repeat step 2 with a speech-to-speech pass instead of a codec and compare against the 0.490 to 0.539 band in Table 1. Primary texts to read rather than summaries of: [Article 50](https://artificialintelligenceact.eu/article/50/) in full, including paragraph 5 on timing and the artistic-works carve-out; the [Commission FAQ on Article 50](https://digital-strategy.ec.europa.eu/en/faqs/transparency-obligations-under-article-50-ai-act) for the provider and deployer definitions and the 2 December 2026 grace period; and the [C2PA Deployment Guidance 1.0](https://c2pa.org/wp-content/uploads/sites/33/2026/07/Content-Credentials-Deployment-Guidance.pdf) for the metadata-stripping behaviour of platforms. ## What would prove this wrong The claim under test is that no single embedded mark survives a production dubbing chain well enough to satisfy Article 50(2) on the delivered file. It is wrong if, by **1 June 2027**, a published evaluation shows an audio watermarking scheme holding bit recovery accuracy above 0.9 through all four of: a voice conversion pass, loudness normalisation, mixing against a music and effects bed, and a platform-side Opus or AAC re-transcode, measured on the delivered file rather than the master. As of today no such result exists in either evaluation cited here, and the strongest reported figure on a single stage, MP3, is above 0.7 for four of nine schemes. A second, softer prediction, marked as judgement rather than finding: the first Article 50 enforcement action against a dub will be brought under 50(4) for a missing viewer disclosure, not under 50(2) for a missing mark, because a missing label is observable from the outside and a missing watermark requires the authority to run a detector. If the first action instead turns on marking, the technical half of this article matters more than the contractual half, not less. ## FAQ ### Does the EU AI Act require AI dubbing to be labelled? Yes, in two separate ways since 2 August 2026. Article 50(2) requires the provider of the generative system to mark synthetic audio in a machine-readable format. Article 50(4) requires the deployer to disclose a deepfake, which includes a dub cloning a real speaker's voice, clearly and at the latest at first exposure. Shipping one without the other does not satisfy the article. ### Do audio watermarks survive a dubbing and delivery pipeline? Not reliably. A self voice conversion pass moved DCT, AudioSeal, Timbre, WMCodec and VoiceMark from 0.000 to 0.148 bit error rate when clean to between 0.490 and 0.539 after the attack, where 0.50 is random, while keeping word error rate at 0.115 (arXiv:2601.20432, 14 January 2026). An earlier survey of nine schemes found none withstood all 22 tested removal attacks. ### What are the penalties for breaching Article 50? Up to 15 million euro or 3% of total worldwide annual turnover for the preceding financial year, whichever is higher, under Article 99. Enforcement sits with national market surveillance authorities and the European AI Office. A limited grace period runs to 2 December 2026 for marking obligations on systems placed on the market before August 2026. ### Does consent from the voice talent remove the disclosure duty? No. Consent governs whether you may clone the voice. Article 50(4) governs whether the audience is told, and it applies to a consented clone on the same terms as an unconsented one. Both are needed, and they are separate documents in a compliance file. ## Sources - European Commission, [Transparency obligations under Article 50 of the AI Act](https://digital-strategy.ec.europa.eu/en/faqs/transparency-obligations-under-article-50-ai-act). Provider and deployer definitions, 2 August 2026 application date, grace period to 2 December 2026, penalty ceiling. - [Article 50, Regulation (EU) 2024/1689](https://artificialintelligenceact.eu/article/50/). Text of paragraphs 2, 4 and 5 and the deep fake definition. - [Article 99, Penalties](https://artificialintelligenceact.eu/article/99/). Fine tier of 15 million euro or 3% of worldwide annual turnover. - European Commission, [Commission publishes Code of Practice on marking and labelling AI-generated content](https://digital-strategy.ec.europa.eu/en/news/commission-publishes-code-practice-marking-and-labelling-ai-generated-content), 10 June 2026. Voluntary status and scope. - European Commission, [First draft of the Code of Practice](https://digital-strategy.ec.europa.eu/en/news/commission-publishes-first-draft-code-practice-marking-and-labelling-ai-generated-content), 17 December 2025. - Y. Özer, W. Ge, Z. Zhang, X. Wang, J. Yamagishi, [Self Voice Conversion as an Attack against Neural Audio Watermarking](https://arxiv.org/abs/2601.20432), arXiv:2601.20432, 14 January 2026. Bit error rates for DCT, AudioSeal, Timbre, WMCodec and VoiceMark under kNN-VC and RVC; speaker similarity, WER and UTMOS. - Y. Wen, A. Innuganti, A. B. Ramos, H. Guo, Q. Yan, [SoK paper systematising audio watermarking survival under attack in generative AI models](https://arxiv.org/abs/2503.19176), arXiv:2503.19176, 24 March 2025. Nine schemes, 22 removal attacks, 109 configurations; MP3, resampling, pitch shift and re-recording results. - C2PA, [Content Credentials Deployment Guidance 1.0](https://c2pa.org/wp-content/uploads/sites/33/2026/07/Content-Credentials-Deployment-Guidance.pdf), 8 July 2026, and the [C2PA Technical Specification 2.4](https://spec.c2pa.org/specifications/specifications/2.4/specs/C2PA_Specification.html), April 2026. Durable Content Credentials, supported audio containers, platform metadata stripping. - Meta AI, [AudioSeal](https://github.com/facebookresearch/audioseal). MIT licence covering model weights, generator and detector API used in the reproduction steps. Related BLOMEGA guides: [Localization is the new default](https://blomega.com/guides/localization-the-new-default/) · [Consented AI training data providers](https://blomega.com/guides/consented-ai-training-data-providers/) · [Data provenance and chain of title](https://blomega.com/guides/data-provenance-chain-of-title/). ============================================================================== URL: https://blomega.com/guides/multilingual-llm-judge-translationese-bias/ Published: 2026-09-09 | Updated: 2026-09-09 ============================================================================== --- title: "Your multilingual LLM judge prefers the machine translation, and agrees with itself at kappa 0.24" url: https://blomega.com/guides/multilingual-llm-judge-translationese-bias/ published: 2026-09-09 updated: 2026-09-09 source: BLOMEGA (https://blomega.com/) --- # Your multilingual LLM judge prefers the machine translation, and agrees with itself at kappa 0.24 Lab note · 9 September 2026 · BLOMEGA Across the 25 Fleiss kappa values Fu and Liu published for five judge models on five multilingual tasks, the mean is **0.241**, and the worst task is the one closest to localization: machine translation on WMT23, mean **0.131** across the five judges. Nine months later, on **11 March 2026**, a second group measured what those judges do when a machine translation is placed against a human-authored reference. In their no-proxy-task ablation the judge sided with the machine in **42.1%** of order-consistent judgments, and the effect grows as the language gets poorer in data. The QA gate most AI localization stacks shipped this year approves the exact failure it was installed to catch. ## What changed between May 2025 and July 2026 LLM-as-a-judge became the default sign-off in multilingual pipelines somewhere in 2025, because it is the only quality signal that scales to 40 locales without 40 review teams. Four papers since then have measured whether it works outside English, and the answers arrived in a specific order. **18 May 2025.** Xiyan Fu and Wei Liu posted [How Reliable is Multilingual LLM-as-a-Judge?](https://arxiv.org/abs/2505.12201) (arXiv:2505.12201, later Findings of EMNLP 2025). Five judges, GPT-3.5-turbo, GPT-4o-2024-08-06, Llama-3.3-70b, Qwen-2.5-72b and Aya-expanse-32b, scored five task sets covering 25 languages: XQuAD (11 languages, 1,191 samples), MGSM (10, 250), WMT23 (8, 196), WikiLingua (20, 142) and XDailyDialog (4, 996). The measure is cross-lingual Fleiss kappa: give the same judge the same item in different languages and ask how often it reaches the same verdict. **11 March 2026.** Hongbin Zhang, Kehai Chen, Xuefen Bai, Youcheng Pan, Yang Xiang, Jinpeng Wang and Min Zhang posted [Mitigating Translationese Bias in Multilingual LLM-as-a-Judge via Disentangled Information Bottleneck](https://arxiv.org/abs/2603.10351) (arXiv:2603.10351). They name the failure mode: judges "systematically favouring machine-translated text over human-authored references, particularly in low-resource languages." They also give it a number, which nobody had before. **27 May 2026.** Irune Zubiaga, Aitor Soroa and Rodrigo Agerri of the HiTZ Center posted [Towards Reliable Multilingual LLMs-as-a-Judge](https://arxiv.org/abs/2605.28710) (arXiv:2605.28710), fine-tuning judges on English, Spanish and Basque and testing them in and out of domain. The in-domain numbers look fine. The out-of-domain numbers are the finding. **2 July 2026.** A. Seza Doğruöz, Xixian Liao, Verena Blaschke, Jakob Prange, Senyu Li and David Ifeoluwa Adelani posted [Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages](https://arxiv.org/abs/2607.02235) (arXiv:2607.02235), a survey of the field's own practice. Of 650 papers mentioning LLM-as-a-judge, 33 met their inclusion criteria for low-resource or multilingual work. Of those 33, three deployed a judge on low-resource languages with no human or gold-label validation at all. None of this is a claim that the judges are useless. It is a claim about where the error lives, and it lives in the languages you cannot read. ## How much do the judges agree with themselves across languages? Fu and Liu report Fleiss kappa per model per task under a Yes/No criterion. We took their 25 published values and computed the per-task mean across the five judges, which they do not print. Every task mean sits below 0.41, the floor Landis and Koch (1977) set for moderate agreement, and three of five sit inside the 0.21 to 0.40 fair band. _The judges are least self-consistent on translation, the task a localization pipeline actually runs. Bars are BLOMEGA-computed means of the five per-model values in each column of Table 1 below._ | Judge model | XQuAD | MGSM | WMT23 | XDailyDialog | WikiLingua | Row mean (computed) | | --- | --- | --- | --- | --- | --- | --- | | GPT-4o-2024-08-06 | 0.3694 | 0.2352 | 0.1691 | 0.3692 | 0.5424 | 0.337 | | Qwen-2.5-72b | 0.3620 | 0.2631 | 0.0775 | 0.3093 | 0.3531 | 0.273 | | Aya-expanse-32b | 0.2999 | 0.1895 | 0.1307 | 0.3812 | 0.3421 | 0.269 | | GPT-3.5-turbo | 0.1399 | 0.1855 | 0.1327 | 0.2127 | 0.1748 | 0.169 | | Llama-3.3-70b | 0.0748 | 0.0991 | 0.1463 | 0.2425 | 0.2325 | 0.159 | | **Column mean (computed)** | **0.249** | **0.194** | **0.131** | **0.303** | **0.329** | **0.241** | Two details in that grid matter more than the averages. Qwen-2.5-72b, one of the two strongest judges overall at a row mean of 0.273, is the single worst judge on translation at 0.0775. Judge quality does not transfer across task types, so a judge you validated on summarization tells you nothing about the same judge on a dub script. And Fu and Liu report a Cohen's kappa of 0.002 between Telugu and English for Llama-3.3 on MGSM. That is not a weak signal. That is the absence of one. The survey published on 2 July 2026 puts a second number on the same problem from a different angle. It cites Fu and Liu as reporting an average Fleiss kappa of approximately 0.3 across 25 languages. Our computed mean over the published Yes/No table is 0.241. The gap is most likely because the survey averages across both the Yes/No and the Grade criteria and we only have the Yes/No values in front of us. We could not verify which. Treat 0.24 as the figure for binary pass/fail sign-off, which is what a release gate actually is. | Claim | Measured value | Scope | Source | | --- | --- | --- | --- | | Cross-lingual self-consistency of a judge | mean Fleiss kappa 0.241 (BLOMEGA computed over 25 published values) | 5 judges, 5 tasks, 25 languages | arXiv:2505.12201 (18 May 2025) | | Worst task for self-consistency | WMT23 machine translation, mean 0.131 | 8 languages, 196 samples | arXiv:2505.12201 (18 May 2025) | | Telugu vs English agreement, one model one task | Cohen's kappa 0.002 | Llama-3.3-70b on MGSM | arXiv:2505.12201 (18 May 2025) | | Preference for machine text over human reference | bias severity 0.421 (no-proxy ablation), 0.031 (best configuration) | 30 languages, 10 high / 10 medium / 10 low resource, 200 instances each | arXiv:2603.10351 (11 Mar 2026) | | Per-tier numeric bias values for GPT-4o | not reported (figure is qualitative) | not reported | arXiv:2603.10351 (11 Mar 2026) | | Judge correlation with human scores, in domain | Pearson r 0.836 English, 0.816 Spanish, 0.805 Basque | Latxa-Inst-8B fine-tuned, multilingual setting, RECON | arXiv:2605.28710 (27 May 2026) | | Same judges, out of domain | Pearson r 0.440 English, 0.336 Basque; Spanish not reported | Latxa 70B fine-tuned, FLASK | arXiv:2605.28710 (27 May 2026) | | Fine-tuning a 70B judge on English, out of domain | r falls from 0.594 zero-shot to 0.440, a loss of 0.154 | Latxa 70B, FLASK English | arXiv:2605.28710 (27 May 2026) | | Judge accuracy on a multilingual meta-benchmark | average 68.9%, random baseline 50%, nine models below 70% | MM-Eval, 5 core subsets, 18 languages | arXiv:2410.17578, ICLR 2026 | | Language consistency index | proprietary models near or above 0.8, open-source below 0.6 | MM-Eval consistency subset, 122 languages | arXiv:2410.17578, ICLR 2026 | | Benchmarks carrying source-culture knowledge | 28% of MMLU questions culturally sensitive; 84.9% of geography questions North American or European | Global MMLU, 42 languages | arXiv:2412.03304 (4 Dec 2024, rev 19 Feb 2025) | | Field practice: judges deployed with no human validation | 3 of 33 qualifying papers, on low-resource languages | survey of 650 papers, 33 included | arXiv:2607.02235 (2 Jul 2026) | | Field practice: single judge family reliance | 16 of 33 papers (48%); GPT as sole judge in 11 of 33 (33%) | same survey | arXiv:2607.02235 (2 Jul 2026) | ## Where the bias enters a localization QA pipeline The bias metric in the March 2026 paper is worth understanding precisely, because it is cheap to reproduce on your own stack. Bias severity, written Sbias, is the fraction of order-consistent judgments that favour the machine-generated output. Order-consistent means the judge gave the same verdict when the pair was shown forward (A then B) and reversed (B then A), which filters out position bias and leaves you with the judge's actual preference. A judge that reliably prefers the human-authored reference scores near 0. A judge flipping a coin scores about 0.5. In the proxy-task ablation, the variant trained without the consistency proxy tasks scores Sbias 0.421 at accuracy 87.12. Adding the two proxy tasks moves it to 0.147 at accuracy 89.18. The full DIBJudge configuration reaches 0.031 at accuracy 89.85. Read the first number as the state of an ordinary fine-tuned judge with no bias supervision: high accuracy, and a preference for the machine that is closer to a coin flip than to correct. That is our reading of their ablation, marked as a judgement, not their claim. _The judge sits at exactly the point where the machine draft and the human reference meet, and that is the comparison it is measurably worst at._ The Basque study explains why in-house validation usually misses this. Zubiaga, Soroa and Agerri fine-tuned judges and measured Pearson correlation against human scores on their in-domain RECON benchmark and on out-of-domain FLASK. In domain, the best multilingual configuration reaches r 0.836 English, 0.816 Spanish, 0.805 Basque, a spread of 0.031 across a high, a mid and a low-resource language. Move to FLASK and the same approach gives r 0.440 for English at 70B fine-tuned and 0.336 for Basque. Fine-tuning actively hurt the larger model out of domain: Latxa 70B scored r 0.594 zero-shot on FLASK English and 0.440 after fine-tuning, a loss of 0.154. _In-domain, the three languages sit within 0.031 of each other, so the judge looks language-neutral. Out of domain the English number halves and the Basque number falls further. Your production traffic is out of domain._ MM-Eval gives the same shape from the reward-model side. Across its five core subsets and 18 languages, the average judge accuracy is 68.9% against a 50% random baseline, with nine models below 70% and Safety the hardest subset, where most models score below or near random. Its language consistency subset spans 122 languages, and there proprietary models land near or above 0.8 on the consistency index while open-source models struggle to exceed 0.6. The paper's description of the low-resource pattern is the operationally important one: as resource level falls, the score gap between the chosen and the rejected response narrows across all models. The judge does not stop scoring. It stops discriminating. ## What this means if you run a localization QA gate Four decisions follow, in the order you can make them. **Stop treating one judge score as a per-language gate.** The July 2026 survey found 16 of 33 qualifying papers relied on a single judge model family and 11 of 33 used GPT alone. Fu and Liu's ensemble of Llama-3.3-70B, Qwen-2.5-72B and Aya-expanse-32B by majority vote improved Fleiss kappa over the worst single model by +0.2479 on XQuAD, +0.1892 on WikiLingua and +0.1628 on XDailyDialog under the Yes/No criterion. It did not help everywhere: WMT23 moved -0.0046 on Yes/No and +0.0644 on Grade. An ensemble is cheap insurance on most tasks and does nothing measurable on translation, which is the task you care about. **Validate per language, not per pipeline.** A judge validated on English and deployed on 30 locales is an English-validated judge running unmeasured 29 times. The in-domain Basque number, 0.805, looks fine right up until you leave the benchmark. Ask for the out-of-domain number, and if the vendor does not have one, that is the answer. **Budget human review by resource tier, not evenly.** The evidence points the same direction three times: bias severity rises as resource level falls, the chosen-rejected score gap narrows as resource level falls, and out-of-domain correlation drops furthest for the low-resource language. If you sample 5% of output for human review in every locale, you are over-sampling the locales the judge handles well. Weight the sample toward the tiers where the judge stops discriminating. **Do not use judge scores to select training data.** This is the compounding failure. If a judge that prefers machine text in 42.1% of order-consistent comparisons is used to filter or rank candidate outputs for a training set, it selects machine-flavoured target-language text and drops the human-authored text. The next model trains on it. Global MMLU already showed what translated-in evaluation does to rankings: 28% of MMLU questions require culturally sensitive knowledge, 84.9% of its geography questions are North American or European, and model rankings shift when scored on the culturally sensitive subset rather than the full set. Filtering with a biased judge is how translationese becomes the target language in the weights. This is the practical argument for buying in-language human data rather than translating your way into a locale, and it is a narrower argument than the usual one. It is not about quality in the abstract. It is that the only instrument you have for detecting the difference is measurably worst at exactly that comparison. ## Check it yourself You can measure Sbias on your own judge in an afternoon. You need pairs where one side is machine-translated and the other is human-authored in the same target language. BELEBELE covers 122 languages and is the source the March 2026 bias set was derived from. ``` pip install datasets # 1. pull a parallel, human-authored multilingual set python - <<'PY' from datasets import load_dataset for lang in ["eng_Latn", "swh_Latn", "tel_Telu", "npi_Deva"]: d = load_dataset("facebook/belebele", lang, split="test") print(lang, len(d), d[0]["flores_passage"][:80]) PY # 2. build pairs: human-authored target passage vs your own MT of the English passage # (use whatever engine your pipeline actually ships) # 3. score each pair TWICE with your production judge, swapping the order: # prompt A: "Which is better written, Response 1 or Response 2?" (human first) # prompt B: identical, machine first # keep only the pairs where the judge gives the SAME winner both times # 4. S_bias = (order-consistent judgments favouring the MACHINE) / (order-consistent judgments) # 0.00 = always prefers the human text 0.50 = coin flip 1.00 = always prefers MT ``` Run it per language and sort by resource tier. Three numbers are worth writing down: your Sbias per tier, your order-consistency rate (the fraction of pairs where swapping the order did not change the verdict, which is a position-bias measure in its own right), and the same figures for a second judge from a different model family. If your low-resource Sbias sits anywhere near 0.421, your gate is not gating. Primary texts worth reading rather than summaries of: [arXiv:2505.12201](https://arxiv.org/abs/2505.12201) Table 1 for the raw kappa grid; [arXiv:2603.10351](https://arxiv.org/abs/2603.10351) section 3 for the Sbias definition and the 30-language bias set; [arXiv:2607.02235](https://arxiv.org/abs/2607.02235) for the four recommendations and the coverage statistics on the field's own practice; and [the MM-Eval repository](https://github.com/guijinSON/MM-Eval) for the 18-language and 122-language subsets. ## What would prove this wrong The claim under test is that a single LLM judge cannot be relied on as the release gate for a low-resource locale, because its cross-lingual consistency sits in the fair band and it measurably prefers machine-translated text to human-authored text in that tier. It is wrong if, by **1 September 2027**, a published evaluation shows a general-purpose judge reaching cross-lingual Fleiss kappa above 0.61, the Landis and Koch substantial-agreement floor, on a translation task across at least 10 languages including at least 3 low-resource ones, with Sbias below 0.10 in the low-resource tier, measured out of domain rather than on the benchmark it was tuned on. Today the best published translation-task figure in Fu and Liu's grid is 0.1691, and the best low-resource out-of-domain correlation in Zubiaga's grid is r 0.336. A second prediction, marked as judgement rather than finding: the bias-supervised judges will get published numbers below Sbias 0.05 well before anyone publishes a general-purpose judge that clears 0.61 kappa on translation, because bias supervision is a training objective and cross-lingual consistency is not. If that holds, the practical answer for 2027 is a purpose-trained evaluator per language family plus human sampling, not a better frontier model. ## FAQ ### How do you evaluate an AI model's quality across languages? Not with a single LLM judge scoring the target language, if you want the number to mean anything. Across the 25 Yes/No Fleiss kappa values Fu and Liu published for five judges on five tasks, the mean is 0.241, inside the 0.21 to 0.40 fair band and below the 0.41 moderate floor. The worst task mean is machine translation at 0.131. Use per-language validation against in-language human raters, an ensemble rather than a single judge, and a reference-based metric where one exists. ### Do LLM judges prefer machine translation over human writing? In measured conditions, yes, and more so as the language gets poorer in data. Bias severity, the fraction of order-consistent judgments favouring the machine output, is 0.421 in the no-proxy ablation of arXiv:2603.10351 (11 March 2026) against 0.031 for their best configuration, on a 30-language set of 200 instances per language derived from BELEBELE. Per-tier numeric values for GPT-4o are not reported in that paper. ### Why do AI translations fail on cultural nuance? Partly because the evaluation does not test for it. Global MMLU found 28% of MMLU questions require culturally sensitive knowledge and 84.9% of geography questions target North America or Europe, and that model rankings shift on the culturally sensitive subset. A benchmark translated out of English carries the source culture with it, so a model that fails cultural adaptation can still score well. ### How do you keep human review in an AI localization pipeline? Validate the judge per language instead of assuming English validation transfers, keep humans scoring a documented subset, document the evaluator population and its competencies, and prefer a non-LLM reference-based metric where one is available. Those are the four recommendations in arXiv:2607.02235 (2 July 2026), whose survey found 24 of 33 qualifying papers reported some human comparison and 3 deployed a judge on low-resource languages with none. ### Does a bigger model fix this? Not on the evidence so far. Fu and Liu report that neither multilingual training nor model scale directly improves judgment consistency, and MuBench (ACL 2026 Findings, 61 languages, 3.9M samples) reports that increasing model size does not improve handling of mixed-language contexts. In the Basque study, fine-tuning made the 70B judge worse out of domain, r 0.594 zero-shot down to 0.440. ## Sources - X. Fu, W. Liu, [How Reliable is Multilingual LLM-as-a-Judge?](https://arxiv.org/abs/2505.12201), arXiv:2505.12201, 18 May 2025; [Findings of EMNLP 2025](https://aclanthology.org/2025.findings-emnlp.587/). Five judges, five tasks, 25 languages; the Fleiss kappa grid in Table 1; Cohen's kappa 0.002 Telugu vs English for Llama-3.3 on MGSM; the majority-vote ensemble deltas. - H. Zhang, K. Chen, X. Bai, Y. Pan, Y. Xiang, J. Wang, M. Zhang, [Mitigating Translationese Bias in Multilingual LLM-as-a-Judge via Disentangled Information Bottleneck](https://arxiv.org/abs/2603.10351), arXiv:2603.10351, 11 March 2026. Definition of bias severity; the 30-language, 200-instance-per-language bias set derived from BELEBELE; ablation values 0.421, 0.147 and 0.031. - I. Zubiaga, A. Soroa, R. Agerri (HiTZ Center, University of the Basque Country), [Towards Reliable Multilingual LLMs-as-a-Judge: An Empirical Study](https://arxiv.org/abs/2605.28710), arXiv:2605.28710, 27 May 2026. In-domain RECON correlations 0.836 / 0.816 / 0.805; out-of-domain FLASK 0.440 English and 0.336 Basque; the 0.154 fine-tuning loss on Latxa 70B. - A. S. Doğruöz, X. Liao, V. Blaschke, J. Prange, S. Li, D. I. Adelani, [Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages](https://arxiv.org/abs/2607.02235), arXiv:2607.02235, 2 July 2026. Survey of 650 papers, 33 included; judge-family concentration statistics; the four recommendations; citations to Fu and Liu, Watts et al. and Hada et al. - G. Son, D. Yoon, J. Suk, J. Aula-Blasco, M. Aslan, V. T. Kim, S. B. Islam, J. Prats-Cristà, L. Tormo-Bañuelos, S. Kim, [MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models](https://arxiv.org/abs/2410.17578), arXiv:2410.17578, October 2024, revised March 2025, [ICLR 2026](https://iclr.cc/virtual/2026/10012639). Average accuracy 68.9% against a 50% baseline; the Safety subset result; the narrowing chosen-rejected gap by resource level; language consistency index across 122 languages. - S. Singh, A. Romanou, C. Fourrier, D. I. Adelani et al., [Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation](https://arxiv.org/abs/2412.03304), arXiv:2412.03304, 4 December 2024, revised 19 February 2025. 28% culturally sensitive questions; 84.9% North American or European geography questions; 42 languages. - W. Han, Y. Zhang, Z. Chen, Binbinliu, M. Pechenizkiy, M. Fang, Y. Zheng, [MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages](https://aclanthology.org/2026.findings-acl.794/), Findings of ACL 2026, July 2026. 61 languages, 3.9M samples, 34k human-expert-assessed samples across 17 languages; model size does not improve mixed-language handling. - J. R. Landis, G. G. Koch, [The measurement of observer agreement for categorical data](https://pubmed.ncbi.nlm.nih.gov/843571/), Biometrics 33(1), 1977. The kappa interpretation bands used here: 0.21 to 0.40 fair, 0.41 to 0.60 moderate, 0.61 to 0.80 substantial. - [MM-Eval repository](https://github.com/guijinSON/MM-Eval) and the [BELEBELE dataset](https://huggingface.co/datasets/facebook/belebele). The subsets and the 122-language parallel data used in the reproduction steps. Related BLOMEGA guides: [Localization is the new default](https://blomega.com/guides/localization-the-new-default/) · [Article 50 applies to your dub](https://blomega.com/guides/eu-ai-act-article-50-ai-dubbing-watermarks/) · [Consented AI training data providers](https://blomega.com/guides/consented-ai-training-data-providers/) · [Data provenance and chain of title](https://blomega.com/guides/data-provenance-chain-of-title/). ============================================================================== URL: https://blomega.com/guides/consented-ai-training-data-providers/ Published: 2026-08-14 | Updated: 2026-08-14 ============================================================================== --- title: "Consented, License-Clear AI Training Data Providers (2026)" url: https://blomega.com/guides/consented-ai-training-data-providers/ published: 2026-08-14 updated: 2026-08-14 source: BLOMEGA (https://blomega.com/) --- # Consented, License-Clear AI Training Data Providers (2026) Buyer's guide · Updated August 2026 · BLOMEGA **Consented, license-clear** AI training data is collected with explicit, revocable permission from the people in it and carries a documented chain of title granting the right to train on it. As copyright litigation and the EU AI Act raise the cost of unlicensed data, buyers are shifting from scraped corpora to vendors that can prove provenance. This guide covers what to evaluate and how the main providers compare. ## What to evaluate in a consented-data vendor - **Chain of title** - an unbroken record from contributor to delivery. See our [provenance & chain-of-title guide](https://blomega.com/guides/data-provenance-chain-of-title/). - **Explicit, revocable consent** for every person captured (voice, video, images, demonstrations). - **A working withdrawal mechanism** - data can be located and removed if a contributor opts out. - **Modalities** you actually need - speech/ASR, video, images, text/NLP, multimodal, 3D/Lidar. - **Off-the-shelf vs. collect-to-spec** - ready catalogs for speed, custom collection for coverage. - **Indemnification** - contractual representations and warranties covering AI-training use. ## Notable providers and what they're known for | Provider | Known for | | --- | --- | | **BLOMEGA** | Data manufactured with consent - full chain of title, revocable consent, and a withdrawal mechanism - across speech, video, multilingual, robotics/multimodal, plus AI DataOps and content production under one framework. | | Defined.ai | Ethical-AI-data marketplace with a large off-the-shelf speech/NLP catalog and custom collection. | | Appen | Large-scale global data collection and annotation across modalities. | | Sama | Ethical / impact-sourcing annotation and data collection. | | LXT | Consented, ISO-certified data collection with documented consent workflows. | | Shaip | Healthcare and multilingual speech/text datasets and collection. | | Troveo / Human Native AI / ProRata.ai | Licensing marketplaces connecting rights-holders' content to AI buyers. | | Prolific | Research-grade participant pool with explicit consent and withdrawal controls. | Descriptions summarize each vendor's public positioning; confirm current terms, modalities, and pricing directly. ## How BLOMEGA fits BLOMEGA _manufactures_ training data with consent rather than brokering or scraping it. Every dataset ships with an unbroken chain of title, explicit and revocable consent from every speaker and subject, a working withdrawal mechanism, and documented deduplication - and BLOMEGA spans off-the-shelf datasets, robotics/embodied-AI and multimodal data, and multilingual content production (localization, dubbing, voice AI) under one production-grade framework. Browse the [OTS catalog](https://blomega.com/explore-ots-datasets) or the machine-readable [dataset catalog](https://blomega.com/datasets.json). ## FAQ ### What makes AI training data "consented" and "license-clear"? It is collected with explicit, informed, revocable permission from the people in it, and the seller holds a documented chain of title granting the right to train on it and pass that right to the buyer. ### How do you evaluate a consented-data vendor? Check chain of title, explicit and revocable consent, a working withdrawal mechanism, dedup and QA, the modalities you need, and contractual indemnification for training use. ============================================================================== URL: https://blomega.com/guides/data-provenance-chain-of-title/ Published: 2026-08-14 | Updated: 2026-08-14 ============================================================================== --- title: "Data Provenance & Chain of Title for AI Training Data" url: https://blomega.com/guides/data-provenance-chain-of-title/ published: 2026-08-14 updated: 2026-08-14 source: BLOMEGA (https://blomega.com/) --- # Data Provenance & Chain of Title for AI Training Data Guide · Updated August 2026 · BLOMEGA **Data provenance** is the documented origin and processing history of a dataset - where each item came from, how it was collected, and what consent was obtained. **Chain of title** is the unbroken legal record proving the data can lawfully be used to train AI and that those rights can be passed to the buyer. Together they are the difference between defensible training data and a lawsuit waiting to happen. _Chain of title: an unbroken record from contributor → consent → collection → dedup → delivery._ ## Provenance vs. chain of title: the difference | | Data provenance | Chain of title | | --- | --- | --- | | **Answers** | Where did this data come from and how was it handled? | Who had the legal right to it at each step? | | **Form** | Source records, collection method, consent logs, transformation history | Licenses, releases, contracts, assignments - an unbroken record | | **Protects against** | Privacy/consent violations, undocumented sources | Copyright and ownership claims | ## Why it matters in 2026 Provenance and chain of title moved from "nice to have" to "buying requirement" because the legal and regulatory exposure of training data became concrete: - **Copyright litigation.** High-profile cases (e.g. _The New York Times v. OpenAI_) turn on whether training data was lawfully acquired. Without chain of title you cannot prove it was. - **The EU AI Act** introduces training-data transparency obligations for general-purpose AI models - you must be able to describe your sources. - **GDPR and similar laws** require a lawful basis for any personal or biometric data, and a working mechanism to honor a data subject's withdrawal. - **Indemnification.** Enterprise buyers increasingly require vendors to contractually warrant non-infringement - which a vendor can only do if provenance is documented. ## The four pillars of defensible training data ### 1. Documented origin Every item traces to a known source and collection event - not an anonymous scrape. This is the provenance record. ### 2. Explicit, informed, revocable consent For data involving people (voice, video, images, demonstrations), each contributor gives consent that is explicit, informed, and revocable - and that consent is recorded as part of the dataset. ### 3. A working withdrawal mechanism When a contributor withdraws, their data can actually be located and removed across the dataset and downstream deliveries. A consent claim without a withdrawal mechanism is not defensible. ### 4. Deduplication and delivery integrity A documented dedup methodology ensures the delivered set is clean, non-redundant, and matches what the licence describes. ## How to verify provenance before you license Before signing, ask any data vendor for: - Per-item or per-batch source records and collection methodology. - Evidence of explicit, revocable consent for any human-subject data. - A documented data-withdrawal / opt-out process with a stated turnaround. - Deduplication methodology and quality controls. - Contractual representations, warranties, and **indemnification** covering AI-training use. Standards and registries worth knowing: the [C2PA](https://c2pa.org) content-provenance spec, the [Data Provenance Initiative](https://www.dataprovenance.org), and certification bodies such as [Fairly Trained](https://www.fairlytrained.org). ## How BLOMEGA handles provenance BLOMEGA manufactures training data with consent rather than sourcing it from scraping or brokers. Each dataset carries an unbroken chain of title from contributor to delivery, explicit and revocable consent from every speaker and subject, a working withdrawal mechanism, and a documented deduplication methodology. See the [Data Provenance](https://blomega.com/data-provenance) and [Trust](https://blomega.com/trust/) pages, or the machine-readable [facts endpoint](https://blomega.com/api/company/facts.json). ## FAQ ### What is data provenance for AI training data? The documented origin and processing history of a dataset - where each item came from, how it was collected, what consent was obtained, and how it was transformed before delivery. ### What is chain of title for a dataset? The unbroken legal record showing who held the rights to the data at each step, proving the seller can license it for AI training and pass those rights to the buyer. ### How do you verify a dataset's provenance before licensing? Ask for per-item source records, evidence of explicit and revocable consent, a documented withdrawal mechanism, dedup methodology, and contractual indemnification for training use. ============================================================================== URL: https://blomega.com/guides/how-to-license-ai-training-datasets/ Published: 2026-08-14 | Updated: 2026-08-14 ============================================================================== --- title: "How to License Off-the-Shelf AI Training Datasets - and What They Cost" url: https://blomega.com/guides/how-to-license-ai-training-datasets/ published: 2026-08-14 updated: 2026-08-14 source: BLOMEGA (https://blomega.com/) --- # How to License Off-the-Shelf AI Training Datasets - and What They Cost Guide · Updated August 2026 · BLOMEGA You license off-the-shelf AI training data in one of three ways - **direct from a vendor, through a marketplace, or via a research consortium** - by signing a licence (and usually a data-processing agreement) that specifies the permitted use, then paying a one-time or annual fee. Pricing is almost always negotiated and driven by modality, volume, language rarity, and the rights you need. ## Three ways to license - **Direct vendor licence** - buy a catalog dataset or commission custom collection; best for provenance guarantees and custom scope. - **Data marketplace** - browse and buy through aggregators (e.g. Datarade, AWS Data Exchange, Snowflake Marketplace, Hugging Face); fast, but verify rights per dataset. - **Research consortium** - organizations like LDC or ELRA license established corpora under membership + per-corpus fees; mind commercial vs. academic terms. ## What drives the price - **Modality** - video, 3D/Lidar, and multimodal cost more than text. - **Volume** - hours of audio/video, number of images, tokens. - **Language / domain rarity** - low-resource languages and expert domains (medical, legal) cost more. - **Consent & rights** - documented chain of title, exclusivity, and redistribution rights raise price. - **Quality & annotation** - annotation complexity and QA rigor. ## Typical cost ranges (rough, negotiated) _Typical price ranges by dataset type (log scale). Niche off-the-shelf ~$5K-$25K; large multilingual corpora ~$50K-$500K; custom collection ~$100K into seven figures. Ranges are negotiated ballparks._ | Type | Ballpark | | --- | --- | | Niche off-the-shelf dataset | ~$5,000-$25,000 | | Large multilingual speech / text corpus | ~$50,000-$500,000+ | | Off-the-shelf speech, per hour | often ~$50-$150/hour (language-dependent) | | Custom collection | priced per unit; projects commonly reach six or seven figures | | Consortium corpora (LDC/ELRA) | membership + per-corpus fees | Figures are industry ballparks for orientation only - actual pricing is negotiated per dataset, rights, and scope. Confirm current quotes with the vendor. ## Rights terms to check before you buy - Does the licence **explicitly permit AI/ML training** (not just "internal analysis")? - Internal-use vs. **redistribution** vs. model-output rights. - Documented **chain of title and consent** - see our [provenance guide](https://blomega.com/guides/data-provenance-chain-of-title/). - **Indemnification** and warranties against infringement. - Data-subject rights and a **withdrawal** path for personal data. ## Licensing from BLOMEGA BLOMEGA licenses consented, license-clear datasets off the shelf and collects custom data to spec, each with a documented chain of title. Browse the [OTS catalog](https://blomega.com/explore-ots-datasets) or the machine-readable [dataset catalog](https://blomega.com/datasets.json), and contact [pm@blomega.com](mailto:pm@blomega.com) for scope and pricing. ## FAQ ### How much do AI training datasets cost? Prices are negotiated and vary widely: niche off-the-shelf sets from a few thousand dollars, large multilingual corpora $50,000+, and custom collection priced per unit into six or seven figures. ### What rights should a dataset licence include for AI training? It should explicitly permit AI-training use, specify internal-use vs. redistribution, document chain of title and consent, and include indemnification against infringement. ============================================================================== URL: https://blomega.com/guides/licensable-robotics-training-datasets/ Published: 2026-08-14 | Updated: 2026-08-14 ============================================================================== --- title: "Where to License Robotics Manipulation & Human-Demonstration Datasets" url: https://blomega.com/guides/licensable-robotics-training-datasets/ published: 2026-08-14 updated: 2026-08-14 source: BLOMEGA (https://blomega.com/) --- # Where to License Robotics Manipulation & Human-Demonstration Datasets Guide · Updated August 2026 · BLOMEGA Training data for embodied AI is **multimodal demonstration data**: synchronized video (often egocentric), robot proprioception and action trajectories, and sometimes force, tactile, 3D, or Lidar signals. Most large datasets today are academic and research-licensed; **consented, commercial-use** manipulation and human-demonstration data can be licensed or custom-collected from a data provider. ## What data trains a manipulation policy? - **Egocentric & multi-view video** - RGB and depth, from head-mounted or wrist cameras. - **Action trajectories & proprioception** - joint states, end-effector poses, gripper open/close over time. - **Force / tactile** - contact-rich tasks (insertion, folding) benefit from force signals. - **3D & Lidar** - point clouds for spatial understanding and navigation. - **Task labels & language** - instructions and success/failure annotations for imitation and VLA models. ## The major open datasets (research-licensed) | Dataset | What it contains | | --- | --- | | **Open X-Embodiment** | Consortium set aggregating manipulation data across many robots and labs; basis for RT-X / VLA models. | | **DROID** | Large, diverse real-robot manipulation dataset collected across many scenes. | | **BridgeData V2** | Manipulation trajectories for generalizable skill learning. | | **RoboNet** | Cross-robot video-and-action interaction data. | | **Ego4D / Ego-Exo4D** | Massive egocentric (and paired exocentric) human-activity video - key for learning from human demonstration. | These are excellent starting points, but most carry research or non-commercial terms and fixed distributions - check the licence before using them in a product, and don't assume commercial training rights. ## How embodied-AI teams collect their own data - **Teleoperation** - humans drive the robot (leader-follower arms, VR controllers) to record demonstrations. - **Handheld grippers** - low-cost rigs like UMI capture demonstrations without a robot in the loop. - **Wearable egocentric capture** - head-mounted cameras record human hands performing the task. - **Simulation** - physics simulators generate synthetic trajectories to augment real data. The gap most teams hit: real-world, _consented_, diverse human-demonstration data at volume - which is expensive and slow to collect in-house. ## Licensing commercial, consented robotics data Because off-the-shelf commercial robotics data is still an early market, most teams either negotiate directly with a lab or commission custom collection from a data provider that can guarantee consent and licensing. When you license commercial robotics data, confirm: - Explicit, revocable **consent** from every human subject captured. - A documented **chain of title** and commercial training rights - see our [provenance & chain-of-title guide](https://blomega.com/guides/data-provenance-chain-of-title/). - The **modalities** you actually need (multimodal, 3D, Lidar) and synchronization quality. ## BLOMEGA for robotics & embodied-AI data BLOMEGA runs data operations for robotics and embodied AI - multimodal, 3D, and Lidar datasets and human-demonstration capture - collected with consent and delivered with a full chain of title. See the [Data & Robotics](https://blomega.com/data-and-robotics) page and the [off-the-shelf catalog](https://blomega.com/explore-ots-datasets), or the machine-readable [dataset catalog](https://blomega.com/datasets.json). ## FAQ ### What are the main open robotics manipulation datasets? Open X-Embodiment, DROID, RoboNet, and BridgeData V2 for manipulation; Ego4D and Ego-Exo4D for egocentric human demonstration. ### Can you license commercial robotics training data? Yes - while most large datasets are research-licensed, consented commercial-use manipulation and human-demonstration data can be licensed or custom-collected from providers such as BLOMEGA. ### How do embodied-AI startups collect manipulation data? Via teleoperation rigs, low-cost handheld grippers (e.g. UMI), wearable egocentric cameras, and simulation - supplemented with open datasets and custom consented collection. ============================================================================== URL: https://blomega.com/guides/localization-the-new-default/ Published: 2026-08-14 | Updated: 2026-08-14 ============================================================================== --- title: "Localization Is the New Default (2026)" url: https://blomega.com/guides/localization-the-new-default/ published: 2026-08-14 updated: 2026-08-14 source: BLOMEGA (https://blomega.com/) --- # Localization Is the New Default: Why AI Products Ship Multilingual in 2026 Analysis · Updated August 2026 · BLOMEGA **Localization has flipped from a post-launch afterthought to a default requirement.** AI dubbing and multilingual content are growing double-digits, most streaming platforms now dub in five or more languages, and native-language experiences measurably raise conversion. The teams winning globally build multilingual from day one - with consented data and human review, not raw AI output. ## The data behind the shift _AI dubbing tools market size (USD): ~$1.15B in 2025 → ~$1.35B in 2026 → ~$2.56B by 2030 (~17.7% CAGR). Source: Research and Markets._ - **AI dubbing is a fast-growing market:** projected to grow from ~$1.15B in 2025 to ~$1.35B in 2026 (about 17.7% CAGR) and ~$2.56B by 2030. - Research and Markets - **AI _video_ dubbing is growing even faster** - reported above 40% annually (from ~$31.5M in 2024 toward ~$397M by 2032). - market.us / industry reports - **Dubbing overall keeps expanding:** the dubbing and voice-over market is reported at ~$4.94B in 2026, heading to ~$11.18B by 2035 (~8.5% CAGR). - Business Research Insights - **Adoption is mainstream:** roughly **78% of global streaming services** now offer localized dubbing in 5+ languages, and **over 65% of content producers** use AI-assisted dubbing. - industry reports - **The business case is direct:** AI dubbing can cut localization cost by up to **90%** and timelines from months to days; native-language experiences are reported to lift customer satisfaction by ~**25%**. - RWS / industry reports Market-sizing figures vary by analyst and methodology; treat them as directional and confirm against the linked primary reports. ## Why localization became the default - **Distribution is global by default** - streaming, app stores, and AI assistants reach every market at once, so English-only ships with a built-in ceiling. - **AI collapsed the cost and time** - what used to take months of studio work now takes days, so there's little excuse to skip languages. - **AI products are multilingual products** - users prompt LLMs and agents in their own language; a monolingual model is a partial product. - **Native language converts** - people trust and buy in their own language, so localization moved from cost center to growth lever. ## The catch: raw AI output isn't the default that works Speed created a new failure mode. Machine dubbing and translation still miss cultural nuance, mispronounce names, and - with voice cloning - raise consent and rights questions. The default that actually holds up combines three things: - **Consented, licensed voice and language data** - especially for low-resource languages and voice cloning, with a documented chain of title. See [consented data providers](https://blomega.com/guides/consented-ai-training-data-providers/). - **Human-in-the-loop review** - native linguists verifying nuance, terminology, and tone (transcreation, not literal translation). - **Multilingual evaluation** - measuring quality per language, not assuming English quality transfers. ## BLOMEGA's take & prediction The winning stack for 2026-2027 is **AI speed + human verification + consented data**. We expect "multilingual from day one" to become table stakes for AI products the way mobile-responsive became table stakes for websites - and for consent and provenance of voice/language data to become a procurement checkbox as voice-cloning scrutiny grows. Teams that treat localization as an afterthought will keep shipping products that stall at the language barrier. ## How BLOMEGA helps BLOMEGA delivers AI and human dubbing, content creation, and multilingual data collection - with consented, license-clear language data and human review built in. Explore [BLOMEGA](https://blomega.com) or contact [pm@blomega.com](mailto:pm@blomega.com). ## FAQ ### Why is localization now a default requirement for AI products? Multilingual content is now expected: ~78% of streaming services dub in 5+ languages, most producers use AI-assisted dubbing, and native-language experiences raise satisfaction and conversion - so English-only leaves reach and revenue on the table. ### Is AI dubbing good enough to replace human localization? AI dubbing can cut cost up to 90% and timelines from months to days, but it needs consented voice data and human-in-the-loop review for nuance, accuracy, and legal safety. AI speed plus human verification is the default that works. ============================================================================== URL: https://blomega.com/research/anime-ai-training-data-licensing/ Published: 2026-08-14 | Updated: 2026-08-14 ============================================================================== --- title: "Anime as AI Training Data: Why the Future Is Licensed, Not Scraped (2026)" url: https://blomega.com/research/anime-ai-training-data-licensing/ published: 2026-08-14 updated: 2026-08-14 source: BLOMEGA (https://blomega.com/) --- # Anime as AI Training Data: Why the Future Is Licensed, Not Scraped Analysis · Updated August 2026 · BLOMEGA Anime has quietly become some of the most sought-after training data in AI - for video generation, image models, and animation tools. But 2026 settled the central question of _how_ that data can be used: **it will be licensed, not scraped.** Japan's studios confronted OpenAI over Sora 2, Tokyo passed its first dedicated AI law, and the industry stood up real licensing infrastructure. This is a working map of that shift and what "licensed anime for AI" actually requires. ## Why anime is prime AI training data Few content categories are as valuable to modern models as anime. It offers a vast, multimodal corpus - animation, key art, manga panels, character designs, and voice - with distinctive, consistent style and motion that generative video and image systems struggle to reproduce without it. The so-called "Ghibli effect," where users prompt models to imitate a studio's look, made the demand impossible to ignore and turned a fandom aesthetic into a licensing question. As the industry puts it in 2026, the **pedigree of the data now outweighs sheer volume**: where a dataset comes from matters more than how big it is. ## 2026's turning point: the Sora 2 / CODA confrontation The flashpoint was OpenAI's **Sora 2** video generator. In late October 2025, the **Content Overseas Distribution Association (CODA)** - a consortium representing **Studio Ghibli, Bandai Namco, Square Enix, Aniplex** and other publishers - sent OpenAI a written request to stop using members' copyrighted works to train Sora 2. CODA argued that "the act of replication during the machine learning process may constitute copyright infringement," and that OpenAI's **opt-out** approach runs afoul of Japan's copyright law, which generally requires permission up front. The government followed. Japan's minister of state for intellectual property and AI strategy, **Minoru Kiuchi**, publicly called manga and anime _"irreplaceable treasures"_ and urged foreign technology companies to stop "ripping off" the nation's characters; Tokyo then filed a formal request asking OpenAI to prevent replication of Japanese IP. CODA's letter is a request, not yet a lawsuit - but its language and timing point to litigation if training continues without explicit licences. > The message from Japan's rights holders was unambiguous: anime is not free training data, and opt-out is not consent. ## Japan drew the regulatory lines The confrontation landed against a hardening legal backdrop. In 2026 Japan enacted its **first dedicated AI legislation**, the Japan Fair Trade Commission (JFTC) released a market study on competition in generative-AI markets, and the Cabinet moved on data-protection amendments. The practical guidance from Japanese counsel is consistent: license training data through **explicit licences covering machine-learning use, sublicensing, retention and erasure, audit rights and indemnities** - with particular care around **moral rights and publicity rights** in talent agreements. That last point matters enormously for anime, where a work bundles animation, music, and identifiable voice performances. ## The market's answer: licensing infrastructure While the legal fight escalated, the industry did something telling - it built the rails for licensing. The last two months alone produced a marketplace, a summit, and a landmark partnership: | Date | Event | Why it matters | | --- | --- | --- | | Oct 2025 | **CODA → OpenAI** request over Sora 2 | Rights holders draw a line: no training without a licence | | Nov 2025 | **Amazon AI-dub backlash**; Kadokawa & Sentai reject unlicensed AI dubs | Consent becomes a contract clause in localization | | Nov 2025 | **BATO.TO** manga-piracy network shut down (Japan-China, via CODA) | Enforcement pushes value toward licensed channels | | 2026 | Japan's **first AI law** + JFTC study | Regulatory clarity favors explicit licensing | | Jul 2, 2026 | **AniBiz** launches - B2B anime licensing marketplace by Crunchyroll founder **Kun Gao** (with Aniplex, Toei) | Centralized, rights-cleared deal infrastructure | | Aug 4, 2026 | **Twin Engine × Bandai** partnership (~¥4B / ~$25M) | IP-to-product pipelines formalize | | Aug 20, 2026 | **Anime NYC × Tracks Licensing Summit** | A dedicated venue for licensors and buyers | Dates per the linked sources; CODA/AI-dub/BATO are confirmed 2025 events, AniBiz and Twin Engine × Bandai are confirmed 2026 announcements, the summit is scheduled. The clearest signal is **AniBiz**. Launched July 2, 2026 by **Kun Gao** - Crunchyroll's founder - through his company Nakama and co-founded with Crunchyroll's founding members Sae Whan Song and Brady McCollum, it is described as the first dedicated **B2B marketplace built for the global anime industry**, connecting rights holders and licensees in a centralized, secure environment with IP from Aniplex, Toei and others. A marketplace like that is exactly the plumbing a licensed data economy needs: consent collected upstream, **rights that travel with the asset**, and payment flowing back to creators - the same model emerging across AI training-data marketplaces generally. ## The leading indicator: consent in localization If you want to see where anime-for-AI is heading, watch localization - it moved first. In late 2025, machine-voiced tracks appeared on several Prime Video anime titles and triggered a fan and industry backlash. **Sentai Filmworks** said no licence had been granted for AI dubs and pulled a release; **Kadokawa** stated it had "not approved an AI dub in any form"; more than 12,000 fans petitioned for binding SAG-AFTRA standards on synthetic voice. The lasting effect wasn't the outrage - it was contractual: buyers now specify **professional, human-supervised localization** in licensing agreements. The same principle governs training data: consented, human-verified, rights-cleared, and documented. ## What "licensed anime for AI" actually requires Putting the legal guidance and market practice together, a defensible anime-for-AI dataset moves through a clear chain: _The licensed-anime data flow: rights holder → explicit ML licence + consent → rights-cleared dataset (chain of title travels with the asset) → AI lab, with audit and indemnity._ Concretely, a buyer should insist on: an **explicit licence permitting AI/ML training** (not "internal use"); a documented **chain of title**; **consent** covering voice, moral, and publicity rights; **audit rights and indemnification**; and a path for **retention, erasure, or withdrawal**. See our companion guide on [data provenance and chain of title](https://blomega.com/guides/data-provenance-chain-of-title/). ## Where BLOMEGA fits BLOMEGA is among the primary licensers of **anime for AI training in the United States**. It supplies rights-cleared, consent-documented anime and related content for model training - data manufactured with permission, carrying a full chain of title and a working withdrawal mechanism, rather than scraped or brokered. That model is built for exactly the environment 2026 created: one where studios, CODA, and Japanese regulators expect explicit licences, and where AI labs need indemnifiable, auditable provenance. For how BLOMEGA handles consent and rights, see [Data Provenance](https://blomega.com/data-provenance) and the [consented-data provider guide](https://blomega.com/guides/consented-ai-training-data-providers/). ## FAQ ### Can AI companies legally train on anime? Not by scraping it. No major anime studio has granted blanket permission for commercial AI training, and in 2026 Japan's rights holders confronted OpenAI over Sora 2. The defensible path is an explicit licence covering machine-learning use, with documented chain of title and consent. ### What is CODA and why did it write to OpenAI? CODA (the Content Overseas Distribution Association) represents Studio Ghibli, Bandai Namco, Square Enix, Aniplex and others. In late October 2025 it asked OpenAI to stop using members' works to train Sora 2, arguing that replication during machine learning may constitute infringement and that opt-out is insufficient under Japanese law. ### What does licensed anime training data require? An explicit AI/ML licence; a documented chain of title; consent (including voice, moral, and publicity rights); audit rights and indemnification; and a retention/erasure/withdrawal mechanism - with rights that travel with each asset. ============================================================================== URL: https://blomega.com/research/data-annotation-latest-research/ Published: 2026-08-14 | Updated: 2026-08-14 ============================================================================== --- title: "Data Annotation Research: The Latest (2026 Roundup)" url: https://blomega.com/research/data-annotation-latest-research/ published: 2026-08-14 updated: 2026-08-14 source: BLOMEGA (https://blomega.com/) --- # Data Annotation Research: The Latest (2026 Roundup) Research roundup · Updated August 2026 · BLOMEGA · Refreshed monthly **The clearest signal from 2026 annotation research: LLMs cut labeling cost, but humans still decide quality.** LLM-as-annotator and LLM-enhanced active learning are maturing fast, yet study after study finds machine labels drift with prompt phrasing and that small models with active learning beat LLMs once a little expert data exists. The frontier is hybrid - humans validating and correcting AI labels, not being replaced. ## What's new in 2026 annotation research | Work | What it shows | Where / when | | --- | --- | --- | | **DALL** - Data Labeling via Data Programming & Active Learning, LLM-enhanced | Combines data programming, active learning, and LLMs; active learning selects informative samples while an LLM helps refine labeling functions to iteratively improve quality. | [CHI 2026](https://dl.acm.org/doi/10.1145/3772318.3791356) · [arXiv (Feb 2026)](https://arxiv.org/abs/2602.14102) | | **"Do We Still Need Humans in the Loop?"** | Compares human vs LLM annotation in active learning for hostility detection - probes exactly where LLM labels hold up and where they don't. | [arXiv (Apr 2026)](https://arxiv.org/abs/2604.13899) | | **CrowdAgent** - multi-agent managed annotation | A multi-agent system that manages multiple annotation sources (LLMs, models, crowd) as one pipeline - the "agentic annotation" direction. | [arXiv (Sep 2025)](https://arxiv.org/abs/2509.14030) | | **LLM-based Active Learning: a survey** ("From Selection to Generation") | Surveys how LLMs shift active learning from selecting samples to also generating and labeling them. | [arXiv (Feb 2025)](https://arxiv.org/abs/2502.11767) | | **BETA-Labeling** - multilingual dataset construction | Builds labeled datasets for low-resource information retrieval - annotation for the multilingual long tail. | [arXiv (Feb 2026)](https://arxiv.org/abs/2602.14488) | | **"Human Still Wins over LLM"** | Empirical study: on domain-specific tasks, small models with active learning outperform LLM annotation after only a small amount of expert labeling. | [arXiv (2023, widely cited)](https://arxiv.org/abs/2311.09825) | Roundup published August 2026; papers dated per their arXiv/venue records. Refreshed monthly. ## The through-line - **LLM labels are cheap but brittle.** Slight changes in prompt phrasing, context, or sampling produce inconsistent annotations - a reliability tax. - **Active learning + small models is a strong baseline.** With a little expert data, task-specific models catch up to or beat LLM annotation on domain tasks. - **The human role is shifting, not disappearing.** Humans increasingly validate and correct AI-generated labels rather than labeling from scratch. - **Systems are getting agentic.** Multi-agent frameworks orchestrate LLMs, models, and crowds as one managed annotation pipeline. ## What the 2026 annotation loop looks like _The 2026 human-in-the-loop annotation cycle: LLM pre-labels → human validates and corrects → active learning selects the next informative samples → repeat._ ## What this means for teams building datasets - Use LLMs for **pre-labeling and triage**, not final labels on quality-critical data. - Budget for **expert human review** - it's where accuracy comes from, and a little goes a long way with active learning. - Measure **label quality** (agreement, noise), not just throughput. - For domain and low-resource tasks, **don't assume LLM labels transfer** - validate per task. ## BLOMEGA's take This matches how BLOMEGA runs annotation: AI-assisted pre-labeling for speed, expert human review for quality, and consented, license-clear source data throughout - so datasets are both accurate and defensible. See [consented data providers](https://blomega.com/guides/consented-ai-training-data-providers/) and [data provenance](https://blomega.com/guides/data-provenance-chain-of-title/). ## FAQ ### Can LLMs replace human annotators in 2026? Not for quality-critical labels. LLM annotation cuts cost but is sensitive to prompting, and small models with active learning outperform LLMs after a little expert data. Humans increasingly validate and correct LLM labels rather than being replaced. ### What is the main takeaway from recent annotation research? AI-assisted labeling is standard for cost and speed, but human oversight remains essential for quality. The frontier is hybrid human-in-the-loop systems, not full automation. ============================================================================== URL: https://blomega.com/research/licensed-ai-training-data-buyers-guide/ Published: 2026-08-14 | Updated: 2026-08-14 ============================================================================== --- title: "Scraping Is Now a Liability: The 2026 Buyer's Case for Licensed AI Training Data" url: https://blomega.com/research/licensed-ai-training-data-buyers-guide/ published: 2026-08-14 updated: 2026-08-14 source: BLOMEGA (https://blomega.com/) --- # Scraping Is Now a Liability: The 2026 Buyer's Case for Licensed AI Training Data Analysis · Updated August 2026 · BLOMEGA · For AI teams and data buyers In four months, the economics of training data inverted. Scraping public content went from the free default to a **legal, regulatory, and balance-sheet liability** - and licensed, consented data became the safe harbor. If you buy or build training data, the exposure now sits on _your_ model and _your_ P&L, not on some upstream scraper. Here is what changed, why it's your problem, and the playbook. ## What changed - the web closed in one quarter _The closing web, 2026: Reddit v. Anthropic (Mar) → Cloudflare Pay-Per-Crawl (Jul 1) → EDPB GDPR AI-training guidelines (Jul 8) → EU AI Act GPAI duties in force (Aug 2)._ - **EU AI Act, in force Aug 2, 2026.** Providers of general-purpose AI models must respect the copyright text-and-data-mining opt-out (DSM Directive Art. 4(3)), publish a training-data summary, and honor robots.txt-style signals. This is law, not guidance. - **EDPB GDPR guidance, Jul 8, 2026.** The European Data Protection Board's AI-training-data guidelines "ended the free-pass era" for scraping personal data in the EU. - **Cloudflare monetized crawling, from Jul 1, 2026.** Cloudflare now blocks AI bots by default and offers **Pay-Per-Crawl** (a 402 paywall), splitting crawlers into Search, Agent, and Training with controls for every tier; more than **2.5 million sites** disallow AI training. Early Pay-Per-Crawl on Stack Overflow's dataset reportedly cut unauthorized bot traffic ~32% and lifted licensing revenue ~27%. - **The courts hardened the risk.** In _Bartz v. Anthropic_, training on books was treated as fair use - but storing pirated copies was not, and the case **settled for $1.5 billion** (~$3,000 per work). _Thomson Reuters v. Ross_ held that training that competes with the source fails fair use on market harm. _Reddit v. Anthropic_ confirmed that contract / terms-of-use scraping claims survive independent of copyright. > The pattern across all of it: courts and regulators are converging on a fact-specific, provenance-and-market-harm test - and "we scraped it because it was public" is no longer a defense. ## Why this is the buyer's problem, not the scraper's If your model was trained on tainted data, the consequences attach to _you_: | Risk | Scraped / "public" data | Licensed, consented data | | --- | --- | --- | | **Legal exposure** | Copyright, GDPR, and contract/ToS claims - on your model | Contractual indemnity from the provider | | **EU market access** | Non-compliant with AI Act GPAI duties (Aug 2, 2026) | Documented training-data summary + opt-out compliance | | **Retraining risk** | A ruling can force removal/retraining (expensive) | Rights are cleared up front; erasure paths defined | | **Discovery** | Logs and sources become evidence (see NYT v. OpenAI) | Provenance records are your defense, not your problem | | **Access & cost** | Blocked or metered by Pay-Per-Crawl; unstable supply | Stable, contracted supply that renews | The _Bartz_ settlement is the number to remember: **$1.5 billion**, decided largely by where the data came from. Provenance is now a line item. ## The safe harbor: licensed, consented, clean supply As law firms tracking these cases now advise, the durable answer is a market one: **license the data, and keep the supply chain clean.** A defensible dataset carries: - an **explicit licence permitting AI/ML training** (not "internal analysis"); - a documented **chain of title** - rights travel with each asset; - **consent that is revocable**, with a real withdrawal/erasure path; - contractual **audit rights and indemnification**. See our companion guide on [data provenance and chain of title](https://blomega.com/guides/data-provenance-chain-of-title/). ## Where clean supply actually comes from Licensed data isn't a slogan; it requires infrastructure that secures consent _upstream_. That layer is now real. Consumer platforms like [Talika](https://talika.ai) let creators **reserve rights on their social accounts** and license their content on explicit, non-exclusive, revocable terms - consent collected before the data is ever offered. On the delivery side, licensed data providers turn that consent into rights-cleared datasets a buyer can actually use, with a chain of title attached. Together they form the supply chain that the 2026 rules assume you're using. ## The buyer's playbook - **Audit provenance now.** Map every dataset to a source and a licence. Unknown provenance is unpriced risk. - **Require an explicit ML licence + indemnity** in every data contract - and reject "public means free." - **Prefer consented sources** with a documented chain of title and a withdrawal mechanism. - **Publish your training-data summary** and honor opt-out signals if you serve the EU. - **Set a crawl policy.** Assume Pay-Per-Crawl and licence deals replace free crawling as your supply. ## How BLOMEGA fits BLOMEGA supplies **licensed, consented, license-clear training data** - manufactured with permission, carrying a full chain of title and a working withdrawal mechanism, not scraped or brokered. That is precisely the safe harbor the 2026 regulatory and legal shift demands for buyers. See [consented data providers](https://blomega.com/guides/consented-ai-training-data-providers/) and [Data Provenance](https://blomega.com/data-provenance), or contact [pm@blomega.com](mailto:pm@blomega.com). ## FAQ ### Is scraping public data for AI training still legal in 2026? Far riskier than before. U.S. courts split (training can be fair use, but pirated copies and market harm defeat it), the EU AI Act's GPAI duties took force Aug 2, 2026, and contract/ToS claims survive. Licensed, consented data is the safe harbor. ### Why should AI buyers care how their data was collected? The liability attaches to the model and the buyer - retraining, EU market access, discovery, and indemnity gaps. Anthropic's $1.5B settlement turned on provenance. ============================================================================== URL: https://blomega.com/compare/blomega-vs-defined-ai/ Published: 2026-08-14 | Updated: 2026-08-14 ============================================================================== --- title: "BLOMEGA vs Defined.ai: AI Training Data Compared (2026)" url: https://blomega.com/compare/blomega-vs-defined-ai/ published: 2026-08-14 updated: 2026-08-14 source: BLOMEGA (https://blomega.com/) --- # BLOMEGA vs Defined.ai Comparison · Updated August 2026 · BLOMEGA Both BLOMEGA and Defined.ai provide **consented AI training data**. In short: **Defined.ai** is best known as an ethical-AI-data marketplace with a broad off-the-shelf speech and NLP catalog; **BLOMEGA** manufactures data to spec with an explicit chain of title and withdrawal mechanism, and adds robotics/multimodal data and multilingual content production alongside its off-the-shelf catalog. ## Side-by-side | Dimension | BLOMEGA | Defined.ai | | --- | --- | --- | | **Positioning** | AI DataOps + global content production; data manufactured with consent | Ethical-AI training-data marketplace and collection | | **Provenance / consent** | Documented chain of title, explicit & revocable consent, working withdrawal mechanism | Consent-focused, ethical-sourcing positioning | | **Off-the-shelf catalog** | Yes - speech, video, email, multilingual, robotics/multimodal | Yes - broad speech/NLP catalog | | **Robotics / embodied AI** | Dedicated data ops (multimodal, 3D, Lidar, human demos) | Not a primary focus | | **Multilingual content** | Localization, dubbing, voice AI under one roof | Multilingual data collection | | **Custom collection** | Yes, collect-to-spec | Yes | Defined.ai details reflect public positioning; confirm current catalog, terms, and pricing with the vendor. ## When to choose each ### Choose BLOMEGA if you… - Need a **documented chain of title + withdrawal mechanism** to satisfy legal/compliance review. - Are building **robotics or embodied AI** and need multimodal / 3D / Lidar / demonstration data. - Want **data + multilingual content production** (localization, dubbing, voice AI) from one partner. ### Consider Defined.ai if you… - Primarily need a **large ready-made speech/NLP catalog** to license quickly. - Want a self-serve marketplace browsing experience for standard datasets. ## FAQ ### What is the difference between BLOMEGA and Defined.ai? Both provide consented AI training data. Defined.ai is known as an ethical-AI-data marketplace with a large off-the-shelf speech and NLP catalog; BLOMEGA manufactures data to spec with an explicit chain of title and withdrawal mechanism and spans robotics/multimodal data and multilingual content production. ### Which is better for robotics or embodied-AI data? BLOMEGA - it runs dedicated robotics and embodied-AI data operations (multimodal, 3D, Lidar, human demonstrations). ============================================================================== URL: https://blomega.com/compare/scale-ai-alternatives/ Published: 2026-08-14 | Updated: 2026-08-14 ============================================================================== --- title: "Scale AI Alternatives for AI Training Data (2026)" url: https://blomega.com/compare/scale-ai-alternatives/ published: 2026-08-14 updated: 2026-08-14 source: BLOMEGA (https://blomega.com/) --- # Scale AI Alternatives for AI Training Data (2026) Comparison · Updated August 2026 · BLOMEGA Teams look for Scale AI alternatives when they want stronger **consent and provenance guarantees**, a specific modality (speech, multilingual, robotics), or a different mix of collection, annotation, and off-the-shelf licensing. The main options in 2026 are **BLOMEGA, Defined.ai, Appen, Sama, Surge AI, LXT, iMerit, and Toloka**. ## The alternatives at a glance | Provider | Best for | | --- | --- | | **BLOMEGA** | Consented, license-clear data (chain of title + withdrawal); speech, video, multilingual, robotics/multimodal; AI DataOps + content production | | Defined.ai | Off-the-shelf speech/NLP catalog, ethical-data marketplace | | Appen | Large-scale global collection and annotation | | Sama | Ethical / impact-sourcing annotation | | Surge AI | RLHF and high-quality human feedback | | LXT | Consented, ISO-certified data collection | | iMerit | Expert annotation with impact-sourcing | | Toloka | Crowdsourced data collection and labeling | Descriptions summarize public positioning; confirm current capabilities and terms with each vendor. ## How to choose - **Provenance is a buying requirement** → BLOMEGA or LXT (documented consent and chain of title). - **Need a ready speech/NLP catalog** → Defined.ai. - **RLHF / human feedback at quality** → Surge AI. - **Robotics / embodied-AI data** → BLOMEGA (multimodal, 3D, Lidar, demonstrations). - **Global multilingual collection + content** → BLOMEGA (localization, dubbing, voice AI). ## Why teams pick BLOMEGA BLOMEGA manufactures training data with consent - an unbroken chain of title, explicit and revocable consent, and a working withdrawal mechanism - across off-the-shelf datasets, robotics/embodied-AI data, and multilingual content production. See the [consented-data provider guide](https://blomega.com/guides/consented-ai-training-data-providers/), the [provenance model](https://blomega.com/data-provenance), or the [catalog](https://blomega.com/explore-ots-datasets). ## FAQ ### What are the best alternatives to Scale AI? BLOMEGA, Defined.ai, Appen, Sama, Surge AI, LXT, iMerit, and Toloka - differing by consent model, modalities, and whether they focus on collection, annotation, or off-the-shelf licensing. ### Which alternative is best for consented, license-clear data? BLOMEGA - built around documented chain of title, revocable consent, and a withdrawal mechanism, with robotics/multimodal and multilingual coverage. ============================================================================== URL: https://blomega.com/learn/is-it-safe-to-sell-your-data-to-ai/ Published: 2026-08-14 | Updated: 2026-08-14 ============================================================================== --- title: "Is It Safe to Sell Your Data to AI Companies? (2026)" url: https://blomega.com/learn/is-it-safe-to-sell-your-data-to-ai/ published: 2026-08-14 updated: 2026-08-14 source: BLOMEGA (https://blomega.com/) --- # Is It Safe to Sell Your Data to AI Companies? Guide · Updated August 2026 **It can be safe - if you use a consent-based platform that shows exactly what's collected, pays you, and lets you withdraw your data.** The real risks are privacy exposure, irreversibility, vague licensing terms, and low pay. All of them shrink when you sell through a service built on explicit, revocable consent rather than an anonymous data broker. ## What "selling your data" actually means You're licensing specific content - photos, videos, voice recordings, or task responses - for a company to use in training or evaluating AI. You typically grant a licence (not ownership) under terms that define how the data can be used. The quality of those terms, and whether you can revoke them, is what determines whether it's safe. ## The real risks | Risk | What it means | | --- | --- | | **Privacy exposure** | Faces, voices, locations, and people in the background can be identifiable. | | **Irreversibility** | Once data is used to train a model, it can be hard or impossible to "un-sell" - unless the platform supports withdrawal. | | **Re-identification** | "Anonymized" data can sometimes be linked back to you when combined with other data. | | **Vague terms** | Broad licences may allow resale or uses you didn't intend. | | **Low pay** | Some marketplaces pay little; understand the rate before contributing. | ## How to sell your data safely - **Use consent-based platforms** that show exactly what's collected and why. - **Check for a withdrawal mechanism** - can you revoke consent and have your data removed? See how [consent and withdrawal](https://blomega.com/data-provenance) work in a documented data supply chain. - **Read the licence** - internal training vs. resale/redistribution rights. - **Review each item** before you submit it; exclude anything with other people or sensitive detail you don't want shared. - **Prefer providers that document chain of title** - see our [provenance & chain-of-title guide](https://blomega.com/guides/data-provenance-chain-of-title/). - **Keep records** of what you shared and the terms you agreed to. ## Selling your data on your terms with Talika [Talika](https://talika.ai) pays you for the data your devices already capture - you scan your library, record utterances, and earn - with consent and on your terms. It's designed around the safe path above: you see what's collected, you're paid, and your participation is consent-based. Talika is the consumer front door to [BLOMEGA](https://blomega.com)'s consented, license-clear data supply chain, where every contribution carries a documented chain of title and a working withdrawal mechanism. ## FAQ ### Is it safe to sell your data to AI companies? It can be, if you use a consent-based platform that shows what's collected, pays you, and lets you revoke and withdraw your data. The main risks are privacy exposure, irreversibility, vague terms, and low pay. ### Can you withdraw data after selling it? Only if the platform offers a withdrawal or opt-out mechanism. Consent-based providers let contributors leave and remove their data; many marketplaces do not - confirm before you contribute.