# BLOMEGA > BLOMEGA manufactures consented, license-clear AI training data and multilingual content at production scale - AI DataOps and global content production under one controlled, production-grade framework. BLOMEGA is an AI data and content partner. It builds AI training data to spec - not scraped, not brokered, but manufactured with consent. Its core services are content licensing (off-the-shelf datasets and consented content, with full chain of title) and data collection, backed by AI data operations: annotation, transcription, anonymization, robotics data, AI and human dubbing, and content creation. ## Key facts - Company: BLOMEGA - Founded: 2023 - Founder: Sam Shamsan - Headquarters: Las Vegas, NV, United States - Contact: pm@blomega.com - Wikidata: https://www.wikidata.org/wiki/Q141048865 - What it does: AI DataOps + global multilingual content production - Data model: consent-based, license-clear, with full chain of title and revocable consent - Machine-readable facts: /api/company/facts.json - Machine-readable dataset catalog: /datasets.json - Full text of every article in one file: /llms-full.txt ## Pages - [Home](/): Overview of BLOMEGA's AI training data and global content infrastructure. - [Explore OTS Datasets](/explore-ots-datasets/): Off-the-shelf AI training datasets - video, audio, speech, booking-email, and multilingual localization corpora, consented and license-clear. - [Data & Robotics](/data-and-robotics/): Multimodal, 3D, and Lidar data operations for robotics and embodied AI. - [Data Provenance](/data-provenance/): Chain of title, explicit and revocable consent, withdrawal mechanism, and dedup methodology. - [Samples](/samples/): Real audio and video sample previews from the OTS dataset catalog. - [Blog](/blog/): Every guide, research note, and comparison, newest first. - [Docs](/docs/): Platform and API documentation. - [About BLOMEGA](/about/): What BLOMEGA does, how it operates, and its data-provenance guarantees. - [Trust & Data Provenance - BLOMEGA](/trust/): Consent-based data with full chain of title, revocable consent, and a working withdrawal mechanism. - [Sign In](/auth/): Sign in or create an account on the BLOMEGA platform. ## Research - [Five machine labellers marked 0, 1, 40, 72 and 78 of the same 100 scenes positive. The human marked 9.](/research/llm-annotators-kappa-near-zero-2026/): Raw agreement of 74.7% to 84.5% on a six-feature annotation scheme, and Cohen's kappa at chance on the feature that mattered. A worked case of what class imbalance hides. - [Once individual accuracy drops below 50%, majority vote makes rare-event labels worse](/research/rare-event-labeling-prevalence-effect-2026/): 290 annotators, 750 white blood cell images, four conditions. Two levers that cost nothing per label: the prevalence of the gold-standard feedback stream, and a recalibration step before aggregation. - [A blind 50/50 split of the annotation budget missed the near-optimal region up to 80% of the time. A $2 proxy run did not.](/research/sft-rl-annotation-budget-near-optimal-region-2026/): How to divide a fixed annotation budget between demonstrations and preference data, measured on a 5-point grid across three model families, four tasks and two RL objectives. - [Persian has 1.7% of English's web pages and 3.4 times its news-NER labels](/research/web-normalized-annotation-density-persian-2026/): A reusable metric for deciding which language and task actually needs annotation spend: divide the labeled-volume ratio by the web-presence ratio, per task, never in aggregate. - [On the same diagnostic rubric, frontier LLM judges agreed with physicians 28% to 68% of the time](/research/grand-rounds-llm-judges-vs-physicians-2026/): 9,217 physician scores, five clinical tasks, eight LLM judges. No prompt-only model matched physician agreement on all five. A few dozen physician-scored cases per task moved a 32B open model by 4 to 15 points. - [HH-RLHF's label noise is 39% or 1.8%, depending on the instrument](/research/hh-rlhf-preference-label-noise-2026/): Cleanlab flags 38.97% of HH-RLHF. An influence pipeline plus Gemini 3.1 Pro confirms 1.77%. On 108 flagged evaluation records, a fine-tuned Qwen3.5-9B disagrees with the human label 62.04% of the time, against 28.9% on the full split. - [An LLM judge matches the majority label, but predicts human disagreement worse than 3 voters](/research/llm-judge-soft-labels-human-disagreement-2026/): Hard labels: Claude-4-Sonnet at or above human F1 on most strata. Soft labels: 2.9x the error of a 20-annotator sample on ChaosNLI, and on Anecdotes the widest gap to 3 human votes is on the items voters agree about. - [Only 49.7% of NVD CWE labels match the code they describe](/research/nvd-cwe-label-audit-2026/): 15,556 CVEs audited against their fix commits: 49.70% exact match, 3.63% contradicted by the evidence, 434 confirmed mislabels. The label quality that security datasets inherit is set by whoever assigned the CWE. - [Unpaid annotation tasks grew 4.2x in ACL papers while crowdsourced ones grew 1.3x](/research/unpaid-annotation-tasks-acl-2018-2025/): The ACL checklist made papers 25.8 points better at saying whether annotators were paid. It moved the unpaid rate among disclosing papers by 0.1 points. - [Anime as AI Training Data: Why the Future Is Licensed, Not Scraped (2026)](/research/anime-ai-training-data-licensing/): The Sora 2 / CODA fight, Japan's AI law, the AniBiz marketplace, and what licensed anime data for AI requires. - [Data Annotation Research: The Latest (2026 Roundup)](/research/data-annotation-latest-research/): The latest annotation research - LLM annotation, active learning, multi-agent labeling - and why humans still decide quality. - [Scraping Is Now a Liability: The 2026 Buyer's Case for Licensed AI Training Data](/research/licensed-ai-training-data-buyers-guide/): Why AI teams are switching from scraped to licensed, consented data - and the playbook to do it. ## Guides - [In IWSLT 2026's first voice-cloning track, the best clone of your speaker got the script most wrong](/guides/cross-lingual-voice-cloning-identity-tradeoff-2026/): The accuracy-versus-identity tradeoff in cross-lingual voice cloning, with all 14 published results, the selection objectives that produced them, and the metric that carried no signal. - [Japanese subtitles scored 12.19 chrF at IWSLT 2026. Re-scored without the word segmenter, 28.18](/guides/japanese-subtitle-scores-segmentation-artifact-2026/): Two measurement problems in automatic subtitling, both quantified by the organisers, and what they mean for anyone benchmarking a subtitling vendor in CJK languages. - [The first-placed system in IWSLT 2026's 2-to-4 second latency class ran at 52.7 seconds on YouTube audio](/guides/live-dubbing-latency-iwslt-2026/): What the 2026 IWSLT simultaneous speech translation results say about live dubbing lag, and why a vendor's 100 millisecond time-to-first-byte is not the number you need. - [130 hours of Mapuzugun bought 0.82 BLEU. 30 hours of Central Kurdish bought 21.09](/guides/low-resource-speech-translation-hours-vs-bleu-2026/): What the IWSLT 2026 low-resource results say about how much speech data you actually need, what kind, and why frontier LLMs lost to a fine-tuned 2022 translation model. - [The best speech recognizer gets 51% of words wrong on Moroccan YouTube speech](/guides/asr-real-world-speech-arabic-southeast-asia-2026/): GigaSpeechBench results for 16 ASR systems on Arabic, Southeast Asian, Japanese and Korean speech taken from YouTube, set against the same systems' FLEURS scores. The first step of every AI dubbing pipeline is where the long-tail locales break. - [The language industry's mid-tier shrank 4.3% in 2025 while its top 10 grew 3.6%](/guides/nimdzi-100-2026-language-industry-mid-tier/): What the 2026 Nimdzi 100 actually says about market size, segment growth, pricing and data-for-AI, with its internal inconsistencies tabulated and a company-level check of the media localization and data-for-AI claims. - [Meta's 8B Omnilingual MT beats Llama 3 70B into mid-resource languages and loses into zero-resource ones](/guides/omnilingual-mt-resource-tier-results/): Omnilingual MT's gains, broken out by how much parallel data a language has, with the thresholds, the long-tail counts and what specialization does not fix. - [Prime Video now changes the picture to fit the dub, starting with Maxton Hall](/guides/prime-video-lip-sync-visual-dubbing-2026/): Visual dubbing moved from creator video to a streaming original on 9 September 2026. What changed, how it compares with YouTube, Meta and ElevenLabs, and the Article 50 question it opens. - [A dubbed minute costs $0.33 to $9.00, and the watermark discount is gone](/guides/cost-of-a-dubbed-minute-2026/): Every published per-minute price for AI dubbing in September 2026, normalized to dollars per minute of source media, with the plan arithmetic shown. - [A 30-trillion-token corpus buys Basque a 160-million-parameter model](/guides/per-language-token-ceiling/): Per-language token counts from HPLT 3.0 and the MaLA corpus, converted into the largest model each language can compute-optimally support. The gap between English and a mid-sized European language is 5,161 to 1. - [Article 50 Applies to Your Dub, and the Watermark It Asks For Dies in Your Mix](/guides/eu-ai-act-article-50-ai-dubbing-watermarks/): AI Act Article 50 has applied since 2 August 2026. The machine-readable mark it requires does not survive voice conversion, and often does not survive your own mix and encode. - [Your multilingual LLM judge prefers the machine translation, and agrees with itself at kappa 0.24](/guides/multilingual-llm-judge-translationese-bias/): The LLM judge gating your localized builds is least consistent on the task closest to localization, and is measurably biased toward machine-translated text in exactly the low-resource languages you added it for. - [Consented, License-Clear AI Training Data Providers (2026)](/guides/consented-ai-training-data-providers/): What to look for in a consented AI training-data vendor, and how the main providers compare. - [Data Provenance & Chain of Title for AI Training Data](/guides/data-provenance-chain-of-title/): What provenance and chain of title mean for AI training data, why they matter legally, and how to verify them before you license. - [How to License Off-the-Shelf AI Training Datasets - and What They Cost](/guides/how-to-license-ai-training-datasets/): How dataset licensing works, what drives price, and typical cost ranges by modality. - [Where to License Robotics Manipulation & Human-Demonstration Datasets](/guides/licensable-robotics-training-datasets/): Open datasets, collection methods, and commercial licensing options for embodied-AI training data. - [Localization Is the New Default (2026)](/guides/localization-the-new-default/): Why localization became a default requirement for AI products - the data, the quality/consent catch, and what to build. ## Comparisons - [BLOMEGA vs Defined.ai: AI Training Data Compared (2026)](/compare/blomega-vs-defined-ai/): How BLOMEGA and Defined.ai compare on provenance, modalities, catalog, and services - and when to choose each. - [Scale AI Alternatives for AI Training Data (2026)](/compare/scale-ai-alternatives/): Ethically-sourced, consent-based alternatives to Scale AI for training data and annotation, and how to choose. ## Learn - [Is It Safe to Sell Your Data to AI Companies? (2026)](/learn/is-it-safe-to-sell-your-data-to-ai/): The real risks of selling your data to AI, and how to do it safely with consent-based platforms. ## Machine-readable resources - [Company facts](/api/company/facts.json): Source-of-truth JSON facts about BLOMEGA. - [Facts schema](/api/company/facts.schema.json): JSON Schema for the company facts document. - [Dataset catalog](/datasets.json): Structured catalog of off-the-shelf datasets. - [OpenAPI](/openapi.json): Describes BLOMEGA's public read-only JSON endpoints. - [API catalog](/.well-known/api-catalog): RFC 9727 linkset for automated API discovery. - [auth.md](/auth.md): Agent access & authentication policy (public API is unauthenticated). - [Service status](/status.json): Health of the public data endpoints. - [Agent integration](/developers/agent-integration.html): How AI agents should consume BLOMEGA data. - [Updates feed (JSON Feed)](/feed.json) / [RSS](/rss.xml): New articles and catalog updates. Every article is also available as Markdown: append `index.md` to its URL, e.g. /research/nvd-cwe-label-audit-2026/index.md Generated from the live page set by scripts/geo-discovery.mjs.