BLOMEGA

Persian has 1.7% of English's web pages and 3.4 times its news-NER labels

Lab note · 16 September 2026 · BLOMEGA

Four slate columns of very unequal height rising from a thin amber horizontal baseline on a dark field, one column far taller than the rest and one falling below the line

A review posted 25 August 2026 takes 34 Persian text resources and divides each one's size, task by task, by Persian's share of the crawled web. Common Crawl CC-MAIN-2026-30 puts Persian on 0.7039% of HTML pages against English's 40.5782%, a ratio of 0.01735. Normalised against that, Persian carries 197.0 times its web-proportional share of news NER labels, 33.8 for dependency parsing, 17.2 for news summarization, and 0.60 for natural-language inference. One of those four is scarce.

What changed, and when

On 25 August 2026, MohammadHossein Mortazavi, Mostafa Salehi and Hadi Veisi posted The Annotation Bottleneck in Persian Text NLP: Persian as an Annotation-Scarce Language (arXiv:2608.24698, 18 pages, 8 tables). It reviews 34 representative Persian text resources available by July 2026 and adds three quantitative cross-checks.

The argument is a definitional one with a metric attached. "Low-resource" collapses several distinct shortages into one label: no raw text, no labeled data, no evaluation sets, no tooling, no access. Persian fails only some of those tests, and the paper's contribution is a way to say which. The metric is web-normalized annotation density (WNAD), and it is small enough to state in one line: take the ratio of the target language's annotated resource size to a comparable English resource for the same task, then divide it by the two languages' web page-share ratio. A value of 1.0 means the labels are exactly proportional to the language's share of the web.

The raw-text side of the argument is settled quickly and is worth having on hand. Persian corpora include Hamshahri at 166,774 categorized newspaper documents, the ParsBERT collection at roughly 3.98 million documents and more than 38 million sentence-like segments, MirasText at about 2.84 million documents and 1.43 billion tokens from more than 250 sites, hmBlogs at nearly 20 million blog posts and over 6.8 billion tokens, and Matina at 72.9 billion preprocessed and deduplicated tokens. A general claim of raw-text absence is not defensible against those numbers.

The evidence table

Two independent July 2026 measurements of web presence, with different denominators. W3Techs counts websites; Common Crawl counts crawled HTML pages by CLD2-identified primary language. The paper is explicit that the two should not be combined, and cites their convergence rather than their sum.

LanguageW3Techs, % of websitesCC-MAIN-2026-30, % of HTML pagesNoteSource
English49.640.5782dominant baseline languageTable 1
Turkish1.61.3455regional comparison, larger web shareTable 1
Persian0.90.7039within roughly the top 20 under both viewsTable 1
Arabic0.60.6548regional script-sharing comparisonTable 1
Two independent July 2026 measurements, different denominators, same shape log scale, because English is two orders of magnitude above the rest EnglishTurkishPersianArabic 49.6%40.5782% 1.6%1.3455% 0.9%0.7039% 0.6%0.6548% 0.1%1%10%100% W3Techs, share of websitesCommon Crawl, share of HTML pages Source: arXiv:2608.24698, Table 1. The two denominators differ and must not be combined; only their convergence is the evidence.
Persian sits within roughly the top twenty identified languages under both views. That is the number the annotation counts get divided by.

Then the same treatment applied to labels. Each row compares one documented Persian resource against one canonical English resource in the same unit, and divides the resulting ratio by 0.01735.

TaskPersian resourceEnglish comparatorFa / EnWNADSource
News NERNSURL-2019, 1,029,822 tokensCoNLL-2003 English, 301,418 tokens3.42197.0Table 8
Dependency parsingPerDT + Seraji, 645,790 tokens13 treebanks on the English UD comparison page, 1,102,940 tokens0.58633.8Table 8
News summarizationpn-summary, 93,207 recordsCNN/DailyMail standard splits, 312,084 pairs0.29917.2Table 8
Natural-language inferenceFarsTail, 10,367 examplesSNLI + MultiNLI, ~1.003M examples0.01030.60Table 8
Web-proportional baseline0.017351.00§8.2
Persian labels per unit of Persian web presence, relative to English log scale; 1.0 = exactly proportional to Persian's 1.735% share of English's page count News NER Dependency parsing News summarization Natural-language inference baseline 1.0 197.033.817.20.60 0.11101001000 web-normalized annotation density (WNAD) The 197.0 is a warning about comparator choice: CoNLL-2003 is one canonical English NER set, not all English NER data. Source: arXiv:2608.24698, Table 8 and Section 8.2. The paper reports no cross-task average, and neither do we.
Two and a half orders of magnitude separate the best-served and worst-served Persian task. That spread, not the average, is the finding.

What the ratio corrects for, and what it does not

The naive version of this question is "does Persian have enough labeled data", and it has no answer, because enough compared to what. Comparing absolute counts to English penalises every language for not being English. Comparing to speaker population ignores that NLP consumes text, not speakers. WNAD picks the denominator that matches what a model actually trains on: crawled web pages in that language.

One denominator, chosen because it is what models read Common Crawl CC-MAIN-2026-30Persian 0.7039% of HTML pagesEnglish 40.5782%Rₑₑ₋ = 0.01735 One task, one unitNSURL-2019: 1,029,822 tokensCoNLL-2003: 301,418 tokensAₑₐ / Aₑₙ = 3.42 WNAD = (Aₑₐ/Aₑₙ) / Rₑₑ₋3.42 / 0.01735 = 197.0 Read the value> 1: denser than the web= 1: proportional< 1: thinner (NLI, 0.60) Three guards the paper puts on its own metricSame unit both sides · never sum or average across tasks · the English side is a named benchmark, not the whole ecosystem Source: arXiv:2608.24698, Section 8.2 and Table 8
The metric is four numbers and a division. Its value is entirely in the discipline around it: per task, same unit, named comparator, no averaging.

What WNAD does not measure is stated as plainly in the paper as the metric itself, and the caveats are the reason the number is usable rather than decorative. A high density says nothing about domain coverage, annotation quality, licensing, documentation or access. The paper's own task-level profile lists the failures that a volume ratio cannot see: concentration in news, incompatible label inventories across resources, limited coverage of medical, legal and colloquial text, and thin support for varieties beyond standard Iranian Persian. Persian news NER scores 197.0 and is still news NER.

The 197.0 is also a lesson in comparator sensitivity, and the authors say so rather than banking the headline. CoNLL-2003 English is 301,418 tokens. It is the canonical English NER benchmark and nothing close to the total of English NER data. Swap the denominator for a larger English NER collection and the ratio falls by whatever factor you chose. This is why the paper reports no cross-task average: the metric is a within-task diagnostic whose comparator has to be named every time it is quoted.

One more thing worth recording about provenance. The manuscript states that large language models, including Claude Sonnet 5, were used to improve the clarity and phrasing of the text. That is a disclosure about the writing, not about the numbers, and the resource counts are attributed to named source publications throughout.

What it means if you are planning where to spend annotation money

"Low-resource" is a purchasing decision disguised as a description. The four Persian tasks here span 0.60 to 197.0 on the same scale, in the same language, in the same year. Any budget allocated to "Persian" as a unit is allocated blind. The unit that carries information is language crossed with task, and the paper's own construction of it costs one crawl statistic and two published resource sizes.

Compute the ratio before you scope the collection. For a language and task you are considering, the inputs are: the target language's page share in the latest Common Crawl snapshot, English's page share in the same snapshot, the size of the best existing target-language resource, and the size of a named English comparator in the same unit. Four numbers, all public. The output tells you whether you are filling a genuine gap or adding to an island that is already dense.

The gaps that matter are structural, not numeric. Persian's shortages, as the paper characterises them, are uneven task and domain coverage, incompatible annotation schemes, access and documentation friction, thin supervision for specialist domains and preference data, and almost nothing beyond standard Iranian Persian. None of those show up in a volume count, and all of them show up in a delivery schedule. Our judgement: for a consented-data programme, the scheme-compatibility and licensing problems are the expensive ones, because they cannot be fixed by collecting more.

Preference data is the named gap. The paper singles out preference data, specialist domains and non-standard varieties as the places where supervision is thin, and those are exactly the categories that post-training now consumes. The NLI figure of 0.60 is the only one of the four measured tasks below the baseline, and inference-style supervision is closer in kind to preference annotation than news NER is.

Re-run it per snapshot. The web baseline moves. This computation is pinned to CC-MAIN-2026-30 and to W3Techs in July 2026, and the paper says the visibility figure should be read as dated rather than as a permanent property of the language.

Check it yourself

Every number in the WNAD table is a division of two published counts. The paper has no arXiv HTML rendering, so the tables come from the PDF.

# the paper (HTML build is absent; PDF resolves)
curl -s -o /dev/null -w "%{http_code}\n" https://arxiv.org/abs/2608.24698    # 200
curl -s -o /dev/null -w "%{http_code}\n" https://arxiv.org/html/2608.24698   # 404
curl -sL -o 2608.24698.pdf https://arxiv.org/pdf/2608.24698 && python3 -c "
from pypdf import PdfReader; r=PdfReader('2608.24698.pdf'); print(len(r.pages), 'pages')"
# 18 pages

# recompute every cell of Table 8 from the two published counts
python3 - <<'PY'
FA_PAGES, EN_PAGES = 0.7039, 40.5782        # CC-MAIN-2026-30, CLD2 primary language
R_web = FA_PAGES / EN_PAGES
print(f"R_web = {R_web:.5f}\n")

rows = [                                     # task, Persian units, English units
  ("News NER",            1_029_822,   301_418),   # NSURL-2019 vs CoNLL-2003 English
  ("Dependency parsing",    645_790, 1_102_940),   # PerDT + Seraji vs 13 English UD treebanks
  ("News summarization",     93_207,   312_084),   # pn-summary vs CNN/DailyMail splits
  ("NLI",                    10_367, 1_003_000),   # FarsTail vs SNLI + MultiNLI
]
for task, fa, en in rows:
    ratio = fa / en
    print(f"{task:20} Fa/En {ratio:7.4f}   WNAD {ratio / R_web:7.1f}")
PY
# R_web = 0.01735
#
# News NER             Fa/En  3.4166   WNAD   197.0
# Dependency parsing   Fa/En  0.5855   WNAD    33.8
# News summarization   Fa/En  0.2987   WNAD    17.2
# NLI                  Fa/En  0.0103   WNAD     0.6

# run it for your own language pair: the page-share table is published per crawl
# https://commoncrawl.github.io/cc-crawl-statistics/plots/languages
# and W3Techs publishes the website-share view
# https://w3techs.com/technologies/overview/content_language

Two limits on what this reproduces. The Persian and English resource sizes are source-reported, taken from each dataset's own publication rather than recounted from the files, so an error in a source paper propagates. And the English comparator on each row is a choice made by the authors; the dependency-parsing row aggregates 13 treebanks while the NER row uses a single benchmark, which is a difference in kind between rows that the single WNAD column does not show.

What would prove this wrong

The claim worth testing is the general one, not the Persian one: that within-task, web-normalized annotation density varies enough across tasks in a single language to make language-level "low-resource" labels useless for planning.

A dated prediction. By 31 December 2027, applying this computation to any language outside the top five by Common Crawl page share, across at least four tasks with named English comparators, will produce a spread of at least one order of magnitude between its highest and lowest task density. If someone runs it on three such languages and finds the per-task densities clustered within a factor of ten, the language-level label is doing more work than we are giving it credit for and this reading is wrong. The result would be most likely to hold for a language whose entire NLP resource base came from one coordinated programme rather than from accreted individual efforts.

Sources

  1. Mortazavi, M., Salehi, M., Veisi, H. The Annotation Bottleneck in Persian Text NLP: Persian as an Annotation-Scarce Language. arXiv:2608.24698, 25 August 2026. Tables 1, 7 and 8, Sections 5.2, 8.2 and 10; PDF, 18 pages.
  2. Common Crawl language statistics, crawl CC-MAIN-2026-30. The page-share figures the normalization divides by.
  3. W3Techs content language usage, July 2026. The independent website-share measurement.
  4. Tjong Kim Sang, E. F., De Meulder, F. Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition. CoNLL 2003. The 301,418-token English comparator behind the 197.0 figure.
  5. BLOMEGA. Unpaid annotation tasks grew 4.2x in ACL papers, on who actually produces the labels these counts describe.
  6. BLOMEGA. Data annotation research: the latest.