BLOMEGA

X's public Community Notes snapshot no longer serves the ratings file, and 13.3% of scored notes have lost their text

Lab note · 21 September 2026 · BLOMEGA

Abstract dark visualisation of a two-axis scatter of crowd rater positions with one axis fading out

X's public Community Notes release no longer contains the ratings file. On 21 September 2026 we probed the newest snapshot, 2026-09-20: notes, noteStatusHistory and userEnrollment return HTTP 206, and every ratings shard returns 404. The 212,900,053 votes that a paper published three days earlier used to refit the bridging algorithm are not obtainable from any path X documents. The files that remain hold 3,297,157 scored notes, and 439,244 of them, 13.32%, have already lost their text.

What the snapshot serves on 21 September 2026

Community Notes publishes a cumulative daily snapshot, and the documentation lists five files: Notes, Ratings, Note Status History, User Enrollment and Note Requests. We probed all five on the 2026-09-20 snapshot, the newest one, with range requests against the documented path.

What the 2026-09-20 public snapshot actually serves. Probed 21 September 2026 with HTTP range requests against https://ton.twimg.com/birdwatch-public-data/2026/09/20/<kind>/<kind>-0000N.zip, the path the documentation and every prior study use. Shard indices 0 to 20 tested for notes, 0 to 9 for the rest.
FileShards foundBytes servedHTTPWhat it carries
notes5 (00000 to 00004)497,800,227 plus 4 shards near 1.9 MB206note text, author id, classification, tweet id
noteStatusHistory1182,976,832206status, scoring model, lock and change timestamps
userEnrollment171,453,051206rater enrollment state
ratingsnonenot served404every helpful and not-helpful vote, the model's input
noteRequestsnonenot served404requests for a note on a post
Any older date (2026-09-10, 2026-08-20, 2026-07-05)nonenot served404X retains the newest snapshot only

Notes, noteStatusHistory and userEnrollment answer with 206 and real byte counts. Ratings answers 404 at every shard index we tried, and so does noteRequests. We also tried five alternative spellings of the ratings path (ratings_00000.zip, ratings-00000.tsv.zip, capitalised directory, bare 00000.zip, noteRatings/) and got 404 from each. Older dates are gone, which is documented behaviour: X keeps the most recent snapshot only.

We cannot say from outside whether this is an outage, a deliberate withdrawal or a path change nobody has announced. We can say what a person following the published instructions gets today, which is no ratings.

What the public files feed, and what is missing on 2026-09-20 Model: r(u,n) = mu + i_u + i_n + f_u . f_n, as published by X and refit in arXiv:2609.21496 notes-0000{0..4}.zip note text, author, topic 206 noteStatusHistory.zip status, decided-by model 206 userEnrollment.zip rater enrollment state 206 ratings-00000.zip every helpful / not-helpful vote 404 noteRequests.zip requests for a note 404 i_n note intercept the term the publication rule reads f_n note factor which camp likes the note f_u rater factor which camp the rater is in i_u rater intercept how generous the rater is estimated only from ratings CURRENTLY_RATED_HELPFUL 279,902 of 3,297,157 scored notes = 8.49% still downloadable The decision is public. The 212.9M votes behind it are not.
The bridging model's four parameter blocks and where each one's evidence lives. Three of the five public files still resolve. The one that carries the votes does not.

Why the ratings file is the whole experiment

Community Notes publishes a note under a post only when raters who usually disagree both call it helpful. To apply that rule the system has to learn who disagrees with whom, which it does with a regularised matrix factorisation over the rating matrix: a global intercept, a per-rater intercept, a per-note intercept, and one latent factor per rater and per note. The note intercept decides publication. The factor decides whether the support crossed a divide.

Three days before our probe, on 18 September 2026, Andreou and Sirivianos at the Cyprus University of Technology published arXiv:2609.21496, which refits that base model on 212,900,053 ratings from the 2026-06-14 snapshot and extends it from one latent dimension to two. Their result is that one axis is too few. A second axis lowers held-out RMSE by 0.0185 on the 0/0.5/1 rating scale, a 6.0% relative reduction and about a quarter of what the first axis buys, while a third adds 0.0008. A permutation null that shuffles rating values reproduces the 50/50 variance split and none of the held-out gain, which is the check that makes the second axis real rather than an artefact of extra parameters.

What the second axis costs is measurable. Among notes that barely split raters on the first axis and carry at least 100 ratings, the published share falls from 71.5% at the low end of second-axis disagreement to 11.7% at the high end. Past a 50-rating threshold, notes that bridge axis one while splitting axis two publish at 9.5% against 18.3% for all notes. Two topics sit furthest out on that second axis: COVID and vaccines, and Ukraine and Russia. Rater positions learned with zero exposure to those topics still predict how the same raters judge them, which rules out the reading that the axis is just subject matter.

The design argument is older than the measurement. Jonathan Warden set out a two-factor version of the Community Notes algorithm on 5 January 2024, on the grounds that the dominant latent factor need not be political polarisation at all, and can be something like expertise. The 2026 paper does not cite it. What the paper adds is the scale: 2.33 million notes, 1.07 million raters, a held-out test and a permutation null.

None of it can be rerun today. The authors already flagged that their 2026-06-14 rating file had vanished by the time they wrote the limitations section, and moved their rank comparison to a 2026-07-05 snapshot holding 55.7M ratings. Now there is no rating file at all.

Three readings of the same public release. The first two are reported in arXiv:2609.21496 (Sections 4.1 and 5.4) after its filter of five or more ratings per note and ten or more per rater. The third is ours, unfiltered, because the filter cannot be applied without the ratings file.
SnapshotRatings servedNotesRatersSource
2026-06-14212,900,0532,334,6301,067,909arXiv:2609.21496, Section 3.1
2026-07-0555,700,000718,000531,000arXiv:2609.21496, Section 4.1
2026-09-20none served (HTTP 404)3,092,305 with text, 3,297,157 scorednot derivable without ratingsour probe, 21 September 2026

13.3% of scored notes have lost their text, and the loss climbs with age

The files that do resolve still carry a usable corpus, so we measured it. The note status history holds 3,297,157 scored notes created between 23 January 2021 and 18 September 2026. The five notes shards hold 3,092,305 notes with text. Joining them, 439,244 status rows, 13.32%, have no text left in the public release.

That is documented behaviour rather than a bug: when a note is deleted, the content leaves every future snapshot while the status row stays. What the documentation does not say is how much of the corpus that removes, and how it skews.

Scored notes whose text is no longer in the public files, by year written Our join of noteStatusHistory (3,297,157 rows) against the five notes shards (3,092,305 rows), snapshot 2026-09-20. 0% 10% 20% 30% 25.7% 2021 23,070 notes 2.91% published 29.6% 2022 27,038 notes 8.18% published 19.9% 2023 455,515 notes 8.45% published 15.6% 2024 1,176,698 notes 8.29% published 9.3% 2025 799,000 notes 8.76% published 9.5% 2026 815,836 notes 8.71% published Older notes lose their text faster. Deletions are not backfilled: a deleted note keeps its status row and loses its content. Notes with no text left publish at 3.50%; notes whose text survives publish at 9.26%.
Text loss by year written, our join of the two files on the 2026-09-20 snapshot. Note volume peaked in 2024 at 1,176,698 scored notes.

The skew is strong. Notes whose text has been withdrawn were rated helpful 3.50% of the time; notes whose text survives, 9.26%. Deletion concentrates in notes that never published, so a corpus assembled from today's files over-represents successful notes by roughly a factor of two and a half on that margin. Anyone training a helpfulness classifier on the public notes file is training on a survivor sample, and nothing in the file says so.

Overall status on the full status history: 2,881,020 notes still need more ratings (87.4%), 279,902 are currently rated helpful (8.49%), 136,235 are currently rated not helpful (4.13%). The Core model decided 2,503,871 of them, the Expansion model 292,835, Expansion Plus 181,073, Core with Topics 147,939, the scoring drift guard 59,119, the first multi-group model 57,935 and the Gaussian model 19,010. Of the notes with text, 2,482,127 (80.3%) are filed as misinformed or potentially misleading and 610,176 as not misleading; 167,219 are media notes and 49,670 are collaborative notes.

Publication rate by script, on the notes that remain

The Cyprus paper reports a language-equity result on the June snapshot: crude publication rates vary by community, most of that variation is how many ratings a note attracts, and after standardising for rating volume only Hindi stays below the global rate. We cannot standardise, because standardising needs the ratings. We can count crude rates on the newest snapshot, which is what a practitioner downloading today would get.

Our count on the 2026-09-20 snapshot: every scored note whose text is still public, bucketed by the first non-Latin script found in the note summary, joined to its current status. Script is not language, and a Latin-script bucket holds English, Portuguese, French, Turkish and more. Compare arXiv:2609.21496, which found 1,921 Greek-script notes publishing at 7.76% on 2026-06-14 with a filter of at least five ratings per note.
Script of note textScored notesRated helpfulPublished share
Latin2,551,954231,1589.06%
Japanese252,56829,10211.52%
Arabic23,3571,8067.73%
Hebrew14,2891,2508.75%
Devanagari4,2682425.67%
Greek2,8291605.66%
Cyrillic2,5252299.07%
CJK2,2881998.70%
Thai2,20824611.14%
Korean1,6271519.28%
Text withdrawn439,24415,3593.50%
All scored notes3,297,157279,9028.49%
Published share by script of the note text (our count, 2026-09-20) 0 2.6 5.2 7.8 10.4 13 percent of scored notes rated helpful Latin (n=2,551,954) 9.06% 231,158 published Japanese (n=252,568) 11.52% 29,102 published Arabic (n=23,357) 7.73% 1,806 published Hebrew (n=14,289) 8.75% 1,250 published Devanagari (n=4,268) 5.67% 242 published Greek (n=2,829) 5.66% 160 published Cyrillic (n=2,525) 9.07% 229 published CJK (n=2,288) 8.7% 199 published Thai (n=2,208) 11.14% 246 published Korean (n=1,627) 9.28% 151 published global 8.49%
Crude published share by script. Our detector takes the first non-Latin script found in the note summary, so it identifies script and not language, and it will bucket an English note quoting Greek characters as Greek.

Japanese is the second largest community in the file, 252,568 notes, and publishes above the global rate at 11.52%. Greek is 2,829 notes at 5.66%, and Devanagari 4,268 at 5.67%, both below. Those are crude rates on a different snapshot with a different filter and a different detector than the paper's, so read the direction and not the decimals: the two smallest-volume communities sit at the bottom, which is the same shape the paper found before standardisation, and the paper's own conclusion is that most of that gap is ratings supply rather than the bridging rule.

What this means if you rely on crowd labels

An audit that depends on a vendor's raw labels has the lifetime of that vendor's export. The largest crowdsourced labelling system in production published 212.9 million ratings in June, 55.7 million in July and none in September, on a path that has not changed. Every published finding about the algorithm now rests on snapshots nobody else can obtain.

The decision is not the data. The status file still tells you what X published. It cannot tell you whether the publication rule worked, because the rule is a function of the votes. If your own labelling pipeline exports only adjudicated outcomes, you have built the same gap into it.

Deletion is not neutral. A 13.32% hole in the corpus that is concentrated in unpublished notes is a hole that flatters the system. Keep a hash and a status row for every withdrawn item, which Community Notes does, and also keep the sampling fraction visible, which it does not.

One axis of disagreement is a modelling choice, not a fact about people. The same warning applies to any consensus label: if your aggregation collapses raters onto a single agreement scale, the disagreement it cannot represent shows up as low confidence, and the annotation gets thrown away rather than flagged.

Check it yourself

The probe is two commands. A 206 means the file exists and the server honoured a range request.

D=2026/09/20   # the newest snapshot; older dates are not retained
for k in notes ratings noteStatusHistory userEnrollment noteRequests; do
  printf "%-20s " "$k"
  curl -s -o /dev/null -w "%{http_code}\n" -r 0-10 \
    "https://ton.twimg.com/birdwatch-public-data/$D/$k/$k-00000.zip"
done
# notes 206 / ratings 404 / noteStatusHistory 206 / userEnrollment 206 / noteRequests 404

The text-loss measurement needs the two files that do download, about 680 MB compressed.

curl -sL -o notes.zip "https://ton.twimg.com/birdwatch-public-data/$D/notes/notes-00000.zip"
curl -sL -o nsh.zip   "https://ton.twimg.com/birdwatch-public-data/$D/noteStatusHistory/noteStatusHistory-00000.zip"
unzip -q notes.zip; unzip -q nsh.zip
python3 - <<'PY'
import csv; csv.field_size_limit(10**9)
have={r['noteId'] for r in csv.DictReader(open('notes-00000.tsv'),delimiter='\t')}
tot=miss=crh=crh_miss=0
for r in csv.DictReader(open('noteStatusHistory-00000.tsv'),delimiter='\t'):
    tot+=1; h=r['noteId'] in have; c=r['currentStatus']=='CURRENTLY_RATED_HELPFUL'
    crh+=c
    if not h: miss+=1; crh_miss+=c
print(tot, miss, round(100*miss/tot,2), round(100*crh/tot,2), round(100*crh_miss/miss,2))
PY

Shards 00001 to 00004 of the notes file add another 49,000 notes; the numbers in this article include them. The Cyprus paper's own artefacts are on OSF at osf.io/n5w4z.

What would prove this wrong

Our claim is narrow: on 21 September 2026, following the documented path, the ratings file is not obtainable, and the notes corpus that is obtainable is missing 13.32% of its rows' text with the loss concentrated in unpublished notes. It is wrong if X is serving ratings from a path we did not try, and a single working URL settles that. It is wrong in the other direction if the ratings file returns and the 2026-09-20 date resolves again, since the snapshot would then have been a gap rather than a withdrawal. We predict that by 1 December 2026 either the ratings file is back at the documented path or the Community Notes documentation is amended to stop listing it, and that the text-loss share of the cumulative corpus is higher than 13.32%, because the notes file only ever loses rows for a given creation year.

Sources

  1. Andreou, Sirivianos. Two Fault Lines: Latent Polarity Geometry in X Community Notes. arXiv:2609.21496v1, 18 September 2026, Cyprus University of Technology. Sections 3.1, 4.1, 4.4, 4.5, 5.4, Tables 1 to 4.
  2. X Community Notes. Downloading data, the documentation that lists the five files and the daily best-effort snapshot. Retrieved 21 September 2026.
  3. Community Notes public data, snapshot 2026-09-20, files notes, noteStatusHistory and userEnrollment, downloaded and counted 21 September 2026. Download page.
  4. Wojcik et al. Birdwatch: Crowd Wisdom and Bridging Algorithms can Inform Understanding and Reduce the Spread of Misinformation. arXiv:2210.15723, 2022. The original bridging model.
  5. Warden. Multidimensional Community Notes, 5 January 2024. A two-factor proposal predating the measurement.

FAQ

Can you still download Community Notes ratings data?

Not on 21 September 2026. The newest public snapshot, dated 2026-09-20, serves notes, noteStatusHistory and userEnrollment with HTTP 206 at the documented path, while every ratings shard index and noteRequests return 404. X retains only the most recent snapshot, so older dates return 404 for all files. The ratings file is the model's only input, so the bridging computation cannot be reproduced from the public release.

How many Community Notes are published?

On the 2026-09-20 snapshot, 279,902 of 3,297,157 scored notes are currently rated helpful, which is 8.49%. Another 136,235 (4.13%) are currently rated not helpful and 2,881,020 (87.4%) still need more ratings. The Core model decided 2,503,871 of the statuses.

Is the public Community Notes corpus complete?

No. Joining the status history to the note text on the 2026-09-20 snapshot, 439,244 of 3,297,157 scored notes (13.32%) have no text left in the public files, because deleted notes keep their status row and lose their content. The loss is age-dependent, from 9.3% of notes written in 2025 to 29.6% of notes written in 2022, and it is concentrated in notes that never published: withdrawn notes were rated helpful 3.50% of the time against 9.26% for notes whose text survives.

Is one axis of disagreement enough for bridging?

arXiv:2609.21496 says no. Refitting X's base model on 212.9M ratings and extending it to two dimensions lowers held-out RMSE by 0.0185, about a quarter of what the first axis buys, while a third dimension adds 0.0008. Among lightly polarised notes with at least 100 ratings, the published share falls from 71.5% to 11.7% as disagreement on the second axis grows.