X's public Community Notes snapshot no longer serves the ratings file, and 13.3% of scored notes have lost their text
X's public Community Notes release no longer contains the ratings file. On 21 September 2026 we probed the newest snapshot, 2026-09-20: notes, noteStatusHistory and userEnrollment return HTTP 206, and every ratings shard returns 404. The 212,900,053 votes that a paper published three days earlier used to refit the bridging algorithm are not obtainable from any path X documents. The files that remain hold 3,297,157 scored notes, and 439,244 of them, 13.32%, have already lost their text.
What the snapshot serves on 21 September 2026
Community Notes publishes a cumulative daily snapshot, and the documentation lists five files: Notes, Ratings, Note Status History, User Enrollment and Note Requests. We probed all five on the 2026-09-20 snapshot, the newest one, with range requests against the documented path.
| File | Shards found | Bytes served | HTTP | What it carries |
|---|---|---|---|---|
| notes | 5 (00000 to 00004) | 497,800,227 plus 4 shards near 1.9 MB | 206 | note text, author id, classification, tweet id |
| noteStatusHistory | 1 | 182,976,832 | 206 | status, scoring model, lock and change timestamps |
| userEnrollment | 1 | 71,453,051 | 206 | rater enrollment state |
| ratings | none | not served | 404 | every helpful and not-helpful vote, the model's input |
| noteRequests | none | not served | 404 | requests for a note on a post |
| Any older date (2026-09-10, 2026-08-20, 2026-07-05) | none | not served | 404 | X retains the newest snapshot only |
Notes, noteStatusHistory and userEnrollment answer with 206 and real byte counts. Ratings answers 404 at every shard index we tried, and so does noteRequests. We also tried five alternative spellings of the ratings path (ratings_00000.zip, ratings-00000.tsv.zip, capitalised directory, bare 00000.zip, noteRatings/) and got 404 from each. Older dates are gone, which is documented behaviour: X keeps the most recent snapshot only.
We cannot say from outside whether this is an outage, a deliberate withdrawal or a path change nobody has announced. We can say what a person following the published instructions gets today, which is no ratings.
Why the ratings file is the whole experiment
Community Notes publishes a note under a post only when raters who usually disagree both call it helpful. To apply that rule the system has to learn who disagrees with whom, which it does with a regularised matrix factorisation over the rating matrix: a global intercept, a per-rater intercept, a per-note intercept, and one latent factor per rater and per note. The note intercept decides publication. The factor decides whether the support crossed a divide.
Three days before our probe, on 18 September 2026, Andreou and Sirivianos at the Cyprus University of Technology published arXiv:2609.21496, which refits that base model on 212,900,053 ratings from the 2026-06-14 snapshot and extends it from one latent dimension to two. Their result is that one axis is too few. A second axis lowers held-out RMSE by 0.0185 on the 0/0.5/1 rating scale, a 6.0% relative reduction and about a quarter of what the first axis buys, while a third adds 0.0008. A permutation null that shuffles rating values reproduces the 50/50 variance split and none of the held-out gain, which is the check that makes the second axis real rather than an artefact of extra parameters.
What the second axis costs is measurable. Among notes that barely split raters on the first axis and carry at least 100 ratings, the published share falls from 71.5% at the low end of second-axis disagreement to 11.7% at the high end. Past a 50-rating threshold, notes that bridge axis one while splitting axis two publish at 9.5% against 18.3% for all notes. Two topics sit furthest out on that second axis: COVID and vaccines, and Ukraine and Russia. Rater positions learned with zero exposure to those topics still predict how the same raters judge them, which rules out the reading that the axis is just subject matter.
The design argument is older than the measurement. Jonathan Warden set out a two-factor version of the Community Notes algorithm on 5 January 2024, on the grounds that the dominant latent factor need not be political polarisation at all, and can be something like expertise. The 2026 paper does not cite it. What the paper adds is the scale: 2.33 million notes, 1.07 million raters, a held-out test and a permutation null.
None of it can be rerun today. The authors already flagged that their 2026-06-14 rating file had vanished by the time they wrote the limitations section, and moved their rank comparison to a 2026-07-05 snapshot holding 55.7M ratings. Now there is no rating file at all.
| Snapshot | Ratings served | Notes | Raters | Source |
|---|---|---|---|---|
| 2026-06-14 | 212,900,053 | 2,334,630 | 1,067,909 | arXiv:2609.21496, Section 3.1 |
| 2026-07-05 | 55,700,000 | 718,000 | 531,000 | arXiv:2609.21496, Section 4.1 |
| 2026-09-20 | none served (HTTP 404) | 3,092,305 with text, 3,297,157 scored | not derivable without ratings | our probe, 21 September 2026 |
13.3% of scored notes have lost their text, and the loss climbs with age
The files that do resolve still carry a usable corpus, so we measured it. The note status history holds 3,297,157 scored notes created between 23 January 2021 and 18 September 2026. The five notes shards hold 3,092,305 notes with text. Joining them, 439,244 status rows, 13.32%, have no text left in the public release.
That is documented behaviour rather than a bug: when a note is deleted, the content leaves every future snapshot while the status row stays. What the documentation does not say is how much of the corpus that removes, and how it skews.
The skew is strong. Notes whose text has been withdrawn were rated helpful 3.50% of the time; notes whose text survives, 9.26%. Deletion concentrates in notes that never published, so a corpus assembled from today's files over-represents successful notes by roughly a factor of two and a half on that margin. Anyone training a helpfulness classifier on the public notes file is training on a survivor sample, and nothing in the file says so.
Overall status on the full status history: 2,881,020 notes still need more ratings (87.4%), 279,902 are currently rated helpful (8.49%), 136,235 are currently rated not helpful (4.13%). The Core model decided 2,503,871 of them, the Expansion model 292,835, Expansion Plus 181,073, Core with Topics 147,939, the scoring drift guard 59,119, the first multi-group model 57,935 and the Gaussian model 19,010. Of the notes with text, 2,482,127 (80.3%) are filed as misinformed or potentially misleading and 610,176 as not misleading; 167,219 are media notes and 49,670 are collaborative notes.
Publication rate by script, on the notes that remain
The Cyprus paper reports a language-equity result on the June snapshot: crude publication rates vary by community, most of that variation is how many ratings a note attracts, and after standardising for rating volume only Hindi stays below the global rate. We cannot standardise, because standardising needs the ratings. We can count crude rates on the newest snapshot, which is what a practitioner downloading today would get.
| Script of note text | Scored notes | Rated helpful | Published share |
|---|---|---|---|
| Latin | 2,551,954 | 231,158 | 9.06% |
| Japanese | 252,568 | 29,102 | 11.52% |
| Arabic | 23,357 | 1,806 | 7.73% |
| Hebrew | 14,289 | 1,250 | 8.75% |
| Devanagari | 4,268 | 242 | 5.67% |
| Greek | 2,829 | 160 | 5.66% |
| Cyrillic | 2,525 | 229 | 9.07% |
| CJK | 2,288 | 199 | 8.70% |
| Thai | 2,208 | 246 | 11.14% |
| Korean | 1,627 | 151 | 9.28% |
| Text withdrawn | 439,244 | 15,359 | 3.50% |
| All scored notes | 3,297,157 | 279,902 | 8.49% |
Japanese is the second largest community in the file, 252,568 notes, and publishes above the global rate at 11.52%. Greek is 2,829 notes at 5.66%, and Devanagari 4,268 at 5.67%, both below. Those are crude rates on a different snapshot with a different filter and a different detector than the paper's, so read the direction and not the decimals: the two smallest-volume communities sit at the bottom, which is the same shape the paper found before standardisation, and the paper's own conclusion is that most of that gap is ratings supply rather than the bridging rule.
What this means if you rely on crowd labels
An audit that depends on a vendor's raw labels has the lifetime of that vendor's export. The largest crowdsourced labelling system in production published 212.9 million ratings in June, 55.7 million in July and none in September, on a path that has not changed. Every published finding about the algorithm now rests on snapshots nobody else can obtain.
The decision is not the data. The status file still tells you what X published. It cannot tell you whether the publication rule worked, because the rule is a function of the votes. If your own labelling pipeline exports only adjudicated outcomes, you have built the same gap into it.
Deletion is not neutral. A 13.32% hole in the corpus that is concentrated in unpublished notes is a hole that flatters the system. Keep a hash and a status row for every withdrawn item, which Community Notes does, and also keep the sampling fraction visible, which it does not.
One axis of disagreement is a modelling choice, not a fact about people. The same warning applies to any consensus label: if your aggregation collapses raters onto a single agreement scale, the disagreement it cannot represent shows up as low confidence, and the annotation gets thrown away rather than flagged.
Check it yourself
The probe is two commands. A 206 means the file exists and the server honoured a range request.
D=2026/09/20 # the newest snapshot; older dates are not retained
for k in notes ratings noteStatusHistory userEnrollment noteRequests; do
printf "%-20s " "$k"
curl -s -o /dev/null -w "%{http_code}\n" -r 0-10 \
"https://ton.twimg.com/birdwatch-public-data/$D/$k/$k-00000.zip"
done
# notes 206 / ratings 404 / noteStatusHistory 206 / userEnrollment 206 / noteRequests 404
The text-loss measurement needs the two files that do download, about 680 MB compressed.
curl -sL -o notes.zip "https://ton.twimg.com/birdwatch-public-data/$D/notes/notes-00000.zip"
curl -sL -o nsh.zip "https://ton.twimg.com/birdwatch-public-data/$D/noteStatusHistory/noteStatusHistory-00000.zip"
unzip -q notes.zip; unzip -q nsh.zip
python3 - <<'PY'
import csv; csv.field_size_limit(10**9)
have={r['noteId'] for r in csv.DictReader(open('notes-00000.tsv'),delimiter='\t')}
tot=miss=crh=crh_miss=0
for r in csv.DictReader(open('noteStatusHistory-00000.tsv'),delimiter='\t'):
tot+=1; h=r['noteId'] in have; c=r['currentStatus']=='CURRENTLY_RATED_HELPFUL'
crh+=c
if not h: miss+=1; crh_miss+=c
print(tot, miss, round(100*miss/tot,2), round(100*crh/tot,2), round(100*crh_miss/miss,2))
PY
Shards 00001 to 00004 of the notes file add another 49,000 notes; the numbers in this article include them. The Cyprus paper's own artefacts are on OSF at osf.io/n5w4z.
What would prove this wrong
Our claim is narrow: on 21 September 2026, following the documented path, the ratings file is not obtainable, and the notes corpus that is obtainable is missing 13.32% of its rows' text with the loss concentrated in unpublished notes. It is wrong if X is serving ratings from a path we did not try, and a single working URL settles that. It is wrong in the other direction if the ratings file returns and the 2026-09-20 date resolves again, since the snapshot would then have been a gap rather than a withdrawal. We predict that by 1 December 2026 either the ratings file is back at the documented path or the Community Notes documentation is amended to stop listing it, and that the text-loss share of the cumulative corpus is higher than 13.32%, because the notes file only ever loses rows for a given creation year.
Sources
- Andreou, Sirivianos. Two Fault Lines: Latent Polarity Geometry in X Community Notes. arXiv:2609.21496v1, 18 September 2026, Cyprus University of Technology. Sections 3.1, 4.1, 4.4, 4.5, 5.4, Tables 1 to 4.
- X Community Notes. Downloading data, the documentation that lists the five files and the daily best-effort snapshot. Retrieved 21 September 2026.
- Community Notes public data, snapshot 2026-09-20, files
notes,noteStatusHistoryanduserEnrollment, downloaded and counted 21 September 2026. Download page. - Wojcik et al. Birdwatch: Crowd Wisdom and Bridging Algorithms can Inform Understanding and Reduce the Spread of Misinformation. arXiv:2210.15723, 2022. The original bridging model.
- Warden. Multidimensional Community Notes, 5 January 2024. A two-factor proposal predating the measurement.
FAQ
Can you still download Community Notes ratings data?
Not on 21 September 2026. The newest public snapshot, dated 2026-09-20, serves notes, noteStatusHistory and userEnrollment with HTTP 206 at the documented path, while every ratings shard index and noteRequests return 404. X retains only the most recent snapshot, so older dates return 404 for all files. The ratings file is the model's only input, so the bridging computation cannot be reproduced from the public release.
How many Community Notes are published?
On the 2026-09-20 snapshot, 279,902 of 3,297,157 scored notes are currently rated helpful, which is 8.49%. Another 136,235 (4.13%) are currently rated not helpful and 2,881,020 (87.4%) still need more ratings. The Core model decided 2,503,871 of the statuses.
Is the public Community Notes corpus complete?
No. Joining the status history to the note text on the 2026-09-20 snapshot, 439,244 of 3,297,157 scored notes (13.32%) have no text left in the public files, because deleted notes keep their status row and lose their content. The loss is age-dependent, from 9.3% of notes written in 2025 to 29.6% of notes written in 2022, and it is concentrated in notes that never published: withdrawn notes were rated helpful 3.50% of the time against 9.26% for notes whose text survives.
Is one axis of disagreement enough for bridging?
arXiv:2609.21496 says no. Refitting X's base model on 212.9M ratings and extending it to two dimensions lowers held-out RMSE by 0.0185, about a quarter of what the first axis buys, while a third dimension adds 0.0008. Among lightly polarised notes with at least 100 ratings, the published share falls from 71.5% to 11.7% as disagreement on the second axis grows.