Unpaid annotation tasks grew 4.2x in ACL papers while crowdsourced ones grew 1.3x
An audit of 2,667 human-annotation tasks across 1,603 ACL-venue papers reports that the share of tasks saying nothing about annotator compensation fell from 62.5% before 2022 to 36.7% after. Condition that on the papers that disclose a compensation status at all and the unpaid share is 24.0% in 2018 to 2021 and 24.2% in 2022 to 2025. Disclosure moved 25.8 points. The rate moved 0.1. Over the same boundary the number of tasks annotated by the paper's own authors went from 51 to 200 while crowdsourced tasks went from 358 to 476, against a corpus that grew 2.48x.
What changed, and when
On 1 June 2026 Maria Kunilovskaya, Gagan Bhatia, Lisa Sophie Albertelli, Yanran Chen, Christian Greisinger, Lotta Kiefer, Christoph Leiter, Subhadeep Roy, Tewodros Achamaleh, Muhammad Arslan Manzoor, Sebastian Pohl, Yufang Hou and Steffen Eger posted Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025 (arXiv:2606.02255). It was revised to v2 on 1 September 2026 and to v3 on 2 September 2026. The arXiv record lists no venue.
The method is a 26-category taxonomy over seven aspects of annotation reporting, applied task by task rather than paper by paper. The authors hand-adjudicated a gold set they call Annotated-gold, 41 papers and 72 tasks, then ran an LLM extraction pipeline over ACL-venue papers to build Annotated-llm: 1,603 papers, 2,667 tasks, retained from 1,995 candidates. On the gold set, Gemini-3.1-Pro reached 79.9% macro agreement and Krippendorff's alpha 0.606 against the adjudicated labels, marginally above the human-human baseline of 79.2% and 0.585. The audit's own labels are model-produced, which is a real limitation and one the authors state.
The comparison that matters here is a 2022 boundary. The ACL Responsible NLP Research Checklist, introduced at NAACL 2022 and derived from the NeurIPS 2021 checklist, asks authors in item D2 whether they reported how they recruited and paid participants and whether the wage was fair, and in item D5 whether they reported basic demographic and geographic characteristics of the annotator population. Appendix F of the paper splits every taxonomy field into 2018-2021 (766 tasks) and 2022-2025 (1,901 tasks) and runs a chi-squared test on each.
Read down that appendix and the checklist looks like it worked. Almost every field moves in the right direction and almost every move is significant. The compensation field moves the most.
The evidence table
Counts and percentages below are transcribed from Appendix F of arXiv:2606.02255v3. Percentages are the paper's; counts are the paper's; the delta and multiple columns are ours.
| Field and value | 2018-2021 n (%) | 2022-2025 n (%) | Delta | Multiple | Source |
|---|---|---|---|---|---|
| Compensation: not reported | 479 (62.5) | 697 (36.7) | -25.8 pt * | 1.46x | Appx F |
| Compensation: paid | 218 (28.5) | 913 (48.0) | +19.5 pt * | 4.19x | Appx F |
| Compensation: free (voluntary) | 69 (9.0) | 291 (15.3) | +6.3 pt * | 4.22x | Appx F |
| Payment rate: specific numeric rate | 139 (18.1) | 605 (31.8) | +13.7 pt * | 4.35x | Appx F |
| Payment rate: general mention only | 79 (10.3) | 308 (16.2) | +5.9 pt * | 3.90x | Appx F |
| Recruitment: crowdsourcing | 358 (46.7) | 476 (25.0) | -21.7 pt * | 1.33x | Appx F |
| Recruitment: the paper's authors | 51 (6.7) | 200 (10.5) | +3.8 pt * | 3.92x | Appx F |
| Annotator training reported | 99 (12.9) | 399 (21.0) | +8.1 pt * | 4.03x | Appx F |
| Age reported | 6 (0.8) | 138 (7.3) | +6.5 pt | 23.0x | Appx F |
| Gender reported | 8 (1.0) | 157 (8.3) | +7.3 pt | 19.6x | Appx F |
| Nation of origin reported | 4 (0.5) | 56 (2.9) | +2.4 pt | 14.0x | Appx F |
| No adjudication procedure | 593 (77.4) | 1,437 (75.6) | -1.8 pt | 2.42x | Appx F |
| All annotation tasks | 766 | 1,901 | n/a | 2.48x | Appx F |
* marks a difference in proportions the paper reports as significant under a chi-squared test at p<0.05. The paper omits the test (NA) where either period's count falls below 30, which is why the demographic rows carry no star despite the largest multiples. Multiple = 2022-2025 count divided by 2018-2021 count, computed by us.
Three quantities in that table are not the paper's. They are ours, computed from its counts, and they are the reason this article exists.
| Derived metric | 2018-2021 | 2022-2025 | Change | How it is computed |
|---|---|---|---|---|
| Unpaid share of tasks that state a compensation status | 24.0% | 24.2% | +0.1 pt | free / (paid + free): 69/287, then 291/1204 |
| Share of "paid" tasks naming an actual amount | 63.8% | 66.3% | +2.5 pt | specific numeric rate / paid: 139/218, then 605/913 |
| Crowdsourced share of all tasks | 46.7% | 25.0% | -21.7 pt | the paper's own marginal, shown for contrast |
Computed by BLOMEGA from Appendix F of arXiv:2606.02255v3 on 10 September 2026. Code in Check it yourself. The first row is the finding: the two conditional rates are flat across a boundary where nearly every marginal in the source table moved significantly.
The denominator does the work
The compensation field has three values and one of them is "we did not say". Every improvement narrative built on that field is a narrative about the third value. Move tasks out of "not reported" and into either of the other two and the field looks healthier whichever one they land in.
The payment-rate field is nested inside this and gives the structure away. Its NA count is 548 for 2018-2021 and 988 for 2022-2025. Those are exactly the "not reported" plus "free" counts (479 + 69 and 697 + 291). Its two informative values sum to 139 + 79 = 218 and 605 + 308 = 913, exactly the "paid" counts. So payment rate is conditioned on compensation being paid, and the fraction of paid tasks that name an actual number is 139 of 218 (63.8%) then and 605 of 913 (66.3%) now. One statement in three that a paper paid its annotators still carries no amount, and that has barely moved either.
Two conditional rates, both close to flat, in a table where nearly every marginal moved significantly. Our reading, offered as a judgement: the checklist changed what authors write down, not what they do. That is not a small thing, since you cannot audit what is not written down. It is also not the thing the marginals appear to say.
Where the new annotation labour came from
The corpus roughly two and a half times itself across the boundary, so raw counts rise everywhere. The informative quantity is each category's growth against the corpus growth of 2.48x.
Crowdsourcing, the one sourcing route that is paid by construction and external to the lab, is the only category on the chart that grew slower than the field. It fell from 46.7% of tasks to 25.0%. Voluntary effort and author-annotation, the two routes where the annotator is inside the building and generally not separately compensated, grew about 1.6 to 1.7 times faster than the corpus.
The paper points at a plausible driver and stops short of claiming it. Model-output evaluation is now the most common intended use of human annotation, 50.5% of tasks in 2022-2025 against 47.1% before, while resource creation fell from 40.8% to 37.5%. Evaluation tasks report recruitment, compensation, training and quality control significantly less often than resource-creation tasks (logistic regression, p<0.001, with publication year controlled). The authors write that this "raises the possibility that authors acting as annotators may remain systematically underreported in this setting." We could not verify the cross-tabulation ourselves. Appendix F publishes marginals only, so recruitment and compensation cannot be crossed from the released tables.
What D2 and D5 actually get you
Checklist item D2 asks for recruitment, payment and a fair-wage discussion. Item D5 asks for basic demographic and geographic characteristics. Four years after the checklist, here is what those two questions have produced across 1,901 tasks.
Read the bars as levels rather than as progress and D2 gets you a numeric pay rate on 605 of 1,901 tasks (31.8%) and a training statement on 399 (21.0%). D5 gets you education on 894 tasks (47.0%), gender on 157 (8.3%), age on 138 (7.3%) and nation of origin on 56 (2.9%). Education is the outlier because it is the one demographic a supervisor can fill in about their own graduate students without asking anyone.
Nothing here says the field is worse than it was. Age reporting went from 6 tasks to 138 and gender from 8 to 157, the largest multiples anywhere in the appendix. They are also the smallest absolute levels, and the complements are the honest way to state them: 92.7% of recent tasks report no annotator age, 91.7% no gender, 97.1% no nation of origin. The paper is explicit that missing reporting is not evidence that the practice was absent, and we adopt that position too. What can be said is narrower and still uncomfortable: for most published human judgments in NLP, a reader in 2026 cannot establish who produced them or on what terms.
What this means if you build or buy annotation
A benchmark's annotator provenance is usually unrecoverable from the paper. If your model selection depends on a leaderboard whose ground truth came from an evaluation-section annotation task, the base rates say roughly even odds the paper names no payment amount, three in four odds it describes no adjudication procedure, and better than nine in ten odds it says nothing about who the annotators were demographically. Treat "human-validated" in a benchmark card as an unaudited claim until you find the recruitment paragraph.
Author-annotated evaluation is a conflict of interest with a growing footprint. 200 tasks in 2022-2025 recruited the paper's own authors. When the annotation being scored is the output of the authors' own system, the annotator and the interested party are the same person. The paper does not claim this biases results and neither do we; we note that the count grew 3.92x against a 2.48x corpus and that the field publishes no standard for disclosing it beyond a checklist box.
A numeric rate is the only compensation claim that survives contact with an auditor. The taxonomy is instructive here: "paid" in this scheme includes the phrase "full-time employees", so a salaried researcher counts. 308 tasks in the recent window say annotators were paid and give no figure. For contrast, Prolific's published participant payment policy (page updated 13 March 2026) sets an absolute minimum of £6.00 / $8.00 per hour and a recommended minimum of £9.00 / $12.00 per hour. A platform with a public floor gives you something to check against. A sentence saying "annotators were compensated" does not.
If you are procuring annotation, ask for the fields the literature omits. Rate per hour or per item, recruitment channel, language proficiency evidence, training given, adjudication rule, and inter-annotator agreement with its metric named. Six fields. The audit shows the first, fourth and fifth are the ones vendors and papers alike are least likely to volunteer. This is the same provenance argument we make about consented training data providers and chain of title, applied one layer down, to the labels rather than the corpus.
Check it yourself
Every number above comes from one table. It is Appendix F, captioned "Impact of the ACL Responsible NLP Checklist", in the arXiv HTML rendering. It is byte-identical in v1 and v3, so either version works.
# the source table, in both versions
open https://arxiv.org/html/2606.02255v3 # Appendix F
open https://arxiv.org/html/2606.02255v1 # same counts
# the two conditional rates, from six numbers in that table
python3 - <<'PY'
pre = {"not_reported": 479, "paid": 218, "free": 69}
post = {"not_reported": 697, "paid": 913, "free": 291}
for lbl, d in (("2018-2021", pre), ("2022-2025", post)):
n = sum(d.values())
disclosed = n - d["not_reported"]
print(f"{lbl}: n={n} disclosed={disclosed} "
f"unpaid_share_of_disclosed={100*d['free']/disclosed:.1f}%")
# 2018-2021: n=766 disclosed=287 unpaid_share_of_disclosed=24.0%
# 2022-2025: n=1901 disclosed=1204 unpaid_share_of_disclosed=24.2%
PY
Two consistency checks worth running while you are in the table. First, the payment-rate row nests exactly inside the compensation row: na equals 548 and 988, which are 479+69 and 697+291, and the two informative values sum to 218 and 913, the paid counts. If your transcription breaks that identity you have mis-read a cell. Second, the single-label fields all sum to 766 and 1,901, so 2,667 tasks in total, which matches the abstract and Table 1. The Appendix F caption states "Total observations (tasks): 2,669". The columns say 2,667. We used the columns.
One thing you cannot check. The paper's footnote 1 states "Code and data are available at https://github.com/NL2G/who-are-the-annotators". Fetched on 10 September 2026, that URL returns HTTP 404 while the parent organisation github.com/NL2G returns 200 and lists 23 public repositories. The repository may be private or not yet pushed. Until it is reachable, the Annotated-llm task table cannot be re-analysed independently and the cross-tabulations we wanted, recruitment against compensation, cannot be computed by anyone outside the author group.
curl -s -o /dev/null -w "%{http_code}\n" https://github.com/NL2G/who-are-the-annotators # 404
curl -s -o /dev/null -w "%{http_code}\n" https://github.com/NL2G # 200
What would prove this wrong
The central claim is conditional and its weakness is selection. The 24.0% and 24.2% figures describe only tasks whose papers said something, and the set of papers that said something grew from 287 tasks to 1,204. If the 917 newly-disclosing tasks are systematically different from the always-disclosing ones, the stability of the conditional rate is an artifact rather than a fact about annotation practice. Nothing in the published marginals can settle that.
The prediction. When github.com/NL2G/who-are-the-annotators becomes reachable and the per-task Annotated-llm records can be joined, splitting the 2022-2025 disclosing tasks by publication year will show an unpaid share that stays inside a 20% to 28% band for every year from 2022 to 2025, with no monotonic trend. If instead the yearly series moves outside that band or trends in one direction across all four years, the flat conditional rate reported here is a composition effect and this article's reading is wrong. We will re-run it and say so.
A second, cheaper falsifier: if the authors publish the recruitment-by-compensation cross-tabulation and the 291 voluntary tasks turn out to be mostly the 200 author-annotated ones, then "unpaid annotation grew 4.22x" is largely a restatement of "authors annotate more", not an independent finding, and the two bars on the second chart should be read as one.
Sources
- Maria Kunilovskaya, Gagan Bhatia, Lisa Sophie Albertelli, Yanran Chen, Christian Greisinger, Lotta Kiefer, Christoph Leiter, Subhadeep Roy, Tewodros Achamaleh, Muhammad Arslan Manzoor, Sebastian Pohl, Yufang Hou, Steffen Eger. Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025. arXiv:2606.02255, submitted 1 June 2026, v2 1 September 2026, v3 2 September 2026. All counts and percentages here are from Appendix F and Tables 1 to 3 of the v3 HTML rendering.
- ACL Rolling Review. Responsible NLP Research Checklist. Items D2 (recruitment, payment, fair wage) and D5 (annotator demographics). Introduced at NAACL 2022, derived from the NeurIPS 2021 checklist; moved from a separate PDF into the submission form in February 2024.
- Prolific. Participant payment rates. Absolute minimum £6.00 / $8.00 per hour, recommended minimum £9.00 / $12.00 per hour. Page last updated 13 March 2026.
- BLOMEGA. HTTP status checks against
github.com/NL2G/who-are-the-annotators(404) andgithub.com/NL2G(200, 23 public repositories), performed 10 September 2026. - BLOMEGA. Conditional rates, growth multiples and the column-sum check, computed from source 1's Appendix F. Method and code in Check it yourself.
BLOMEGA supplies consented, rights-cleared human data, with the annotator terms written down. Related reading: the 2026 annotation research roundup, consented AI training data providers, and data provenance and chain of title. Contact [email protected].