Five LLM runs agree at 0.995, and the measure correlates 0.029 with what it was built to measure
A study posted on 17 September 2026 scored European Commission AI Act consultation submissions with a large open model, five times over. The five runs agree at an intraclass correlation of 0.994 to 0.996. Against the survey answers filed by the same stakeholders, the same scores correlate at Pearson 0.029, 0.176 and 0.097, with intraclass correlations of 0.013 to 0.080 and a systematic offset of about two points on a five-point scale. The QA number almost every LLM annotation pipeline reports is the first one.
What changed, and when
On 17 September 2026 Veronika Batzdorfer (Karlsruhe Institute of Technology) and Carlo R. M. A. Santagiustina (Inria Paris Centre and medialab, Sciences Po) posted Reproducibility is not construct validity: LLM measurement of institutionally situated communication (arXiv:2609.19866v1, 28 pages).
The design is what makes it usable, and it is rare. The European Commission's AI Act consultation collected both a structured survey and a free-text submission from the same respondents, web-crawled from the Better Regulation Portal: N = 857 submissions across three rounds between 20 February 2020 and 27 April 2021, of which 348 (40.6 percent) carry both. That gives an independent measurement of the same construct, on the same unit, from the same source. Almost no LLM annotation study has one.
Qwen3.5-397B-A17B annotated each document on three continuous 0 to 5 scales: safety concern, rights concern and explainability trust. Five independent passes. Then the two measurements were compared.
The evidence table
Everything below is on the same 348 linked cases, except the inter-run block, which is over the full annotation set. Three coefficients, all bounded by 1, all measuring something different.
| Statistic | What it compares | Safety | Rights | Explainability | Source |
|---|---|---|---|---|---|
| ICC(C,1) | LLM run against LLM run, 5 passes | 0.995 | 0.996 | 0.994 | Appendix Table 8 |
| ICC(C,k) | the 5-run mean, same comparison | >0.999 | >0.999 | >0.999 | Appendix Table 8 |
| Pearson r | LLM score against the survey item | 0.029 | 0.176 | 0.097 | Table 1 |
| Spearman rho | LLM score against the survey item | 0.015 | 0.144 | 0.066 | Appendix Table 8 |
| Lin's CCC | same, penalising location and scale bias | 0.013 | 0.080 | 0.041 | Table 1 |
| ICC(2,1) | same, absolute agreement | 0.013 | 0.080 | 0.041 | Table 1 |
| ICC(3,1) | same, consistency | 0.027 | 0.146 | 0.093 | Table 1 |
| Cohen's d | standardised mean difference, LLM minus survey | 1.05 | 0.98 | 1.18 | Table 1 |
| Bland-Altman bias | raw offset, on a 0 to 5 scale | 1.98 | 1.79 | 2.15 | Appendix Table 8 |
| Bonferroni-adjusted p | for the Pearson correlation, alpha = 0.0167 | 1.0000 | 0.0031 | 0.2118 | Appendix Table 8 |
All values from arXiv:2609.19866v1, N = 348 for every row except the inter-run block. Limits of agreement are [-1.71, 5.67] for safety, [-1.78, 5.36] for rights and [-1.42, 5.72] for explainability, spanning roughly four of the five available scale points.
Only the rights-concern correlation survives Bonferroni adjustment, at r of 0.176 and adjusted p of 0.0031. Safety concern, the construct the whole exercise was aimed at, sits at r of 0.029 with an adjusted p of 1.0000. A discrimination check makes it worse: the LLM's combined safety-and-rights score correlates with the survey's explainability item at 0.162 and with its own target at 0.177. The measure is not separating the constructs it was prompted to separate.
Two comparisons, only one of which anyone reports
The reason the two numbers can be so far apart is that they are not two estimates of one thing. Inter-run ICC measures whether the model is a stable function of its input. It is a property of the model and the decoding settings, and at temperature-stable settings on a long document it approaches 1 almost trivially. It says nothing about what the function is computing.
Convergent validity measures whether that function tracks the construct. Nothing in the pipeline forces it to. A well-behaved, perfectly reproducible model can be measuring how much risk language a document contains while the survey measures how much risk the filer believes exists, and those are different quantities that happen to share a name.
The Bland-Altman numbers show this is not a threshold that could be calibrated away. The offsets are 1.98, 1.79 and 2.15 points on a 0 to 5 scale, close to or above one standard deviation in every case. Limits of agreement span about four of the five available points. You cannot subtract a constant and recover the survey measure, because the disagreement is not a constant.
The gap sorts by who is speaking
If the divergence were measurement noise it would be unstructured. It is not. Business associations voice more concern in the public consultation than in their survey answers by a standardised 1.00, 95 percent interval [0.68, 1.33] on 38 cases. Public authorities go the other way at minus 0.70, interval [-1.32, -0.07] on 15. Academics sit at minus 0.20, NGOs at minus 0.15, EU citizens at minus 0.22 and trade unions at minus 0.79, all with intervals crossing zero on small samples.
It sorts geographically too. Anglo-Saxon countries show modest public amplification at plus 0.29 over 65 submissions, Continental Europe sits at plus 0.01 over 163, Southern Europe at minus 0.49 over 36, the Nordics at plus 0.20 over 24 and Eastern Europe at minus 0.35 over 13. Divergence scores show positive spatial autocorrelation across countries, Moran's I of 0.347 at p equal to 0.036, computed with queen contiguity weights over 23 countries with 9,999 random permutations. Stakeholders in neighbouring countries diverge in similar directions.
This is the part that reframes the finding. Some of the gap is model error, and some of it is that a public submission and a private survey answer are different speech acts with different audiences. A validation protocol that cannot separate those two will keep attributing communicative strategy to model failure, or the reverse. The authors' own framing is that LLM text measures capture institutional role and communicative context as well as the nominal construct, and that treating the divergence purely as measurement failure throws away a real finding.
One thing survives the divergence intact: survey-reported concern stays strongly associated with support for explainability across all divergence terciles, at roughly 116, 115 and 116 cases each. Whatever is going wrong in the text-based measure, the underlying survey relationship is stable.
What this means for anyone running LLM annotation at scale
Stop reporting inter-run agreement as a quality metric. At 0.995, it is close to a measurement of your decoding settings. It belongs in a methods appendix as a sanity check that the pipeline is deterministic enough to reproduce, not in the results as evidence the labels are right. This is the same pattern we documented for LLM annotator agreement statistics: models agreeing with each other is cheap and tells you about shared priors.
Budget for a validation instrument, not just a validation sample. The usual defence is to hand-check a few hundred items. That measures agreement between your model and your own annotators, who read the same text under the same framing. What this study had instead was a second, independent measurement of the construct on the same units. Where an independent instrument exists, a survey, a behavioural outcome, a registry field, an audited decision, use it. Where none exists, say so, and treat the label as a measure of text features rather than of the construct.
Check discrimination, not only correlation. A proxy that correlates 0.162 with the wrong survey item and 0.177 with the right one is measuring a general salience dimension. Cross-construct correlations cost nothing to compute once you have the validation data and catch this specific failure, which a single per-construct correlation hides.
Expect the gap to be structured by who produced the text. A divergence of plus 1.00 for one stakeholder type and minus 0.70 for another means an LLM-derived score pooled across sources carries a source-dependent bias, not a uniform one. Any downstream comparison across institution types, countries or platforms inherits it. Our judgement: stratifying validation by document source should be the default, and it almost never is.
Two limits worth stating. This is one consultation, one construct family and one model, and the authors say so. The extent to which the reproducibility-validity gap generalises is not established here. And Qwen3.5-397B-A17B may be Anglo-calibrated in ways that specifically miss Southern and Eastern European precautionary register, which would explain part of the regional pattern without any communicative-strategy story at all.
Check it yourself
The consultation data is public through the Commission's Better Regulation Portal, and the coefficient comparison is a few lines once you have two measurements on the same units. The point of the snippet is the contrast, not the absolute values.
open https://arxiv.org/abs/2609.19866 # Tables 1 to 4, Appendix Tables 5 to 9
# the source: European Commission Better Regulation Portal, AI Act consultation
# 857 submissions, 3 rounds, 20-02-2020 to 27-04-2021
# 348 (40.6%) linkable to structured survey responses from the same stakeholder
# model: Qwen3.5-397B-A17B, three 0-5 scales, 5 independent passes
# LLM annotation flags recorded across N = 4,285 annotations
python3 - <<'PY'
rows = [ # construct, inter-run ICC(C,1), Pearson r vs survey, ICC(2,1), Cohen's d, BA bias
("safety concern", 0.995, 0.029, 0.013, 1.05, 1.98),
("rights concern", 0.996, 0.176, 0.080, 0.98, 1.79),
("explainability trust", 0.994, 0.097, 0.041, 1.18, 2.15),
]
print("%-22s %8s %8s %8s %7s %8s" % ("construct","run-run","r vs","ICC(2,1)","d","offset"))
for name, icc, r, icc21, d, bias in rows:
print("%-22s %8.3f %8.3f %8.3f %7.2f %6.2f pts" % (name, icc, r, icc21, d, bias))
print("%-22s reproducibility is %.0fx the convergence" % ("", icc / max(r, 1e-9)))
PY
# construct run-run r vs ICC(2,1) d offset
# safety concern 0.995 0.029 0.013 1.05 1.98 pts
# reproducibility is 34x the convergence
# rights concern 0.996 0.176 0.080 0.98 1.79 pts
# reproducibility is 6x the convergence
# explainability trust 0.994 0.097 0.041 1.18 2.15 pts
# reproducibility is 10x the convergence
The ratios in that output are ours, not the paper's, and they are a presentation device: two coefficients on the same 0 to 1 range are not on the same scale in any deeper sense. The falsifiable content is in the six published coefficients, which are reproducible from the paper's tables and from the consultation data.
What would prove this wrong
The strongest alternative explanation is that the survey is the bad instrument. Self-reported concern on a structured questionnaire is not ground truth either, and if the survey items are themselves poor measures of the constructs, a near-zero correlation says less than it appears to. The paper cannot rule this out, and neither can we.
A dated prediction: by 31 December 2027, a replication using a second independent instrument on these same 348 stakeholders, such as an audited coding of the submissions by trained human analysts blind to the survey, will report LLM-to-human agreement above 0.5 while LLM-to-survey convergence stays below 0.2. That pattern would confirm the gap is a construct gap between what text expresses and what a filer believes, not model error. If the LLM-to-human agreement also comes back below 0.2, the model is simply reading the documents badly and the institutional-context interpretation is unnecessary.
A second falsifier: run the identical protocol with three different model families. If inter-run ICC stays above 0.99 for all of them while convergence with the survey varies widely, the reproducibility number is confirmed as uninformative. If convergence is stable across families and low, the constructs themselves do not transfer between a public submission and a survey response, which is a finding about the data rather than about LLM annotation.
Sources
- Batzdorfer, V., Santagiustina, C. R. M. A. Reproducibility is not construct validity: LLM measurement of institutionally situated communication. arXiv:2609.19866v1, 17 September 2026. Tables 1 to 4, Appendix Tables 5 to 9, Section 4.
- European Commission. Better Regulation Portal. Source of the 857 AI Act consultation submissions and the linked structured survey responses, three rounds between 20 February 2020 and 27 April 2021.
- Lin, L. I. A Concordance Correlation Coefficient to Evaluate Reproducibility. Biometrics 45(1), 1989. The CCC reported at 0.013 to 0.080 here.
- Bland, J. M., Altman, D. G. Statistical Methods for Assessing Agreement Between Two Methods of Clinical Measurement. The Lancet, 1986. The bias and limits-of-agreement analysis.
- BLOMEGA. When LLM annotators agree with each other and with nobody else.
- BLOMEGA. Data annotation research: the latest.