BLOMEGA

Five machine labellers marked 0, 1, 40, 72 and 78 of the same 100 scenes positive. The human marked 9.

Lab note · 16 September 2026 · BLOMEGA

Five vertical slate columns of wildly different heights against one short amber column on a dark background, representing five machine label counts against a human count

A report posted on 12 September 2026 scored a rule-based detector and four language models against blind human labels on one Turkish corpus. On the scheme's hardest label, the five machines marked 0, 1, 40, 72 and 78 of 100 scenes positive; the human marked 9. Cohen's kappa against the human was 0.000, 0.015, 0.019, 0.027 and 0.185. Overall raw agreement for the same five runs was 74.7% to 86.3%, which is the number a dashboard would have shown.

What changed, and when

On 12 September 2026, Levent Bulut (independent researcher) posted Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus (arXiv:2609.13936v1, CC BY-NC-ND 4.0). The manuscript is dated August 2026 and is written as an empirical reliability report rather than a method paper.

The object under test is the Objective Projection corpus, 500 annotated Turkish-English scene pairs. Since version 7 every scene carries an applied_rules field written by apply_rules.py, a bilingual rule-based heuristic flagging six craft features: two prohibitions (explicit emotion labelling, simile) and four positive techniques (materialized metaphor, micro-focus, temporal anchor, atmosphere contradiction). Those flags had been used to describe the corpus and select examples. Nobody had checked whether a human agreed with them.

The report opens with a declaration of three conflicts of interest: the author designed the six rules, served as the sole human rater in Study 1, and owns the dataset whose annotation layer the result damages. One of the scored systems, Claude, also assisted in preparing the analysis scripts and the manuscript. We flag these because the report does, and because they bound what the numbers can carry: the scoring is deterministic and reproducible from published files, the interpretation is not disinterested.

Three studies ran. Study 1 (n=120) scored the detector against blind labels from the scheme's author. Study 2 (n=100, a scene set with no overlap with Study 1) scored the detector plus Gemini 2.5 Flash and Grok against an independent non-expert volunteer whose labels were locked before any machine ran. Study 2b re-ran Study 2 unchanged with Claude Fable 5 (High) and ChatGPT 5.5. A sixth system, Gemini 3.6, returned thirty identical label rows and was rejected under a pre-registered degenerate-output rule.

The evidence table

Study 2 and Study 2b scored five labellers against the same locked human reference on the same 100 scenes. Values transcribed from Tables 2 and 3 of arXiv:2609.13936v1. "H+" is the human positive count out of 100. Kappa is Cohen's kappa with present as the positive class and the human as reference; n/a means expected agreement equalled 1, so kappa is undefined.

RuleH+Detector
+ / κ
Gemini 2.5 Flash
+ / κ
Grok
+ / κ
Claude Fable 5
+ / κ
ChatGPT 5.5
+ / κ
Source
Emotion label00 / n/a0 / n/a0 / n/a0 / n/a0 / n/aTables 2, 3
Simile10 / 0.0000 / 0.0000 / 0.0000 / 0.0000 / 0.000Tables 2, 3
Materialized metaphor972 / 0.0151 / 0.1850 / 0.00078 / 0.02740 / 0.019Tables 2, 3
Micro-focus9681 / 0.0229 / 0.00882 / -0.070100 / 0.00092 / 0.296Tables 2, 3
Temporal anchor9982 / -0.01993 / -0.01895 / -0.017100 / 0.00098 / -0.014Tables 2, 3
Atmosphere contradiction440 / 0.0002 / 0.0516 / 0.02055 / 0.26942 / 0.184Tables 2, 3
Overall raw agreement74.7%75.7%86.3%81.0%84.5%Tables 2, 3

Read the bottom row, then the materialized-metaphor row. Grok posts the highest raw agreement in Study 2, 86.3%, and its kappa on the scheme's central feature is exactly 0.000 because it marked zero scenes present. ChatGPT 5.5 posts 84.5%, the highest figure in any of the three studies, at kappa 0.019. Claude Fable 5 reached 29% raw agreement on that single rule, the lowest cell anywhere in the report, by marking 78 scenes present against the human's 9.

Materialized metaphor: scenes marked present, out of the same 100 Grok Gemini 2.5 Flash Human rater ChatGPT 5.5 Rule-based detector Claude Fable 5 0 (κ 0.000) 1 (κ 0.185) 9 (reference) 40 (κ 0.019) 72 (κ 0.015) 78 (κ 0.027) 0255075100 scenes marked present (n = 100, identical scene set for every row) Source: arXiv:2609.13936v1, Section 6.1 and Tables 2 and 3
One written definition, one scene set, six raters. The detector and Claude sit above the human by a factor of eight; Grok and Gemini sit below it by an order of magnitude.

Study 1 is the control that makes the rest legible. It scored the detector against the scheme's own author on a disjoint set of 120 scenes, and the surface rules behaved exactly as a surface rule should.

Rule (Study 1, n=120)Human +Detector +TPFPFNPrecisionRecallκSource
Emotion label21221000.1671.0000.265Table 1
Simile222001.0001.0001.000Table 1
Materialized metaphor63955045130.5260.7940.004Table 1
Micro-focus11896942240.9790.797-0.032Table 1
Temporal anchor1201001000201.0000.8330.000Table 1
Atmosphere contradiction121751270.2940.4170.258Table 1

Simile reduces to a short list of Turkish function words (gibi, sanki, adeta) and scores kappa 1.000. Materialized metaphor, on the only near-balanced distribution in the whole report (63 present, 57 absent), scores 0.004 at 51.7% raw agreement. Overall raw agreement across all 720 cells in Study 1 was 81.5%. That single number would have passed any review.

A labeller that always says absent scores 91% on a 9-in-100 label

The arithmetic is not subtle and it is worth writing out, because it is the whole finding. Grok marked zero of 100 scenes as containing a materialized metaphor. The human marked 9. Grok is therefore right on 91 scenes and wrong on 9, which is 91% raw agreement. Cohen's kappa asks a different question: how much of that 91% would two raters have hit by accident, given how often each of them says present? A rater that never says present has an expected agreement of exactly 91% too, so the observed agreement buys nothing and kappa is 0.000.

Two ways to miss the same 9 scenes (materialized metaphor, n = 100) Grok: 0 marked present TP 0 FP 0 FN 9 TN 91 human present human absent said present said absent raw agreement 91% Cohen's κ 0.000 expected agreement is also 91%, so nothing is left over to score Claude Fable 5: 78 marked present TP 8 FP 70 FN 1 TN 21 human present human absent raw agreement 29% Cohen's κ 0.027 near-perfect recall, precision 0.103: a different rule, not a worse one Source: arXiv:2609.13936v1, Tables 2 and 3. Both labellers ran on the identical locked human reference.
Grok's 91% and Claude's 29% sit at opposite ends of a raw-agreement scale and at the same place on a kappa scale, which is the point.

Five of the six rules in Study 2 had human positive counts of 0, 1, 9, 96 and 99. On four of those, kappa carries almost no information and raw agreement is a report on the marginal distribution. Only atmosphere contradiction, at 44 of 100, sits in a range where agreement can be read at all. On that rule, and only that rule, two models are clearly above chance: Claude Fable 5 at kappa 0.269 and ChatGPT 5.5 at 0.184. The report records this as a correction to its own earlier Study 2 conclusion that machines could not detect the feature.

The design that produced these numbers is worth carrying over, independently of the corpus it was run on.

Three studies, two human references, two disjoint scene sets Objective Projection v7.2500 Turkish-English pairsapplied_rules fromapply_rules.py, 6 featuresCC BY-NC-ND 4.0 Study 1, n = 120human = scheme authorraw 81.5%, simile κ 1.000 n = 100, disjoint setseed 2026 draw from 180,seed 2027 shuffle, ids strippedindependent rater, labels locked Study 2detector, Gemini 2.5 Flash,Grok · raw 74.7 / 75.7 / 86.3% Study 2bClaude Fable 5, ChatGPT 5.5raw 81.0 / 84.5% Pre-registered rejection ruleGemini 3.6 returned 30 identical label rows and was discarded, not scored Re-run stability: ChatGPT reproduced 58 of 60 re-run labels, Claude 60 of 60. All runs via public web interfaces. Source: arXiv:2609.13936v1, Sections 2.2 to 2.5
The locked reference and the pre-registered rejection rule are the two cheap parts of this design, and they are what make the spread in Section 6.1 readable rather than arguable.

What it means if you buy, sell or ship an annotation layer

A corpus-level agreement number is not a quality claim. This study's best headline figure is 86.3%, produced by a labeller that returned zero positives on the feature the corpus exists to demonstrate. If a vendor or a dataset card quotes one agreement percentage, ask for it per label, alongside the positive count of the human reference for that label. Without the marginal distribution the percentage is not interpretable, and the report is explicit that its own five skewed rules make raw agreement misleading.

Machine labellers disagree with each other, not just with humans. On micro-focus, Gemini marked 9 scenes and Grok marked 82 against a human count of 96, from the same written definition and the same prompt blocks. Gemini and Grok agreed with each other on 85.7% of cells overall, which again is the marginal distribution talking. Our judgement: a pre-labelling step that silently switches model or model version is a silent change of label semantics, and should be versioned like a schema migration.

Consistency is not validity. ChatGPT reproduced 58 of 60 labels on re-run and Claude 60 of 60. Both were stable. Both were at chance against the human on the rule that mattered. A stability metric measures whether the labeller is reproducible; it says nothing about whether it is applying your definition.

Build a balanced slice before you evaluate. Four of the six rules here cannot be tested on this corpus at any sample size, because 96 to 100 percent of scenes carry them. The design lesson the report draws is the one to steal: sample the evaluation set for balance on the target feature, or accept that the feature is untested. This costs a sampling pass and it is the difference between an interpretable kappa and a decorative one.

Two humans, not one. Every number in the report is agreement with one particular person. That is enough to show the machines are not applying the same rule as each other, and not enough to say whether the rule is hard or broken. For a commercial annotation programme the same gap shows up as a dispute you cannot adjudicate, which is why double-scoring a slice is a cheaper purchase than it looks.

Check it yourself

The labels, prompt blocks and scoring scripts are in the evaluation/ directory of the dataset. Unlike many reliability claims, the central table can be recomputed without asking anyone for access.

# dataset metadata: gating, license, freshness, downloads
curl -s https://huggingface.co/api/datasets/leventbulut/objective-projection \
 | python3 -c 'import json,sys; d=json.load(sys.stdin); print(d["gated"], d["cardData"]["license"], d["lastModified"], d["downloads"])'
# False cc-by-nc-nd-4.0 2026-09-15T18:53:47.000Z 337   (checked 16 Sep 2026)

# the paper, and the table values quoted above
curl -sL -o 2609.13936.pdf https://arxiv.org/pdf/2609.13936   # Tables 1, 2, 3

# reproduce the kappa arithmetic behind Grok's 91% on materialized metaphor
python3 - <<'PY'
def kappa(tp, fp, fn, tn):
    n  = tp + fp + fn + tn
    po = (tp + tn) / n
    pe = ((tp+fp)*(tp+fn) + (fn+tn)*(fp+tn)) / (n*n)
    return po, pe, (po - pe) / (1 - pe) if pe < 1 else float('nan')

for name, cell in [("Grok            ", (0, 0, 9, 91)),
                   ("Claude Fable 5  ", (8, 70, 1, 21)),
                   ("ChatGPT 5.5     ", (4, 36, 5, 55)),
                   ("rule detector   ", (7, 65, 2, 26))]:
    po, pe, k = kappa(*cell)
    print(f"{name} raw={po:.3f}  expected={pe:.3f}  kappa={k:+.3f}")
PY
# Grok             raw=0.910  expected=0.910  kappa=+0.000
# Claude Fable 5   raw=0.290  expected=0.270  kappa=+0.027
# ChatGPT 5.5      raw=0.590  expected=0.582  kappa=+0.019
# rule detector    raw=0.330  expected=0.320  kappa=+0.015

Two caveats on what is recomputable, both stated in the report. The per-scene label files for Gemini 2.5 Flash and Grok were lost; their confusion counts in Table 2 were recovered arithmetically from the surviving kappa and agreement values, uniquely for eleven of twelve cells (Grok on atmosphere contradiction admitted TP=3 or TP=4, resolved to 3 by the original write-up). The detector's Study 2 labels are a reconstruction from the published applied_rules field through the published identifier mapping, and reproduce the original column exactly. The scene-level Gemini-Grok agreement figure of 85.7% cannot be recomputed and is reproduced from the original analysis.

We ran that snippet on 16 September 2026. It reproduces all four published kappas to three decimal places from the confusion counts alone: 0.000, 0.027, 0.019 and 0.015. The expected-agreement column is the part worth staring at. Grok's is 0.910, identical to its observed agreement, which is the whole of its 91%. Claude's is 0.270 against an observed 0.290, so 78 positive calls buy two points of signal over a coin weighted the same way.

What would prove this wrong

The claim we are making is narrower than the paper's: that on a judgement-heavy binary label with a skewed human distribution, raw agreement and self-consistency can both look healthy while chance-corrected agreement is zero, and that this is a general property of such evaluations rather than an artefact of one Turkish corpus.

A dated prediction. By 30 June 2027, a replication of the materialized-metaphor rule on the same 100 scenes with a second independent human rater, using the operational rewrite the report asks for in Section 8, will not produce a prompt-only language model at Cohen's kappa above 0.40 against either human. If someone posts that result, with the published scene set and a locked reference, this article's reading is wrong and the definition, not the machines, was the binding constraint. The second rater is the measurement that decides it, and no further model run substitutes for it.

Sources

  1. Bulut, L. Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus. arXiv:2609.13936v1, 12 September 2026. Tables 1, 2 and 3, Sections 2, 4.1, 6, 7 and 8; PDF.
  2. Objective Projection dataset card, Hugging Face. Metadata retrieved 16 September 2026: not gated, CC BY-NC-ND 4.0, last modified 15 September 2026, 337 downloads.
  3. Feinstein, A. R., Cicchetti, D. V. High agreement but low kappa: I. The problems of two paradoxes. Journal of Clinical Epidemiology 43(6), 1990. The first kappa paradox, which the report cites for its skewed-distribution reading.
  4. Pangakis, N., Wolken, S., Fasching, N. Automated annotation with generative AI requires validation. arXiv:2306.00176, 2023. Cited in the report as the task-by-task validation argument.
  5. BLOMEGA. Two audits of HH-RLHF disagree on how much of it is mislabelled, on what happens downstream when a label layer is wrong.
  6. BLOMEGA. Data annotation research: the latest.