On the same diagnostic rubric, frontier LLM judges agreed with physicians 28% to 68% of the time
GRAND-ROUNDS, released on 11 September 2026, pools 9,217 scores from 11 physicians. On its 19-point Landmark Diagnostic Cases, Claude Opus 4.6 agreed with physician scores on 28% of 185 test responses, GPT-5 on 30% and Gemini 3.1 Pro on 68%. Physicians agreed with each other on 67%. A Qwen3-32B judge fine-tuned on 2 to 46 physician-scored cases per task reached 61%, and no prompt-only judge matched physician agreement on all five tasks.
What changed, and when
On 11 September 2026 Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur, Byron Crowe, Anthony M. Pettinato, Aashna P. Shah, Adrian D. Haimovich, Liam G. McCoy, Daniel Restrepo, Ethan Goh, Jonathan H. Chen, Laura Zwaan, Katherine E. Goodman, Daniel J. Morgan, Raja-Elie E. Abdulnour, Adam Rodman and Arjun K. Manrai posted Scaling Clinical Judgment to Evaluate Medical AI (arXiv:2609.12822, v1).
The paper does two things. It harmonises physician grading from seven published studies into one benchmark, GRAND-ROUNDS: 9,217 physician scores over 5,250 response-rubric entries, with responses written by 160 clinicians and nine AI models, scored by 11 physicians across six tasks. Five tasks are used in the judge experiments: NEJM clinicopathologic conference (CPC) diagnosis scored with the Bond score, the Landmark Diagnostic Cases scored on a 19-point rubric, the Grey Matters Management Cases, NEJM Healer consultation notes scored with R-IDEA, and BIDMC emergency-department triage scored with the Bond rubric. It then tests eight LLMs as judges against those scores and trains a physician-calibrated judge, PrecepTron-32B, by LoRA on Qwen3-32B.
Agreement is defined as a judge score within 1 point of the physician score on the 0 to 5 Bond scale (CPCs, BIDMC ER), or within 10% of the normalised score elsewhere. The physician baseline is computed on the subset of test entries that at least two physicians scored independently. The secondary metric is quadratic-weighted Cohen's kappa.
The evidence table
Accuracy against physician scores on the held-out test set, with quadratic-weighted kappa in brackets, transcribed from Table 1 of arXiv:2609.12822v1. n is test entries per task (Grey Matters: 258 cases, 1,715 question-level entries). Physician n is the double-scored subset.
| Judge | NEJM CPCs n=669 | Landmark n=185 | Grey Matters n=258 | NEJM Healer n=248 | BIDMC ER n=719 | Source |
|---|---|---|---|---|---|---|
| Claude Opus 4.6 | 74% (0.56) | 28% (0.48) | 84% (0.88) | 82% (0.81) | 91% (0.77) | Table 1 |
| Gemini 3.1 Pro | 83% (0.62) | 68% (0.80) | 80% (0.88) | 78% (0.71) | 86% (0.73) | Table 1 |
| GPT-5 | 87% (0.65) | 30% (0.54) | 71% (0.86) | 75% (0.58) | 87% (0.72) | Table 1 |
| Gemma-3-12B | 87% (0.61) | 55% (0.66) | 17% (0.60) | 69% (0.42) | 89% (0.59) | Table 1 |
| Llama-3.1-8B | 75% (0.32) | 43% (0.38) | 44% (0.62) | 46% (0.12) | 75% (0.45) | Table 1 |
| Mistral-3-8B | 53% (0.38) | 54% (0.72) | 62% (0.72) | 63% (0.54) | 75% (0.53) | Table 1 |
| Qwen3.5-9B | 87% (0.62) | 49% (0.43) | 74% (0.83) | 65% (0.47) | 90% (0.65) | Table 1 |
| Qwen3-32B (base) | 82% (0.55) | 46% (0.62) | 72% (0.75) | 76% (0.54) | 87% (0.68) | Table 1 |
| PrecepTron-32B (LoRA, 2 to 46 cases per task) | 92% (0.71) | 61% (0.80) | 80% (0.78) | 81% (0.80) | 91% (0.60) | Table 1 |
| Physician vs physician | 92% (0.68) n=481 | 67% (0.92) n=115 | 95% (0.90) n=216 | 77% (0.83) n=241 | 88% (0.66) n=719 | Table 1 |
| Ensemble of GPT-5, Claude, Gemini | 83% | 44% | not reported | not reported | not reported | Results text, Fig. 3 |
Read the Landmark column once more. A 12-billion-parameter open model, Gemma-3-12B, agrees with physicians on 55% of responses; GPT-5 and Claude Opus 4.6 manage 30% and 28%. Then read the Grey Matters column, where Gemma collapses to 17% and Claude leads the field at 84%. The ranking of judges is task-specific, and the paper says so: "which tasks differs by model".
A few dozen physician scores per task buy more agreement than a larger model does
PrecepTron is a recipe, not a new architecture. Take a 20% training split by case, LoRA fine-tune Qwen3-32B on the physician scores for that task (2 to 46 unique cases), resample so every score level is balanced, and evaluate on the held-out 80%. The base model gained between 4 and 15 points on every task. On the Landmark Diagnostic Cases the calibration set was 93 physician-scored responses spanning just 2 cases, and it moved agreement from 46% to 61% (kappa 0.62 to 0.80). Balanced resampling mattered: in the ablation without it, NEJM Healer kappa fell from 0.72 to 0.51 at similar accuracy, which the authors read as collapse toward the most common scores. Prompting with the same physician examples did not substitute. Few-shot prompting with five scored examples took NEJM Healer from 76% to 67%, and an optimised prompt (GEPA) took Grey Matters from 72% to 63%. The model runs on a single 80GB GPU, which is the paper's case for keeping patient text inside a hospital.
The frontier models are not only far from physicians on some tasks. They are far from each other, and the direction of their bias flips by task. That is why a rubric in the prompt is not enough: the rubric does not carry the panel's calibration.
What it means for anyone running expert evaluation or buying expert labels
Do not pick an LLM judge by general reputation. On this benchmark, the best prompt-only judge is Claude Opus 4.6 on three tasks (Grey Matters, NEJM Healer, BIDMC ER), Gemini 3.1 Pro on one (Landmark), and on CPCs GPT-5, Gemma-3-12B and Qwen3.5-9B tie at 87%. Before a judge scores anything that matters, score a sample with your own experts and report judge-expert agreement next to expert-expert agreement. The paper's own checklist asks for exactly that.
The expert-label budget shifts from volume to calibration. Our judgement: 2 to 46 physician-scored cases per task is a small, specific purchase, and it moved agreement further than switching to a much larger model did. The work that remains human is writing the rubric, scoring the calibration set, and double-scoring enough items to know what "agreement" can be. GRAND-ROUNDS has 481 double-scored CPC entries out of 669 in test; Landmark has 115 of 185.
Accuracy and kappa can disagree, so report both. Physicians on Landmark agree 67% by the within-10% rule but at kappa 0.92. PrecepTron's BIDMC ER accuracy rose from 87% to 91% while kappa fell from 0.68 to 0.60. A vendor quoting one agreement number is quoting half the picture.
Denominators are not matched. Judges are scored on all test entries; the physician baseline uses the double-scored subset. The paper states this. It means "reached physician level" is a comparison across different samples, most visibly on Landmark (185 against 115).
Check it yourself
The data card says CC-BY-4.0. The data is gated.
# dataset metadata: license, last modified, gating
curl -s https://huggingface.co/api/datasets/tbuckley/GRAND-ROUNDS \
| python3 -c 'import json,sys; d=json.load(sys.stdin); print(d["gated"], d["cardData"].get("license"), d["lastModified"])'
# manual cc-by-4.0 2026-09-09T14:17:25.000Z
# code link given in the paper, and the project site
curl -s -o /dev/null -w "%{http_code}\n" https://github.com/2v/PrecepTron # 404 on 15 Sep 2026
curl -s -o /dev/null -w "%{http_code}\n" https://preceptron.net # 200
# gaps from Table 1 (judge minus physician baseline, percentage points)
python3 - <<'PY'
md = {"CPC":92,"Landmark":67,"Grey":95,"Healer":77,"BIDMC":88}
j = {"Claude Opus 4.6":[74,28,84,82,91],"Gemini 3.1 Pro":[83,68,80,78,86],
"GPT-5":[87,30,71,75,87],"PrecepTron-32B":[92,61,80,81,91]}
for k,v in j.items():
print(f"{k:16}", [a-b for a,b in zip(v, md.values())])
PY
# Claude Opus 4.6 [-18, -39, -11, 5, 3]
# Gemini 3.1 Pro [-9, 1, -15, 1, -2]
# GPT-5 [-5, -37, -24, -2, -1]
# PrecepTron-32B [0, -6, -15, 4, 3]
On 15 September 2026 the Hugging Face API reports gated: manual, meaning access requires approval, and the dataset-viewer API refuses unauthenticated reads. We did not request access, so we have not recomputed Table 1 from the scores. BIDMC ER case text is excluded from the release because it contains protected health information. The trained adapters are listed at huggingface.co/collections/tbuckley/preceptron. The Table 1 numbers are in the arXiv PDF, which is the only version posted (the HTML rendering returned 404).
What would prove this wrong
The claim is that prompt-only LLM judges are not interchangeable with a physician panel and that small calibration sets close most of the gap. It would be wrong if a prompt-only judge, given only the published rubric, matched or exceeded the physician baseline on all five GRAND-ROUNDS tasks at once.
A dated prediction: by 31 March 2027, no paper or leaderboard will report a prompt-only judge that meets the physician baseline on all five tasks of GRAND-ROUNDS as defined in arXiv:2609.12822 (92%, 67%, 95%, 77% and 88%). Grey Matters at 95% is the hardest bar, and the best prompt-only score today is 84%. One such report, with the test split as released, falsifies this.
Sources
- Buckley, T. A., Kanjee, Z., Brodeur, P. G., et al., Rodman, A., Manrai, A. K. Scaling Clinical Judgment to Evaluate Medical AI. arXiv:2609.12822v1, 11 September 2026. Table 1, Results, Methods, Code and Data Availability; PDF.
- GRAND-ROUNDS dataset card, Hugging Face. Metadata retrieved 15 September 2026.
- preceptron.net project site, and PrecepTron model collection. Retrieved 15 September 2026.
- BLOMEGA. Unpaid annotation tasks grew 4.2x in ACL papers, on how rarely annotation papers report adjudication and agreement.
- BLOMEGA. Data annotation research: the latest.