53% of "objective" physical descriptions carry a significant gender association
The fairness guideline that tells a captioning model to write "a defined jawline" instead of "he" has never been tested on readers. It has now. Across 14,706 ratings from 304 annotators on 316 physical attributes, 53% of the attributes are rated as significantly more associated with one gender than the others. Sixteen language models asked the same question recover the pattern only partially, compress their ratings into the 3 to 5 band, and abstain disproportionately on the non-binary category.
What changed, and when
On 14 September 2026 Yingjia Wan, Lin L. Lin and Elisa Kreiss (UCLA Communication and Computer Science) posted How Humans and LLMs Read Gender into "Gender-Neutral" Physical Descriptions (arXiv:2609.16366v1), a COLM 2026 paper. It introduces GAPA, Gender Associations of Physical Attributes.
The guideline under test is widely adopted: when a system describes a person whose identity cannot be confirmed, it should avoid inferring categorical labels (gender, race, disability) and describe observable attributes instead. The stated reasoning is that the inference is handed back to the reader rather than asserted by the system. The untested half of that reasoning is whether the reader then infers nothing.
GAPA's 316 attributes come from three separate channels: LLM generation, human elicitation, and extraction from contemporary fiction. Each trial asked "How likely is it for someone to say that a woman / man / non-binary person has a given physical attribute?" on a 7-point scale. The design is cross-classified: a participant rates a given attribute for exactly one target gender, which removes the direct cross-gender comparison a within-subject design would invite. Each participant completed 55 trials, and 304 participants (152 female, ages 18 to 77, mean 41.36) passed at least four of five attention checks.
One arithmetic note the paper does not reconcile: 304 participants at 55 trials is 16,720 trials, but 14,706 ratings are reported, which is 15.5 per attribute-and-gender cell rather than the "average of 17" the paper states. The 17 matches the trial count, not the retained rating count. This does not change any result; it does mean the per-cell sample is closer to 15 than 17 when you are judging how much weight a single attribute's mean can carry.
The evidence table
Every result in the paper that separates the three gender categories puts them in the same order. Women are rated highest and predicted best, non-binary lowest and predicted worst, by humans and by models alike.
| Measure | Woman | Man | Non-binary person | Source |
|---|---|---|---|---|
| Dataset-level association, relative to the highest category | highest | beta = 0.26 (SE 0.04, z = 6.77) | beta = 0.43 (SE 0.04, z = 11.68) | Section 4.3 |
| Human rater reliability (leave-one-rater-out and ICC) | highest | intermediate | lowest | Figure 4, reported ordinally; no values given |
| Trained proxy predictor, Pearson r vs held-out human ratings | 0.807 | 0.717 | 0.590 | Figure 7 (overall 0.764, RMSE 0.633) |
| Zero-shot LLM alignment, 16 models | at or above the typical-rater baseline | weakest of the three | at or above the typical-rater baseline | Figure 5 (right), Section 5.3 |
| Worst model (GPT-OSS-20B), Pearson r | not reported separately | not reported separately | -0.043 | Section 5.3 (overall r = 0.173) |
| Abstention from answering | rare | rare | concentrated here | Figure 6c |
All values from arXiv:2609.16366v1. Cells marked "not reported separately" are not in the paper; we have not estimated them.
Two numbers frame the model results. Claude-Opus-4.6 aligns best with the averaged human ratings at an average Pearson r of 0.669; GPT-OSS-20B is weakest at 0.173, and its non-binary correlation is slightly negative at -0.043. Proprietary models do not separate cleanly from open-weight ones here, which the authors flag as a departure from earlier gender-bias benchmarks. Instruction-tuned variants beat their own base models on RMSE but not on Pearson r, meaning post-training moved the ratings closer in absolute terms without improving the ordering of which attributes are more gendered than which.
The attribute-level structure is not a single axis. Across all 316 attributes the correlation between the man-association and the woman-association is only r = -0.16 (p = .004), and the sign depends entirely on where the attributes came from: strongly negative for human-elicited attributes, weakly negative for LLM-generated ones, and positive for attributes pulled out of novels. Non-binary ratings correlate positively with both women (r = 0.48) and men (r = 0.31), so an attribute being common for a category does not make it diagnostic of that category.
That sign flip is the most operationally useful thing in the paper. If you sample your attribute vocabulary by asking people, you get a vocabulary that opposes the two binary categories. If you sample it from published fiction, you get one where the same attributes go up and down together. Any bias audit that draws its stimulus list from one channel is measuring that channel's sampling as much as the phenomenon.
The guideline does not delete the label, it moves it where you cannot audit it
The substitution is supposed to end the chain: no inferred label emitted, no inferred label received. The measurement says the chain continues at the reader, with a graded signal instead of a categorical one. "A full goatee", "an Adam's apple" and "broad shoulders" land in the man-distinctive group; "plump lips", "a defined waist" and "a graceful figure" in the woman-distinctive group. Nothing about that is surprising as folk knowledge. What is new is that it is measured, on a fixed vocabulary, with a reliability estimate attached.
The consequence for anyone building datasets is specific. A categorical label is auditable: you can count it, balance it, redact it, and compute a disparity over it. A distribution of gender associations induced in readers by a description is none of those things by default. The paper's answer is to make it measurable, by releasing a predictor trained on the human ratings (olmo2-7b-base with a regression head, r = 0.764 on the held-out split) and running it over character descriptions in LitBank's 100 pre-1923 novels.
The abstention finding is the one that bites an annotation pipeline directly. Refusals are not spread evenly across the three targets: they concentrate on the non-binary category, and almost entirely in instruction-tuned open models plus Gemini-3-flash, while the matching base models show little of it. Post-training made those models decline to answer for one group and answer for the other two. If an LLM is doing first-pass annotation on identity-adjacent data, that is a missingness pattern correlated with the protected attribute, which is worse than a wrong label because it is invisible in accuracy metrics.
What it means for annotation guidelines and LLM pre-labeling
"Describe, do not infer" is a guideline about the writer, not about the reader. It is still a defensible rule: refusing to assert an unverifiable claim about a person is different from a reader forming an impression. But if the stated goal is that no gendered information reaches the reader, GAPA says the goal is not met, and a guideline that claims it is will mislead the team following it. Write the rule as what it does: it removes an unverifiable assertion, it does not neutralise the description.
Do not let a model pre-label the category it refuses to rate. Sixteen models, one prompt, and the abstention lands on non-binary. In a human-in-the-loop setup where the model proposes and a person confirms, that pattern turns into systematically thinner coverage for one group, then into a training set that under-represents it, with no error signal anywhere in the pipeline. The check is cheap: count non-responses per class before you count accuracy.
Expect an LLM annotator to compress your scale. Model ratings cluster in the 3 to 5 band and rarely reach 6 or 7, while human ratings spread across the range. A pipeline that thresholds at "strongly associated" will find almost nothing if a model produced the ratings. This is the same failure we measured from a different angle in LLM judges and human disagreement: the model gets the ordering roughly right and the distribution wrong.
Sample your stimulus vocabulary from more than one channel. Human-elicited and fiction-extracted attributes produce correlations of opposite sign (-0.61 and +0.27). A single-source list is a single-source result. Three channels cost the authors more work and are the reason the paper can say which conclusions are stable.
Budget more raters for the categories people agree on least. Reliability is lowest for non-binary targets on both statistics the paper reports, and the trained predictor tops out at r = 0.590 there against 0.807 for women. If per-item agreement is your quality gate, a fixed number of raters per item will deliver an uneven quality floor across categories.
Check it yourself
The dataset and the trained predictor are both released, so the central claim can be re-tested on your own attribute list rather than taken on trust.
# dataset, code and the 316 attributes with per-gender means
open https://github.com/Yingjia-Wan/GAPA
# the released predictor of human gender associations
open https://huggingface.co/alisa-yingjia-wan/gapa-predictor-olmo2-7b
# the paper, with Figures 1 to 8
open https://arxiv.org/html/2609.16366v1
# the arithmetic note in "what changed"
python3 - <<'PY'
participants, trials, ratings, attributes = 304, 55, 14706, 316
print("trials offered :", participants*trials)
print("ratings retained :", ratings)
print("per attribute : %.1f" % (ratings/attributes))
print("per attribute-gender: %.1f (paper states ~17)" % (ratings/attributes/3))
PY
# trials offered : 16720
# ratings retained : 14706
# per attribute : 46.5
# per attribute-gender: 15.5 (paper states ~17)
To reproduce the abstention finding on your own model list, the operational definition is in Section 5.3: a response counts as an abstention when it declines the premise ("non-binary people can have any eye shape, just like anyone else") or ends without a rating, including the case where the model emits 1 as a refusal. That last case matters, because a refusal coded as 1 is indistinguishable from a real low rating unless you keep the raw text.
What we could not check: the per-ranking-pattern counts in Figure 2 (how many attributes fall into each of the six possible orderings of the three categories, and what share of each is significant) are rendered as a figure, and the extracted values do not map unambiguously onto the labels. We have left them out rather than guess.
What would prove this wrong
The result rests on one US-based annotator pool, one language, one 7-point elicitation, and attributes presented stripped of visual context. The strongest counter-evidence would be a replication where attributes are shown inside a full caption with a real image, and where the share of attributes carrying a significant gender association falls below the level you would expect by chance at p < 0.05. That would mean the isolation of the attribute, not the attribute, produced the association.
A dated prediction: by 31 December 2027, no published replication of the GAPA elicitation in another language or annotator population will report fewer than 30% of physical attributes carrying a significant gender association. A replication below 30% falsifies the generality we are reading into this.
Sources
- Wan, Y., Lin, L. L., Kreiss, E. How Humans and LLMs Read Gender into "Gender-Neutral" Physical Descriptions. arXiv:2609.16366v1, 14 September 2026, published as a conference paper at COLM 2026. Sections 3 to 6, Figures 1 to 8; HTML version.
- GAPA dataset and code: github.com/Yingjia-Wan/GAPA. Predictor: alisa-yingjia-wan/gapa-predictor-olmo2-7b.
- Bamman, D., Sims, M., et al. LitBank, the 100-novel corpus used for the scaled-up analysis in Section 6.2.
- BLOMEGA. An LLM judge matches the majority label, but predicts human disagreement worse than 3 voters.
- BLOMEGA. Data annotation research: the latest.