Reliable but Design-Sensitive: Instrument Uncertainty in LLM Annotation
LLM annotators can appear highly reliable when a single pipeline is repeated, yet produce different labels when researchers make equally defensible design choices. The authors tested seven LLMs across 12 task designs with three independent runs each, labeling 3,000 tweets for offensive language and hate speech.
Repeating the same model with the same task design yielded high agreement, with a median Fleiss' κ of 0.91. But when the task design changed for the same tweets, agreement dropped to a median Cohen's κ of 0.76 — a decline attributable to design choices rather than model instability.
Task design and model choice inflated the variance of estimated prevalence by factors of 76.7 for offensive language and 110.6 for hate speech, compared with sampling variance alone. Variation across LLM task designs reached 560–572 basis points, versus 270–331 basis points across five human instrument versions.
Confidence scores did not resolve the problem: they tracked repeated model outputs more closely than agreement with human labels, and grouping six tweets into a single prompt lowered mean offensive-language confidence by 660 basis points. The authors label this design- and model-driven variation "instrument uncertainty" and argue it can only be measured by comparing reasonable task designs — repeating one setup or relying on confidence scores cannot substitute for that test.