Back to Search View Original Cite This Article

Abstract

<p>Large language models (LLMs) are increasingly used to code open-ended survey responses, but their validity against expert annotation remains uncertain, especially outside English. We evaluated four open-weight LLMs for deductive coding of Russian responses from a psychological human-robot interaction experiment (N = 33): DeepSeek-R1-32B, Qwen3-32B, Gemma3-27B, and Mistral-Small-3.2-24B. Ten psychology-trained coders provided a graded expert reference for seven predefined codes. Each model was run under singular-code prompts, an exploratory conservative variant, and fixed-sequence prompts mirroring the expert memo, with ten repeated runs aggregated into graded code-presence estimates. Agreement was assessed with graded concordance, calibration, human-equivalence, and threshold-sensitivity diagnostics. The strongest configurations by concordance — DeepSeek-R1-32B with singular-code prompting (CCC = 0.71) and Qwen3-32B with fixed-sequence prompting (CCC = 0.69) — reached agreement with the expert aggregate comparable to the disagreement among individual experts, although both relied on reasoning mode, which was confounded with model identity. Several models over-attributed codes relative to experts, and prompt format and aggregation strategy materially affected agreement. We recommend validating LLM coding as a measurement procedure, with reported calibration and uncertainty, rather than as automatic expert replacement in small applied psychological datasets.</p>

Show More

Keywords

expert graded agreement models llms

Related Articles

PORE

About

Connect