Back to Search View Original Cite This Article

Abstract

<p>Language models are increasingly used to produce quantitative judgements about organisations from publicly available text, yet such systems differ fundamentally in whether they consult external evidence at the point of judgement or draw only on knowledge encoded during training. This study compared two systems that estimate the same 24 dimensions of employee experience using the same base language model and differ only in how evidence is gathered: a multi-stage retrieval approach, assembling each score through staged web retrieval and sequential scoring over two to three hours per organisation, and a single-prompt approach, producing all dimensions in one call from parametric knowledge. Because the two share a base model, the comparison isolates the contribution of retrieval. Evidence came from two independent samples, comprising 35 large, well-known companies and 152 organisations spanning insurance, retail, healthcare, education and professional services. Both approaches proved reliable across repeated runs (ICC(A,1) = .84 to .97). Convergence between them was moderate, with per-dimension correlations averaging .79 and .63 in the two samples and rising to .81 and .72 after correction for attenuation, and was consistently stronger for an organisation's overall score than for individual dimensions (corrected r = .91 and .85). Performance varied systematically by dimension in both samples: agreement was highest on dimensions describing an organisation's public identity, such as brand, mission and rewards, and lowest on those describing internal working experience, such as the direct manager, career progression, and diversity and inclusion. Retrieval contributed most where public text carries least, suggesting that gathering evidence determines which parts of a construct can be measured rather than simply improving accuracy.</p>

Show More

Keywords

organisations evidence dimensions retrieval from

Related Articles

PORE

About

Connect