Abstract
<p>Ordinal data are widespread in psychiatric research, especially in the form of Likert scales, but can be challenging to adequately model due to their ordered and discrete nature. In this work, we investigate the role of model specification for ordinal data in large-scale psychometric applications, focussing on the methodology of model-based clustering. Methods deliberately designed for ordinal data (ordinal likelihood) demonstrate the undesirable behaviours of instability and computational expense in large-scale data sets, while in comparison other more generic methods (e.g., Gaussian or multinomial likelihoods) offer preferable scalability and usability in the applications considered. In two “big data”-scale UK biobank data sets, using the PHQ-9 and a CIDI-SF-derived symptoms of depression, we examine the robustness of model-based clustering to likelihood choice (Gaussian vs Multinomial vs Ordinal). We further pursue an extensive, systematic simulation approach to understand the data properties driving the computational behaviour. Our results simultaneously uncover clinically relevant depression subtypes, and illuminate the pragmatic decision-making landscape spanning theoretical well-specification and practical convenience in two real-world large-scale applications and extensive simulations. The findings further highlight the demand for more robust and scalable tools designed for unsupervised learning of large-scale ordinal data.</p>