Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Clustering is widely used to identify meaningful groups in survey data, yet standard validation metrics assess geometric coherence and not practical relevance. As a result, statistically well-separated clusters may offer little insight into real-world outcomes. We introduce Cross-Validated Prediction Accuracy (CVPA) as a metric that evaluates clustering solutions by testing how well cluster membership predicts independent demographic or behavioural variables in out-of-sample data. We apply CVPA to three survey datasets that span climate discourse, tidal literacy, and lived experience during COVID-19. Across cases, CVPA distinguishes clusters that are merely mathematically compact from those that capture a socially meaningful structure. We demonstrate that it is possible to obtain similar CVPA scores from computer science clustering methods and manual clustering performed by experts (0.609 vs 0.605, 0.571 vs 0.572, and 0.539 vs 0.532 for each dataset respectively). Furthermore, we show that transformer-based embeddings consistently outperform traditional lexical representations in preserving interpretable semantic structure. By linking cluster evaluation directly to predictive utility, CVPA reframes validation as a test of empirical relevance, providing a principled and domain-agnostic framework for computational segmentation in policy, public health, and social research.</p>

Show More

Keywords

cvpa clustering meaningful survey data

Related Articles

PORE

About

Connect