Abstract
<title>Abstract</title> <p>With the fast expansion of genomic datasets across large geographic areas, understanding fine-scale population structure is crucial for estimating allele frequencies and addressing stratification. Clustering approaches are widely used to identify such structures by grouping individuals based on genetic similarities. However, choosing an appropriate method from the wide range of available algorithms can be challenging. Here, we present a comparative study of clustering algorithms, examining how spatial stratification and different subsampling schemes affect clustering results. We focused on the impact of pre-processing genetic data, different dimensionality reduction strategies, and compared model-based (Mclust, FineSTRUCTURE) and graph-based (Leiden) clustering algorithms under both known and unknown number of groups. We found that the most accurate clustering algorithm depends on the strength of structure and the criteria for quantifying the accuracy: FineSTRUCTURE finds the most spatially coherent clusters but is computationally costly. Mclust is user-friendly and often provides the most accurate partitions, but when used outside of optimal setting it can give misleading results. Leiden is fast but highly sensitive to its hyper-parameters and data pre-processing decisions. Finally, uneven sampling of individuals biases both cluster inference and hence downstream geographic mapping of rare, localized variants. We offer practical guidance for choosing clustering pipelines according to study goals and the degree of population structure.</p>