Abstract
<title>Abstract</title> <p>OSS platforms like GitHub serve as a primary data source for MSR research. As the platform is widely used by different users, spanning from student to developer, not all repositories are actual engineered projects. Therefore, to avoid such noise, researchers often apply several criteria, which may fundamentally change the sample demographics. To understand such biases, this study aims to uncover the hidden cost originating from these arbitrary thresholds or criteria. We analyzed 1.57 million repositories from the SEART platform and constructed several datasets from the thresholds often applied in MSR research. We identify the maintenance bias in these filtering processes, which masks the true abandonment (73.42%) realities of OSS projects. Also, these strategies favor some ecosystems and governance styles. Moreover, the sampling strategy also distorts the relationship between variables, suffering from relational biases. Therefore, to avoid such biases, the researchers should shift towards stratified sampling and refining the criteria for noise detection.</p>