Abstract
<title>Abstract</title> <p>Integrating data from multiple sources, such as survey samples and administrative records, offers significant potential for enhancing statistical inference, yet it remains vulnerable to simultaneous misclassification and selection errors. Methodological research typically approach these errors in isolation, leading to a fragmented understanding of data quality. This paper addresses this issue by introducing a Bayesian modeling framework that captures a composite error that jointly accounts for misclassification and selection errors when estimating population proportions. Our proposed approach utilizes a binary latent class mixture model to estimate indicator validity and measurement error, which are then incorporated into a latent selection error estimator. We evaluate our model’s performance through a simulation study across varying levels of error and an empirical application using linked microdata from the Current Population Survey (CPS) and the American Time Use Survey (ATUS). Our results indicate that the model robustly identifies the composite error across all different scenarios. However, the successful decomposition into its selection and measurement components hinges on the use of valid indicators, relevant auxiliary variables, and informative priors, underscoring the necessity detailed data curation and the incorporation of subject-matter expertise to achieve a fully granular interpretation of bias when using multi-source data.</p>