Abstract
<p>The explosion in digitization of large-scale administrative data and advances in record linkage have opened up new possibilities for studying core topics in population research, including social mobility, early-life determinants of longevity, and shifts in ethnoracial identity. Much attention has been paid to improving record linkage algorithms, but methodological guidance for researchers analyzing linked data is limited. We propose a general framework to explain how different types of linkage errors—false matches and missed matches—impact inference with linked data and introduce a bias-correction method. We then conduct three empirical case studies on canonical topics in population research: social mobility, the education–longevity gradient, and shifts in ethnoracial identification. These case studies demonstrate the settings in which linkage errors do and do not meaningfully impact substantive research findings. We conclude with practical recommendations for researchers performing inference with linked data and provide an accompanying R package implementing the methods introduced in this paper.</p>