Abstract
<title>Abstract</title> <p> <bold>Background</bold> : Incomplete recording in longitudinal electronic health records may bias epidemiologic estimates when binary outcomes are measured repeatedly within hierarchical data structures. <bold>Methods</bold> : We extended the clustered logistic factor (CLF) multiple imputation framework to longitudinal periodontal charting data, a high-dimensional setting in which up to 192 binary site-level outcomes are nested within teeth, individuals, and repeated visits. Because periodontal examinations share statistical features common to other clustered clinical measurements such as eye-level, lesion-level, or joint-level outcomes, the framework may have broader epidemiologic applicability. A patient-level random intercept extension (CLF-RE) was developed to account for between-individual heterogeneity across repeated observations. CLF and CLF-RE were compared with multiple imputation by chained equations using logistic regression (MICE-logreg), multilevel binary imputation (MICE-2l.bin), and K-nearest neighbors (KNN) across simulations varying sample size (n = 50–200), missingness mechanisms, and levels of tooth/site dropout. <bold>Results</bold> : Incomplete records substantially biased downstream analyses. Under missing-completely-at-random scenarios, incomplete data attenuated associations of smoking, diabetes, and age with disease progression by up to fourfold; CLF-RE recovered most of this attenuation, reducing bias 3- to 15-fold relative to observed data. CLF, CLF-RE, and MICE-logreg performed comparably, achieving site-level discrimination (AUC 0.77–0.85) and reducing patient-level disease burden bias from -77 to -117 sites (observed data) to within ±7–13 sites after imputation. KNN performed poorly (AUC 0.57–0.72), whereas MICE-2l.bin was computationally infeasible beyond 100 individuals in highly dimensional settings. Applied to a Singapore longitudinal dental registry of 721 patients, incomplete periodontal charting underestimated the prevalence of mild-to-moderate disease burden by 7–11 percentage points. <bold>Conclusions</bold> : CLF provides a computationally scalable multiple imputation approach for longitudinal epidemiologic studies with high-dimensional clustered binary outcomes and may be broadly applicable to other electronic health record settings with repeated hierarchical measurements. </p>