Abstract
<title>Abstract</title> <p>Repeated non-invasive prenatal testing (NIPT) measurements from the same participant can be distributed across training and test sets when records are split at random, potentially inflating apparent model performance. We evaluated this risk in a de-identified repeated-measure benchmark dataset comprising 1082 male-fetus records from 267 participant codes and 605 female-fetus records from 147 codes. The development-stage Bayesian-optimized XGBoost analysis produced a record-level R² of 0.794 for Y-chromosome fraction prediction and a record-level ROC-AUC of 0.952 with an F1 score of 0.931 for female-fetus aneuploidy-label classification; BMI-stratified probability curves also indicated later candidate testing windows at higher BMI. To assess robustness to repeated measurements, we conducted 30 repeated five-fold experiments comparing record-level cross-validation with participant-grouped cross-validation for nonlinear regularized regression, a 4% fetal-fraction threshold model, and aneuploidy-label classification, together with a participant-mean identity-leakage stress test. Median Y-fraction regression R² decreased from 0.068 under record-level validation to 0.046 under participant-grouped validation, while the participant-mean predictor declined from R²=0.432 to -0.006. Threshold-model ROC-AUC decreased from 0.588 to 0.563, and aneuploidy-label ROC-AUC decreased from 0.677 to 0.666. These findings show that strong record-level performance may partly reflect participant-specific information and should not be interpreted as evidence of generalization to unseen participants. Participant-grouped validation should be considered a minimum requirement for repeated-measure NIPT modeling, and prospective external validation remains necessary before clinical use.</p>