Abstract
<jats:p>Background: Machine-learning models for polycystic ovary syndrome (PCOS) and other conditions frequently report near-perfect diagnostic performance, but retrospective datasets assembled from routine clinical practice can encode diagnostic-group membership in how data were acquired rather than in disease biology, and this acquisition-related information can be indistinguishable from genuine clinical signal under conventional validation. Objective: To determine, using a real-world PCOS cohort as a case study, whether high classification performance reflected clinically meaningful information or artifacts of data provenance, schema structure, and measurement-acquisition workflow, and to develop a generalizable audit framework for detecting such artifacts in retrospective medical machine learning. Methods: We analyzed 1,331 retrospective records (1,286 PCOS, 45 controls) from a single endocrine-gynecology database. A layered acquisition-bias framework compared classification performance using (i) raw and harmonized missingness patterns alone, (ii) measured values with and without explicit missingness indicators, and (iii) ascertainment-balanced feature sets with and without age. Logistic regression and random forest were evaluated using repeated stratified cross-validation, bootstrap resampling, label-permutation testing, and calibration analysis, and the framework was validated against a semi-synthetic experiment with known ground truth. Results: Diagnostic status was perfectly predicted (ROC-AUC = 1.000) from missingness patterns alone, before any clinical value was examined, and this persisted after semantic harmonization of duplicated source columns. Performance declined progressively as acquisition-sensitive information was removed, from near-ceiling in raw and harmonized value models to a mean ROC-AUC of approximately 0.80-0.82 in the most restrictive ascertainment-balanced, age-excluded representation. The semi-synthetic experiment reproduced this pattern under known data-generating conditions, confirming that harmonization removes schema-fragmentation artifacts but not workflow-driven acquisition bias. Conclusions: Apparent diagnostic performance in this cohort was substantially attributable to diagnostic workflow and data-acquisition structure rather than to a stable, transportable biological signal. The layered audit framework generalizes beyond PCOS and offers a practical tool for detecting acquisition-related leakage in retrospective clinical machine-learning studies.</jats:p>