Abstract
<title>Abstract</title> <p>Background The pulmonary embolism rule-out criteria (PERC) are intended for patients with a low pretest probability of pulmonary embolism (PE), which makes retrospective application outside that setting difficult to interpret. We examined the classification performance of retrospectively reconstructed PERC and internally validated an interpretable model for the PE label recorded in an emergency-department VTE database. Methods We analyzed 1,297 adult emergency-department encounters (1,258 patients) recorded from June 2023 to May 2024 in a single-center VTE database. All positive labels were CTPA-confirmed; negative labels were assigned from recorded clinical diagnoses, without longitudinal follow-up. PERC was reconstructed from eight components. A ridge-penalized logistic model using 13 prespecified variables was compared with logistic models using PERC score, NEWS2, or both. Performance was assessed using 20 repeats of five-fold stratified group cross-validation, with encounters from the same patient retained in one fold. Results PE was recorded in 148 encounters (11.4%). Against the recorded outcome, reconstructed PERC had sensitivity 99.3% (95% CI 96.3–100.0), specificity 9.7% (95% CI 8.1–11.6), and negative predictive value 99.1% (95% CI 95.2–100.0). Across validation repeats, the ridge model had median ROC AUC 0.803 (2.5th-97.5th percentiles 0.795–0.806), median precision-recall AUC 0.496, and median Brier score 0.077. Using averaged out-of-fold probabilities, its ROC AUC exceeded the PERC-score model by 0.216 (95% CI 0.156 to 0.274). Adding NEWS2 to the PERC-score model slightly reduced discrimination (AUC difference − 0.015, 95% CI -0.027 to -0.003). Conclusions In this unselected emergency-department VTE database, reconstructed PERC identified nearly all recorded PE cases but rarely classified an encounter as negative. Its performance cannot be taken as evidence of rule-out safety because low pretest probability, uniform verification of negative labels, and longitudinal follow-up were unavailable. The ridge-logistic model improved internal prediction of the database-recorded outcome, supporting prospective multicenter evaluation with predefined patient selection and outcome ascertainment before clinical use.</p>