Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p> Chest X-ray interpretation is a fundamental task in medical imaging research, yet reproducible benchmarking remains challenging due to weak labels, class imbalance, and potential data leakage. We present <bold>Cethraian-X</bold> , an open-source benchmark for multi-label chest X-ray classification using the NIH ChestX-ray14 dataset. The benchmark emphasizes methodological rigor through leakage-clean evaluation, multi-seed reproducibility, calibration analysis, explainability with Grad-CAM++, and comprehensive performance reporting. A DenseNet-121 baseline was evaluated across multiple random seeds using standardized training and evaluation protocols. Performance was assessed using macro ROC-AUC, macro Average Precision, calibration metrics, confidence intervals, and explainability analyses. Duplicate and leakage audits were performed to verify dataset integrity, and sensitivity analyses quantified the impact of removing identified duplicates. Results demonstrate stable performance across independent runs while highlighting persistent challenges for several disease classes and model calibration. Cethraian-X provides publicly available code, documentation, and reproducibility resources to facilitate transparent comparison of future chest X-ray classification methods. The benchmark is intended as a reproducible research resource rather than a clinical decision-support system and does not claim clinical validation or deployment readiness. </p>

Show More

Keywords

chest xray benchmark using calibration

Related Articles

PORE

About

Connect