Abstract
<title>Abstract</title> <p>Background: Chest radiograph classification models can achieve useful predictiveperformance, but limited interpretability reduces their clinical transparencyand practical trustworthiness. This study evaluated whether clinically groundedconcept construction improves concept bottleneck modeling for thoracic diseaseclassification and compared label-free and supervised bottleneck formulations. Methods: We used the NIH ChestXray14 dataset of 112,120 frontal chestradiographs. Disease-specific concept sets were generated through a multi-stagepipeline combining expert-informed terms, automated concept expansion, andUnified Medical Language System application programming interface retrievaland filtering. Label-free concept bottleneck models were trained across multipleconvolutional and contrastive language-image pretraining backbones and textencoder configurations. A separate supervised EfficientNet-B1 concept bottleneckmodel was trained with text-derived concept prototypes. The performanceof the developed method was evaluated using training loss, validation loss, perclassROC-AUC, and macro ROC-AUC. Concept-set quality was assessed withcoverage, quality, diversity, relevance, and semantic diversity scores. Results: Unified Medical Language System-enriched concept sets consistentlyoutperformed simpler basic concept sets in the label-free setting. The best labelfreemodel achieved a macro mean area under the receiver operating characteristiccurve of 0.702, whereas the best basic concept-set configuration achieved 0.691.The supervised concept bottleneck model achieved a macro mean area under the receiver operating characteristic curve of 0.8196, approaching the correspondingnon-bottleneck baseline of 0.8402. The highest class-wise performance wasobserved for emphysema, cardiomegaly, pneumothorax, and edema. Qualitativeexplanation examples showed clinically meaningful concept activations and alsohighlighted opportunities for further concept refinement. Conclusions: Clinically grounded concept construction strengthens conceptbottleneck modeling for chest radiograph classification. Supervised bottleneckmodels currently provide the strongest balance between discrimination and interpretability,whereas label-free models remain promising when concept annotationsare unavailable. These findings support further work on concept refinement,external validation, and clinician-centered evaluation of explanation quality.</p>