Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Objectives To evaluate ChatGPT's diagnostic accuracy in classifying Lumbar spine magnetic resonance imaging (MRI). Materials and Methods Data were collected retrospectively via questionnaire from patients who underwent Lumbosacral spine MRI between October and November 2025. ChatGPT, along with two radiologists serving as a dual reference standard, independently classified the appropriateness. Agreement was measured using weighted Cohen’s kappa. For binary analysis, the three-level American College of Radiology criteria were dichotomized as Appropriate or Inappropriate. Binary performance was evaluated through Sensitivity, Specificity, Positive Predictive Value, and Negative Predictive Value. Results 160 patients (mean age, 48 ± 15 years, male percentage, 58%). ChatGPT showed moderate agreement with radiologists on appropriateness (κw = 0.58 vs. Radiologist 1 (R1), 0.50 vs. Radiologist 2 (R2)). When both radiologists classified ChatGPT’s reasoning quality as high (≥ 4, 79% of cases), exact agreement was high (83.3% vs R1 and 90.4% vs R2) and opposite discordance on appropriateness decision was low (2.3% vs R1 and 4.3% vs R2). Binary classification (Appropriate vs. Inappropriate) demonstrated high sensitivity (85% with both Radiologists) and specificity (92% with R1 and 61% with R2), along with moderate NPV (57% with R1 and 64% with R2), showing a tendency toward conservative MRI use. ROC curves showed AUC of 0.891 (95% CI, 0.824–0.958) vs. R1, indicating good to excellent binary discrimination, and 0.751 (95% CI, 0.661–0.841) vs. R2, indicating fair to good binary discrimination. Conclusion ChatGPT demonstrated moderate agreement with radiologists on MRI appropriateness using ACR criteria. Higher radiologist-rated reasoning scores were associated with reliable ChatGPT decisions.</p>

Show More

Keywords

radiologists binary chatgpt appropriateness agreement

Related Articles

PORE

About

Connect