Abstract
<title>Abstract</title> <p>Background Borrmann type IV advanced gastric cancer (B-IV AGC) is a diffusely infiltrative cancer that mimics hypertrophic gastritis on endoscopy, with superficial forceps biopsy frequently non-diagnostic, contributing to delayed diagnosis and poor prognosis. Existing endoscopic artificial intelligence (AI) studies have predominantly evaluated curated single-image datasets, an approach poorly suited to a diffuse infiltrative entity. We aimed to develop a whole-case patient-level AI decision-support model that interprets every available endoscopic image from each examination, and to compare it with expert endoscopists in a head-to-head reader study. Methods In this retrospective two-center diagnostic modeling study, we combined an endoscopy foundation model (GastroFM) with attention-based multiple instance learning (ABMIL), aggregating all examination images per patient without representative-image selection. Three cohorts (n = 4,860) were assembled: a primary B-IV AGC-suspicious cohort (1,137 B-IV AGC; 319 suspicious benign mimickers), a screening cohort (292 B-IV AGC; 3,033 benign gastritis), and an independent external validation cohort (47 B-IV AGC; 32 controls). Three expert endoscopists re-read the primary held-out test set blinded to the reference standard. Discrimination, binary diagnostic metrics, paired model–reader comparisons, calibration, and decision-curve analyses were assessed. Reporting followed TRIPOD + AI guidelines. Results At a prespecified probability threshold of 0.5, the model achieved AUROC 0.965 (95% CI 0.936–0.986), sensitivity 97.0%, and specificity 79.6% in the primary cohort, matching the majority-reader decision (p = 0.143) and outperforming each individual expert. The model identified 163 of 168 pathology-confirmed B-IV AGC cases — five more than the majority-reader decision (158) and 8 to 24 more than each individual expert (range 144–155). Discrimination was even stronger in the screening cohort (AUROC 0.998); external validation confirmed out-of-center generalization, although the modest sample size (n = 79) precludes saturated performance claims. Decision-curve analysis showed positive net benefit across clinically relevant thresholds, exceeding treat-all and treat-none strategies. Conclusions This whole-case AI model identified additional pathology-confirmed B-IV AGC cases beyond individual expert endoscopists, matched majority-expert performance, and produced clinically meaningful net benefit. Prospective evaluation as a diagnostic-escalation decision-support aid is warranted to reduce missed and delayed B-IV AGC diagnoses in routine practice.</p>