Abstract
<title>Abstract</title> <p> <bold>Background</bold> : Fully autonomous, FDA-approved artificial intelligence (AI) systems for diabetic retinopathy (DR) screening—LumineticsCore (formerly IDx-DR) and EyeArt—demonstrated high sensitivity and specificity in pivotal regulatory trials. Performance in real-world United States clinical practice, outside controlled trial conditions, has not been comprehensively pooled. <bold>Objective</bold> : To systematically review and meta-analyze the real-world diagnostic accuracy of fully autonomous, FDA-approved AI systems for DR screening in U.S. clinical settings, excluding pivotal and regulatory trial data. <bold>Methods</bold> : Database-restricted searches of PubMed/MEDLINE, Cochrane Library, and IEEE Xplore were performed (Embase was inaccessible). Eligible studies were U.S.-based, real-world evaluations of LumineticsCore/IDx-DR or EyeArt that reported sensitivity and/or specificity against a human-grader reference standard. Screening and data extraction were performed by a single reviewer. Risk of bias was assessed with QUADAS-2. Sensitivity and specificity were pooled using a random-effects (DerSimonian–Laird) model on the logit scale. A bivariate exploration, likelihood ratios, diagnostic odds ratio, sensitivity analysis excluding the largest study, and GRADE certainty assessment were also performed. <bold>Results</bold> : Five studies (total N ≈ 108,517 screening encounters/eyes) met inclusion criteria. After verification against source publications (one extraction error corrected), pooled sensitivity was 95.8% (95% CI: 88.2–98.6%; I² = 94.1%) and pooled specificity was 83.5% (95% CI: 74.6–89.8%; I² = 96.8%). Positive likelihood ratio was 5.82, negative likelihood ratio 0.05, and diagnostic odds ratio 115. Specificity ranged from 60.3% to 91.1% across studies. Excluding the largest study did not change the qualitative conclusions. GRADE certainty was rated Low for both outcomes. <bold>Conclusions</bold> : In real-world U.S. deployment, fully autonomous FDA-approved AI systems maintain high sensitivity but show substantial, setting-dependent variability in specificity. Pivotal-trial performance figures should not be assumed to generalize uniformly. Site-level validation remains important. </p>