Abstract
<title>Abstract</title> <p>Generative artificial intelligence (AI)-powered simulated patients may expand access to adaptive and repeatable clinical training, yet evidence concerning their educational effectiveness, technical implementation, and responsible use remains fragmented. To characterize this rapidly evolving field, we conducted a PRISMA-ScR-guided scoping review and meta analysis of empirical studies published between 1 January 2016 and 1 June 2026, searching PubMed, OpenAlex, Scopus, Web of Science, and Springer Nature. Of 2,103 records identified, 157 studies met the inclusion criteria. Among these, 125 (79.6%) were published after 2025, indicating a marked recent increase in research activity on AI-powered simulated patients. Prelicensure medical education was the most common setting, represented in 67 articles, while applications focused primarily on history taking, communication, clinical reasoning, and formative assessment. Technical designs ranged from prompt-engineered conversational agents to multimodal, retrieval-augmented, and multi-agent systems, often incorporating automated feedback and human oversight. OpenAI GPT models were used most frequently, appearing in 86 articles, whereas Anthropic Claude, DeepSeek, Meta Llama, and Google Gemini were each reported in fewer than ten articles. Despite a growing number of randomized and nonrandomized comparative evaluations, most studies were short-term, single-site investigations focused on feasibility, learner experience, or immediate performance. Consequently, evidence regarding long-term retention, generalization to unfamiliar cases, transfer to clinical practice, and reproducibility across settings remained limited. Meta analysis yielded pooled estimates favoring generative AI-powered simulated patients for history taking and information gathering (Hedges' g=0.85, 95% confidence interval (CI): 0.38--1.32), communication and empathy (g=1.03, 95% CI: 0.15--1.91), OSCE global assessment and overall competence (g=0.77, 95% CI: 0.09--1.46), and clinical reasoning and diagnostic performance (g=0.93, 95% CI: -0.03--1.90). Safety, bias, privacy, assessment validity, and governance were frequently acknowledged, but empirical safety testing, incident reporting, equity analyses, and reproducible technical reporting remained uncommon. Generative AI-powered simulated patients therefore show promise for scalable, structured clinical practice and feedback. However, routine high-stakes implementation will require stronger longitudinal and multisite evidence, validated assessment approaches, transparent technical and safety reporting, and sustained educator oversight.</p>