Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p> <bold>Background:</bold> Generative AI chatbots are widely used for health information, but their performance in safety-sensitive cardiovascular queries is unclear. This study evaluated five publicly accessible chatbots on myocarditis-related questions involving urgent symptoms, post-infectious or post-vaccination concerns, exercise restrictions, prognosis, and sudden cardiac death risk, assessing safety, accuracy, empathy, reliability, transparency, quality, and readability. <bold>Methods:</bold> This cross-sectional, comparative, text-level evaluation assessed ChatGPT Plus 5.4, Gemini 3.0 Pro, DeepSeek-V3.1, Doubao, and Copilot. A total of 52 myocarditis-related public consultation questions were developed from Google Trends, online health forums, public question-and-answer platforms, and semi-structured patient interviews. Each question was submitted verbatim as a single-turn prompt in a new independent chat session between April 20 and April 26, 2026. Responses were evaluated against evidence-based reference standards derived from contemporary myocarditis guidelines, consensus statements, and systematic reviews. Five blinded independent raters assessed safety, accuracy, empathy, reliability, information quality, transparency, and readability. Safety was analyzed as a binary outcome; accuracy and empathy were scored on 5-point Likert scales; DISCERN, EQIP, JAMA benchmark criteria, GQS, and six readability indices were used for multidimensional assessment. Between-model comparisons were performed using Cochran’s Q test and Friedman tests, with Kendall’s W reported as effect size. Unsafe responses were additionally examined qualitatively. <bold>Results:</bold> Safe response rates were high across models (90.4%–94.2%), with no significant between-model difference in safety ( <italic>Cochran’s Q</italic>  = 1.077, <italic>df</italic>  = 4, <italic>P</italic>  = 0.898). Qualitative analysis of 21 safety-flagged responses identified five recurrent mechanisms of potential unsafety: under-triage of serious cardiopulmonary symptoms, over-reassurance regarding vaccine-associated myocarditis, premature or insufficiently individualized exercise guidance, incomplete treatment framing, and ambiguous escalation advice. Accuracy differed significantly across models ( <italic>P</italic>  &lt; 0.001; <italic>Kendall’s W</italic>  = 0.582), with Gemini, Copilot, and DeepSeek achieving the highest median scores. Empathy also differed significantly ( <italic>P</italic>  &lt; 0.001; <italic>Kendall’s W</italic>  = 0.180), with DeepSeek performing best. Significant between-model differences were observed for DISCERN, EQIP, JAMA, GQS, and all readability indices (all <italic>P</italic>  &lt; 0.001). DeepSeek performed favorably in DISCERN and EQIP, Gemini, Copilot, and DeepSeek achieved the highest global quality scores, and Copilot had the highest JAMA score; however, transparency remained limited across models. No model met the prespecified sixth-grade readability target for patient-facing information. <bold>Conclusion:</bold> In this controlled text-level evaluation, most chatbot responses were rated as safe, but unsafe or potentially misleading outputs occurred across all models, especially in scenarios involving chest pain triage, vaccine-associated myocarditis, treatment framing, and return to exercise. These tools may support general patient education but should not be used for autonomous triage, diagnostic exclusion, treatment decisions, or return-to-exercise clearance without clinician review and plain-language adaptation. <bold>Clinical trial number:</bold> not applicable. </p>

Show More

Keywords

readability safety accuracy empathy copilot

Related Articles

PORE

About

Connect