Abstract
<jats:p>Individuals increasingly use conversational AI systems for symptom guidance. Whether multi-turn interactions improve clinical triage standard alignment remains uncertain. We conducted a retrospective, cross-sectional evaluation of 255 cases from three physician-reviewed sources: clinically authored vignettes (n=39), and real-world emergency department (N=76) and nurse line cases (n=140). CGPTH single-turn generated triage recommendations after only receiving an initial symptom description, reflecting potential typical use. Second, CGPTH multi-turn simulated nurse triage by asking subsequent questions before generating triage recommendations. Against nurse line standards, 52.9% single-turn and 55.7% multi-turn use prompting agreed exactly; clinician adjudicated disposition agreement was 54.1% and 48.2%. Discordant case recommendations represented lower acuity against nurse triage (natural: 70.8% under-triage, P < 0.0001; multi-turn: 69.0%, P < 0.0001). These findings suggest conversational interaction does not ensure safe triage-disposition alignment. Alongside aggregate agreement, clinical AI systems evaluations for symptom guidance should measure ordinal distance from standards, error direction, and additional dialogue conditions that affect recommended care standards.</jats:p>