Abstract
<title>Abstract</title> <p>Large language model (LLM) outputs shift with context, but how far a clinically minor detail can move them has not been measured against clinicians under controlled conditions. In this randomized controlled clinician-comparator study, 100 physician-validated cases appeared in paired versions, with or without a brief travel, exposure, or occupational cue suggesting a cue-linked lure diagnosis. The primary endpoint, cue-linked lure selection, was compared between 22 LLMs (92,000 usable responses) and 47 clinician participants (1,128 responses on 21 matched case families). Baseline LLMs selected the lure in 84.9% of matched cue responses versus 27.8% for clinician participants, an extra cue effect of 57.6 percentage points (95% CI 50.2-65.2); cue accuracy fell to 12.4% versus 54.6%. Base-rate guardrail prompting reduced lure selection to 28.4%; rare-or-serious prompting raised it to 93.2%. We present contextual cue susceptibility as a safety signal for clinical AI, especially for agentic systems acting on context with little human review.</p>