Back to Search View Original Cite This Article

Abstract

<jats:p>Background: Delayed Code Stroke activation contributes to worse outcomes in acute stroke. Emergency Department (ED) triage notes contain free-text clinical information that could enable automated, real-time pathway activation. We evaluated the diagnostic accuracy of a multi-pass large language model (LLM) pipeline for identifying patients meeting Code Stroke criteria from ED triage notes. Methods: A retrospective cross-sectional study was conducted at Monash Medical Centre, Melbourne, Australia. De-identified triage notes from 3,023 ED presentations over a one-month period (September-October 2023) were analysed. The pipeline applied sequential passes for translation, stroke symptom identification, mimic exclusion, baseline functional status, temporal window classification, and symptom resolution. Six locally deployed language models were evaluated. Performance was assessed against two reference standards: neurologist-labelled diagnosis and documented ED Code Stroke activation. Primary outcomes were sensitivity and specificity; secondary outcomes included PPV, NPV, and Gwet's AC1. Reliability of the neurologist reference standard was assessed by blinded independent re-review of a stratified random sample of 200 presentations by a second neurologist. Results: Of 3,023 presentations, 136 were neurologist-labelled positive. Agreement between the primary and a blinded second neurologist on a 200-note reliability sub-sample was almost perfect (raw agreement 95.0%, Cohen's K; 0.900, 95% CI 0.838-0.959). The cohort included 140 ED Code Stroke activations (median age 69, IQR 56-81 years), of whom 83 (59.2%) had confirmed stroke diagnosis. Sixteen patients (11.4%) underwent endovascular clot retrieval and 4 (2.9%) received thrombolysis. The best-performing model (Qwen 2.5 14B) achieved sensitivity 0.890 (95% CI 0.826-0.932), specificity 0.993 (0.989-0.996), PPV 0.858 (0.791-0.906), and NPV 0.995 (0.991-0.997). Pairwise McNemar testing demonstrated statistically superior overall accuracy for Qwen 2.5 14B over Llama 3.1 8B, Phi-4 14B, and Mistral 3 14B (all p&lt;0.001 after Holm correction), with no significant difference detected versus Nemotron-Nano-12B-v2 or Qwen 3 14B. Conclusions: A locally deployed language model demonstrates acceptable sensitivity and specificity for automated Code Stroke identification from free-text triage notes. Performance was comparable across the two best models, suggesting that capable open-weight models in this parameter range may be sufficient to proceed with ongoing internal testing and external validation. The pipeline operates without internet connectivity or model retraining on patient data, supporting feasibility for real-world ED integration.</jats:p>

Show More

Keywords

stroke code triage notes model

Related Articles

PORE

About

Connect