Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p> Stroke is the leading cause of death and disability in Indonesia, yet clinical NLP tools for stroke triage support remain limited by the absence of domain-specific annotated resources for patient-generated health text. We introduce StrokeID-NER, a stroke-specific named entity recognition (NER) dataset for informal Indonesian patient queries, comprising 1,742 queries annotated with 6,765 entity spans across four clinically grounded entity types — Symptom, Diagnosis, Temporal, and Risk Factor — with inter-annotator agreement ( <italic>κ</italic>  = 0.857). We fine-tuned and evaluated five NER systems on our corpus, spanning monolingual and multilingual transformer models at two scales and a GPT-5.4 zero-shot baseline. Fine-tuned XLM-RoBERTa-large achieved the best performance (macro F1 = 0.7823), outperforming the zero-shot baseline by 11.4 points, with the largest advantage concentrated in Symptom and Risk Factor entities requiring recognition of informal Indonesian, Javanese regional terminology, and clinical context. Hapax rates of 72.0–86.9% across entity types and 68.3% cross-model false negative overlap indicate that performance gains are constrained by lexical coverage rather than model architecture, suggesting vocabulary expansion as the most effective path to improvement. StrokeID-NER, the annotation guidelines, and all trained models are released publicly at https://doi.org/10.5281/zenodo.20305007. </p>

Show More

Keywords

entity stroke clinical annotated strokeidner

Related Articles

PORE

About

Connect