Back to Search View Original Cite This Article

Abstract

<sec> <title>BACKGROUND</title> <p>Electronic health records (EHRs) contain Protected Health Information (PHI) that poses significant privacy and regulatory risks. Cloud-based large language model (LLM) solutions require data to leave institutional boundaries, raising legal and ethical concerns.</p> </sec> <sec> <title>OBJECTIVE</title> <p>This study aimed to develop and validate a fully local de-identification pipeline based on HIPAA Safe Harbor criteria, comparatively evaluating 11 open-source local LLMs on Turkish clinical text.</p> </sec> <sec> <title>METHODS</title> <p>Clinical documents from Turkish healthcare institutions (n=221, 488 pages) were digitized from PDF to Markdown via a vision-language model OCR pipeline and processed through a zero-shot de-identification prompt. A manually annotated gold standard of 11,768 PHI spans across 18 HIPAA entity categories was constructed and used for evaluation. Model performance was assessed using precision, recall, F1 score, and Matthews Correlation Coefficient (MCC).</p> </sec> <sec> <title>RESULTS</title> <p>Binary PHI F1 scores ranged from 0.56 to 0.96. Qwen3.5-27B-NonThinking achieved the highest performance (Binary F1=0.96, Macro F1=0.917, MCC=0.876), establishing itself as the Pareto-dominant model without extended chain-of-thought reasoning. Four models met the Excellent MCC threshold (&gt;0.75) required for HIPAA-regulated deployment. Structurally regular entities (EMAIL, IP, SSN, URL) were detected consistently across models, while contextually dependent categories (DEVICE, HEALTHPLANID, OTHERID) showed high inter-model variance. The domain-specific MedGemma-27B ranked last (F1=0.56, MCC=0.12), indicating that clinical pre-training does not inherently confer advantages in structured PHI detection. A "Performance Cliff" below rank 8 revealed that MCC collapsed while F1 remained superficially acceptable, underscoring F1's insufficiency as a sole evaluation criterion.</p> </sec> <sec> <title>CONCLUSIONS</title> <p>The developed prototype demonstrates the feasibility of privacy-preserving de-identification of Turkish clinical records using local large language models, offering a practical solution for institutions seeking to comply with data protection requirements while leveraging AI-assisted clinical research. The comparative evaluation of 11 models revealed significant differences in performance across identifier categories, highlighting the importance of model selection and language-specific post-processing in non-English clinical contexts. Future work should focus on fine-tuning models on annotated Turkish EHRs and expanding validation across diverse hospital settings.</p> </sec> <sec> <title>CLINICALTRIAL</title> <p>N/A</p> </sec>

Show More

Keywords

clinical model models turkish performance

Related Articles

PORE

About

Connect