Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Ultrasound interpretation requires clinicians to integrate visual evidence with anatomical context and diagnostic reasoning, yet both human interpretation and current vision-language models (VLMs) often remain vulnerable to variability in perception and reasoning pathways. Despite rapid progress in medical and ultrasound VLMs and their evaluation benchmarks, existing systems still rarely make the protocol--system--organ--diagnosis workflow explicit. Here we present SonoReasoner, an ultrasound VLM that formulates static ultrasound interpretation as hierarchical clinical reasoning. We construct SonoFlow-1M, a 1.03-million-image ultrasound corpus organized by clinical hierarchy, derive SonoCorpus, a 7-million-scale reasoning dataset for supervised reasoning initialization, and introduce SonoVQA, a clinician-curated benchmark comprising 1,503 cases and 14,905 question-answer pairs across six reasoning levels. SonoReasoner combines supervised reasoning initialization with group relative policy optimization (GRPO)-based clinical alignment, and is evaluated on SonoVQA-based clinical VQA, diagnosis classification, lesion localization and report generation against proprietary general-purpose, open-weight and medical-domain VLMs. Across these evaluations, SonoReasoner improves task performance while producing inspectable reasoning pathways grounded in anatomical context and sonographic evidence. These findings suggest that explicitly modeling ultrasound interpretation as hierarchical clinical reasoning can improve both performance and reasoning consistency in ultrasound VLMs.</p>

Show More

Keywords

reasoning ultrasound clinical interpretation vlms

Related Articles

PORE

About

Connect