Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Purpose To evaluate current multimodal large language models (LLMs) for detecting hepatocellular carcinoma (HCC) on CT and compare their performance with human readers under controlled conditions. Materials and Methods We retrospectively evaluated 211 portal venous phase CT images (106 HCC, 105 controls). Multimodal LLMs (ChatGPT-4o, ChatGPT-4.5, Perplexity, and ChatGPT-5.1) were tested under zero-shot conditions using a standardized prompt. ChatGPT-5.1 was additionally evaluated with brief anatomical guidance. Outputs were recorded as binary decisions. Diagnostic performance was compared with a senior radiologist and an abdominal imaging fellow. Statistical comparisons were performed using McNemar’s test with Bonferroni correction. Results The radiologist and fellow achieved the highest accuracy over multimodal LLMs for HCC detection on single-slice portal venous phase CT, reaching 0.915 (193/211) and 0.910 (192/211) respectively. Among multimodal LLMs, ChatGPT-5.1 with brief anatomical guidance performed best, achieving sensitivity of 0.792 (84/106), specificity of 0.876 (92/105), and accuracy of 0.834 (176/211). Its accuracy was significantly higher than other multimodal LLMs (all p &lt; 0.05) but lower than both human readers (both p &lt; 0.05). ChatGPT-5.1 without anatomical guidance achieved sensitivity of 0.670 (71/106), specificity of 0.810 (85/105), and accuracy of 0.739 (156/211), the addition of basic anatomical cues to the initial standardized prompt improved these values by 12.2, 6.6, and 9.5 percentage points, respectively. Main causes of error for humans and the best-performing multimodal LLM were isoattenuating HCCs for false negatives and vascular misidentification for false positives. Conclusion Among the multimodal LLMs evaluated, ChatGPT-5.1 with simple anatomical prompt augmentation achieved the highest accuracy for portal venous phase HCC lesion detection, but remained inferior to radiologists. Current multimodal LLMs demonstrated imperfect visual lesion detection capability on constrained tasks, and the performance was sensitive to prompt contextualization and anatomical orientation. Their use in imaging settings should be accompanied by meticulous prompt design, rigorous evaluation, clear characterization of error patterns and appropriate human supervision.</p>

Show More

Keywords

multimodal llms anatomical chatgpt51 prompt

Related Articles

PORE

About

Connect