Back to Search View Original Cite This Article

Abstract

<sec> <title>BACKGROUND</title> <p>AI use in clinical decision making is increasingly frequent. How treatment recommendations generated by large language models (LLMs) compare with real-world treatment decisions remains poorly characterized.</p> </sec> <sec> <title>OBJECTIVE</title> <p>To assess whether AI-generated treatment recommendations are rated as appropriate as real-world MDT decisions in newly diagnosed lung cancer and to test the feasibility of a blinded comparative method.</p> </sec> <sec> <title>METHODS</title> <p>We designed a pilot study consisting of a blinded comparison of treatment recommendations made by four LLMs (ChatGPT, Claude, Gemini and Grok, accessed via consumer chat interfaces in June 2026) against real-world treatment decisions in a series of 27 consecutive, newly diagnosed lung cancer cases that received treatment in a single institution, from June 2024 to May 2026. A standardized prompt was used to ask AI to generate treatment recommendations. These were extracted in a standardized form, along with the actual treatment received in each case, totaling 135 treatment recommendations; a single blinded reviewer, presented with the recommendations in randomized order per case, graded all of them from 1 (inappropriate) to 5 (fully appropriate) and checked three flags (guideline-concordant, potentially unsafe, local/supportive treatment omitted). Scores were compared across sources using the Friedman rank-sum test.</p> </sec> <sec> <title>RESULTS</title> <p>The average score was highest for the real-world MDT decision at 4.52, followed by ChatGPT 4.48, Claude 4.44, Grok 4.41 and Gemini 4.37. No statistically significant difference was found (Friedman χ²(4)=1.99, p=0.74). No recommendation received a score lower than 3. Flags were similar across groups.</p> </sec> <sec> <title>CONCLUSIONS</title> <p>This pilot study did not detect a significant difference in rated appropriateness of treatment recommendations made by AI compared with the treatment actually received by patients. The methodology proved feasible but these findings are hypothesis-generating and require confirmation in larger studies, with more cases and multiple reviewers.</p> </sec>

Show More

Keywords

treatment recommendations realworld received decisions

Related Articles

PORE

About

Connect