Back to Search View Original Cite This Article

Abstract

<sec> <title>BACKGROUND</title> <p>In multi-agent medical diagnostic systems, diagnostic tasks are distributed across specialized agents and connected through inter-agent information sharing. Whether different inter-agent information-sharing conditions improve final diagnostic performance or primarily alter intermediate diagnostic outcomes remains unclear.</p> </sec> <sec> <title>OBJECTIVE</title> <p>To determine whether different inter-agent information-sharing conditions affect final diagnostic performance and intermediate diagnostic outcomes within a standardized multi-agent diagnostic workflow.</p> </sec> <sec> <title>METHODS</title> <p>We conducted an offline repeated-measures evaluation of 297 cases from three datasets using three large language models (LLMs). Each case–model pair was tested under 11 conditions: the single-agent baseline (SA); complete structured sharing (C1); structured sharing with explicit reasoning (C2); controlled field-level sharing (C3); minimal R3 coordination (C4); and six exploratory agent-role ablation conditions (A1–A6), yielding 9,801 case–model–condition evaluations. Prespecified analyses compared C1–C4 with SA and evaluated C2 versus C1, C3 versus C1, and C4 versus C3; C3 served as the reference for A1–A6. Clinical equivalence of the R2 candidate diagnoses, R5 final diagnosis, and final top-3 diagnoses was independently assessed by two physicians blinded to the base LLM and experimental condition, with disagreements adjudicated by a third physician. Accuracy was the principal outcome; top-3 hit rate, MRR@3, R2 differential diagnosis coverage, and token use were secondary outcomes; diagnostic error propagation and correction outcomes and agent-level output measures were exploratory.</p> </sec> <sec> <title>RESULTS</title> <p>C1 had the highest observed accuracy among C1–C4 (60.47% vs 59.36% for SA), but no multi-agent condition showed a statistically supported improvement in accuracy, top-3 hit rate, or MRR@3. Structured sharing with explicit reasoning (C2) increased mean tokens per case relative to C1. In exploratory analyses of agent-level output measures, controlled field-level sharing (C3) increased the R4 conflicting-evidence and coverage-gap count relative to C1 without statistically supported changes in diagnostic error propagation and correction outcomes or final diagnostic performance. Each agent-role ablation condition reduced mean tokens per case by 5,743–12,784 relative to C3 (all Holm-adjusted P=.02) without statistically supported differences in final diagnostic performance; removing R1 alone or with R3 also reduced the R4 count (both Holm-adjusted P=.01).</p> </sec> <sec> <title>CONCLUSIONS</title> <p>Within the evaluated datasets, base LLMs, prompts, and standardized multi-agent diagnostic workflow, the inter-agent information-sharing conditions altered selected intermediate outcomes and token use but did not produce statistically supported gains in final diagnostic performance. Exploratory agent-role ablation findings do not establish noninferiority, equivalence, or general dispensability of any role. Further evaluation in interactive clinical settings is required before extending these findings beyond offline diagnostic tasks.</p> </sec>

Show More

Keywords

diagnostic final sharing outcomes conditions

Related Articles

PORE

About

Connect