Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p> <bold>Background</bold> : Multi-agent collaboration systems are increasingly being applied in the medical field, yet a systematic synthesis of the overall research landscape in this area remains lacking. <bold>Objective</bold> : To systematically characterize the application landscape of multi-agent collaboration research in medicine, including publication trends, task type distribution, quantitative effects in medical question answering (MedQA) tasks, methodological reporting quality, and cost characteristics. <bold>Methods</bold> : We systematically searched PubMed, Embase, Cochrane Library, Web of Science, and China National Knowledge Infrastructure from June 2018 to May 16, 2026. Studies applying multi-agent collaboration frameworks to medical or health-related domains were included. Bibliometric methods were used to analyze publication trends, country distribution, and journal impact. Meta-analysis was performed for MedQA tasks using odds ratio (OR) as the effect measure. We also conducted an exploratory study to replicate the comparison of economic benefits between multi-agent systems and single LLMs as described in the original paper. Reporting completeness was assessed using the Chatbot Assessment Reporting Tool (CHART) guideline. Cost data were extracted and summarized. <bold>Results</bold> : Of 49 included articles, publications grew rapidly from Q3 2024 to Q2 2026 (23 in 2025, 23 in 2026). China (34.7%) and the United States (32.7%) contributed most articles; most appeared in JCR Q1 journals, with npj Digital Medicine publishing the most (n=5). Meta-analysis of 30 comparisons from 8 studies showed multi-agent systems showed higher performance LLMs in MedQA (p &lt; 0.001). Paired subgroup analysis from the single study with item-level data confirmed the net advantage (p &lt; 0.001). CHART assessment revealed systematic reporting deficiencies: full prompts (37.0%), query date/location (0%), and pre-registration (2.2%). Cost-effectiveness varied by architecture; one study favored multi-agent systems, while another favored single models. <bold>Conclusion</bold> : Multi-agent collaboration significantly improves MedQA accuracy over single LLMs, a finding robust across sensitivity analyses. However, reporting transparency and reproducibility remain critically deficient; research is geographically concentrated in China and the United States; and cost-effectiveness is architecture-dependent. Future research should prioritize standardized reporting, prospective multicenter validation, and systematic cost-utility analyses. </p>

Show More

Keywords

multiagent reporting collaboration systems research

Related Articles

PORE

About

Connect