Abstract
<sec> <title>BACKGROUND</title> <p>Systematic reviews and meta-analyses (SRMAs) represent the highest bodies of evidence in clinical research, but the process is labor-intensive, averaging an estimated $141,194.80 and 11 months in the U.S.A. Large language models (LLMs) have emerged as tools to assist, and potentially automate, SRMA production. Reported LLM performance varies by task, model, and methodological design. Many studies evaluate only many-shot, task-specific performance requiring prompt engineering inaccessible to most researchers, limiting scalability. Conversely, fully agentic SRMA generation has proven unreliable, underscoring the need to identify performance bottlenecks requiring human intervention while automating the remaining tasks.</p> </sec> <sec> <title>OBJECTIVE</title> <p>To evaluate the performance of the GPT-5 Pro series across all major SRMA steps using a zero-to-few-shot pragmatic design, ensuring reproducibility for non-AI specialists.</p> </sec> <sec> <title>METHODS</title> <p>Full data for the search string, article screening, data extraction, risk of bias (ROB) assessment and data synthesis steps of 5 in-house SRMAs were extracted and processed. From October 2025 to February 2026, we prompted the GPT 5 Pro series through the ChatGPT interface, to perform each aforementioned SRMA step, given the original author data from the previous step. All prompts were designed using simple, open-sourced templates. For search string, article screening, and data synthesis steps, GPT outputs were compared to original author data. For data extraction and ROB assessment, GPT outputs were compared to a novel GPT-human hybrid reference standard. For each step, a review of human and LLM performance literature was undertaken to better contextualize results.</p> </sec> <sec> <title>RESULTS</title> <p>GPT’s search string sensitivity in capturing relevant references to be included in the final SRMAs was 85% [80-88] CI95%, landing within the range of reported human performances. GPT’s abstract screening sensitivity and specificity were respectively 80% [76-84] CI95% and 85% [84-86] CI95%, roughly comparable to humans. However, GPT’s full-text screening underperformed in sensitivity compared to humans, 46%, [41-52] CI95%, while maintaining a similar specificity, 99% [99-99] CI95%. GPT outperformed humans in data extraction and ROB assessment accuracy, respectively, 93% [90-95] CI95% and 83% [79-86] CI95%. Data synthesis’ meta-analytical results were meaningfully similar between GPT and original authors, but not always strictly identical. All tasks were completed 10 to 100 times faster by GPT than by its human counterparts.</p> </sec> <sec> <title>CONCLUSIONS</title> <p>This study has limitations inherent to its zero-shot design, which often underperforms in highly “prompt-dependent” tasks, such as search string generation and article screening. Furthermore, in the absence of a better reference standard for evaluating GPT in certain steps, outputs were compared to original author data, which cannot be considered the gold-standard. In a zero-to-few shot pragmatic design, the GPT 5 Pro series performs comparably to humans in search string generation, abstract screening and data synthesis, while underperforming in full-text screening and outperforming humans in data extraction and ROB assessment.</p> </sec> <sec> <title>CLINICALTRIAL</title> <p>Clinical trial number: not applicable.</p> </sec>