Abstract
<jats:p>Large language models (LLMs) can ease the work of screening titles and abstracts for systematic reviews, but obtaining reliable results requires researchers to make practical choices about which LLMs to use, how to combine their scores into a ranking, and how far down that ranking to read. We aimed to identify a general-purpose workflow that screens accurately, minimises human review effort, and generalises across environmental literature corpora. We ran an ensemble of five open-source LLMs across ten human-annotated systematic reviews from the field of ecology and environmental science spanning 19,777 studies. We then asked: (1) how well an ensemble of LLMs ranks relevant papers above irrelevant ones, and (2) where a human reviewer should stop working down that ranked list. A four-LLM ensemble chosen without any labels came close, on every review, to the best ranking achievable with that review's annotations (mean Average Precision 0.64 versus 0.66). We tested different rules for when to stop human review, finding the SAFE stopping rule recovered greater than 95% of relevant records on all ten reviews while requiring a human to screen 55% of the corpus on average. The paper offers a complete workflow that can be adopted for new, unlabelled reviews, using open-source LLMs small enough to run on a high-end consumer laptop, and we provide it as an open-source R package.</jats:p>