Abstract
<p>Researchers have begun to delegate extraction and judgement research tasks to tools built on large language models (LLMs), including the annotation of text, screening of papers, extraction of numerical values, and identification of supporting passages. Existing guidance establishes that these tools should be validated, documented, and tested for robustness. Researchers still need practical guidance, however, on how precisely different tools can be most appropriately evaluated. This paper provides a step-by-step workflow for evaluating such tools. This involves firstly defining the tool’s output, identifying the desired comparator, and specifying what would count as satisfactory performance – from these decisions, the appropriate analyses and requirements for comparison can be determined and conducted. I provide a worked example of this evaluation approach based on a simulated systematic review extraction-and-judgement tool.</p>