Abstract
<title>Abstract</title> <p>This study develops artifacts to help researchers extract digital sustainability (DS) data effectively and efficiently from large textual resources such as sustainability reports, and code the data using established Information Systems literature frameworks addressing the challenges associated with manual data collection, processing, and coding. We adopt the design science research methodology to develop the artifacts. We identify the research problem and motivation and define the objective for the solution the artifact attempts to address. Then we describe the iterative process through which we design, develop, and demonstrate artifacts. Finally, we evaluate the artifacts by comparing data collected and coded via the newly designed semi-automatic process with data collected and coded manually by expert researchers. The results reveal that digital sustainability data can be extracted in greater quantity and in less time using text analytics supported by our DS dictionary. The results also reveal that our generative AI prompts coded DS data with a high degree of agreement with expert manual coded data. The study offers researchers a scalable process for building large digital sustainability datasets from textual sources. The artifacts help reduce manual effort in data extraction and coding while preserving the need for expert review and quality assurance. The study contributes to a novel design science approach that combines text mining and generative AI to support digital sustainability research. It provides validated process, dictionary, and prompt artifacts that enable scalable, theory-driven extraction and coding of digital sustainability initiatives from textual data sources.</p>