Abstract
<jats:p>Even though solvent extraction has been the industry-standard method for reprocessing spent nuclear fuel for decades, ligand design and discovery for effective metal separations remain confined to lengthy and hazardous trial-and-error experimentation and could largely benefit from novel computational approaches. In parallel, the development of machine learning models for chemical discovery is fundamentally limited by the availability of high-quality, structured datasets derived from scientific literature. In this work, we evaluate large language model (LLM)-based approaches for automated extraction of ligand properties from solvent extraction literature, with a focus on structured, rule-constrained information identification. We compare prompt engineering and fine-tuning strategies using a curated dataset of 737 annotated paragraphs, supplemented by a synthetic dataset designed to probe generalization and mitigate overfitting. To address the challenge of evaluating semantically equivalent but textually variable outputs, we introduce an LLM-based grading framework that scores extraction quality on a continuous scale using domain-informed criteria. Fine-tuned models outperform prompt-only approaches, improving average accuracy from 0.739 to 0.958 (unadjusted), with an adjusted score of 0.784 after accounting for null entries. These results demonstrate that fine-tuned LLMs can reliably perform structured chemical information identification from scientific text, and the proposed evaluation methodology provides a practical and extensible approach for benchmarking text-based outputs in solvent extraction workflows for metal separations.</jats:p>