Abstract
<jats:p>Plant volatiles have long been used as taxonomic characters, especially in chemotaxonomy. Exploring the utility of chemotaxonomy has been widely regarded critical for drug discovery, despite the knowledge that phytochemicals are often evolutionarily labile. Despite this, the use of chemotaxonomy has been restricted primarily due to the complex interpretations behind translating chemical characters into phylogenetically operational characters. In this paper, we propose a method that integrates phytochemistry, machine learning, and trait mapping to examine whether plant volatiles retain signals for classification across broad taxonomic ranks and which components of the volatile metabolites carry these signals. Leveraging a global dataset comprising 2,139 volatiles across 429 plant species, we trained classifiers on presence-absence data and molecular fingerprints that can predict species to their taxonomic ranks. Incorporating structural features improves performance, suggesting that plant volatiles may encode lineage-specific chemical signatures. Rather than single diagnostic markers, combinations of volatiles informed taxonomic predictions, indicating biosynthetic constraints within volatile clusters. Fingerprint-derived clusters showed a lineage-dependent pattern when mapped onto phylogeny. The proposed method provides an integrated approach to revisit chemotaxonomy and trait evolution to understand the chemical diversity across lineages. This approach can also support comparative chemical prediction, especially for understudied closely related plant groups.</jats:p>