Abstract
<jats:p>Tandem mass spectrometry has become central to untargeted metabolomics. The translation of unknown spectra into biological insight depends on assigning chemical identities to detected metabolites. Structural characterization typically begins with mass spectral library matching, in which experimental spectra are compared against reference libraries and candidate annotations are ranked by their spectral similarity to the query. As spectral libraries and experimental datasets grow, however, more candidates achieve comparable similarity scores for a single query, and similarity scores give no indication of how reproducible a candidate match is or how sensitive it is to the underlying fragment evidence. Existing false-discovery-rate approaches can indicate annotation error at the dataset level but do not provide a per-match estimate of reliability. Here, we introduce a SpecReBoot-inspired query-focused bootstrapping approach that resamples the fragment evidence of each query spectrum. This approach relies on recomputing query similarity to candidate library spectra across bootstrap replicates, which provides a statistical distribution of scores rather than a single value. From this distribution we define the match support, a per-match reliability estimate quantifying the reproducibility of a match under spectral perturbation, together with measures of ranking stability that describe how often a candidate remains among the top-ranked matches across replicates. Applied to a forensic drug-of-abuse case, match support distinguished previously identified annotations from high-scoring false positives: a distinction cosine similarity failed to make. Furthermore, match support values remained stable as the reference library was expanded, whereas ranking stability metrics shifted significantly. In a cross-instrument endogenous metabolite library search, match support further revealed metric-specific annotation behavior, identifying metabolites consistently supported across different similarity metrics, while flagging annotations whose reliability depended strongly on the chosen scoring metric. Benchmarking against a natural-product reference library demonstrated that ranking based on match support values promoted true matches by four ranks on average compared with cosine-based ranking, without promoting analogs. Under controlled spectral perturbation experiments, match support flagged incorrect annotations with an AUROC of 0.75, whereas the cosine similarity score alone of the same match reached only 0.56. Query-focused bootstrapping thus provides a practical, per-match measure of annotation reliability, bringing the field a step toward reliable annotations at scale. We anticipate that incorporation of our annotation reliability scoring into computational metabolomics workflows will further promote the growth of spectral libraries and enhance their applicability across scientific disciplines.</jats:p>