Back to Search View Original Cite This Article

Abstract

<jats:p>ChEMBL computes two per-activity quality annotations—data_validity_comment, which records that a measurement is unlikely to be correct, and potential_duplicate, which marks a probable duplicate—and has published both since ChEMBL release 15. We ask a question that a prior-art sweep dated 2 August 2026 found unasked: what fraction of those annotations survives the distribution path that benchmark and model builders actually use? Auditing ChEMBL_36 against a 2026-07-28 PubChem bulk snapshot, we census the ChEMBL → PubChem → bulk-extract path and find that the two flags die at different hops. data_validity_ comment survives deposition essentially intact—present on 51,006 of 51,015 assays that should carry it (99.9824 %), with all nine misses attributable to a single dated re-deposition event—and then reaches 0.000 % survival in bioactivities.tsv.gz, the summary file whose thirteen fixed columns encode no validity field. potential_duplicate never arrives at all: zero of 1,873,601 ChEMBL-deposited assays carry a duplicate column in any held PubChem artefact. The loss is not benign at the tails. The flag rate is 1.4394 % file-wide but 3.99 % across IC50/EC50/Ki/Kd/MIC, and its consequence for a consumer is concentrated where potency decisions are made: a bulk-file user ranking ChEMBL_36 IC50 data most-potent-first is handed 794 distinct compounds drawn from 322 assays and 224 separate publications in the first 1,000 rows, every one of them a row ChEMBL had annotated as suspect, and none of them carrying that annotation in the file the user pulled. ChEMBL’s range rule is why those rows are flagged; that the warning does not travel with them is the finding. Auditing derived datasets yields a three-category taxonomy—Filters, Inherits, and Unattributableof which the third is the least discussed and the most consequential: across 227,743,379 rows in three large public releases, 3.64 % retain a per-row link back to a source activity, leaving 219,458,520 rows whose flag status cannot be determined by anyone, including their own authors. That figure concerns per-row flag attribution specifically—not scientific quality, and not the conclusions drawn from these releases, which carry substantial provenance of other kinds—and it mixes evidence of two strengths, since one release declares the flag columns itself while the other two are reached by assay-grain joins reconstructed here. We make no claim that public bioactivity data is low quality, that ChEMBL failed, or that any named dataset produced wrong scientific conclusions; the finding concerns transit alone. We release flagaudit, an MIT-licensed instrument that reports whether a held dataset contains flagged rows—or reports that this cannot be determined, which is a distinct answer and is emitted as one.</jats:p>

Show More

Keywords

chembl which flag rows quality

Related Articles

PORE

About

Connect