Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>A major challenge in hazard assessment of drug and material discovery is the identification of potentially toxic compounds. To assess potential toxicity of the compounds, different assays are conducted; however, they are usually costly and time-consuming. Therefore, machine learning (ML) is often used to estimate the effect of a compound on a specific endpoint based on molecular features. However, finding an appropriate ML model, compound representation, and dimensionality reduction (DR) method remains challenging. Therefore, we benchmark a wide range of ML algorithms, molecular representations, and DR approaches to identify effective combinations for predicting toxicity across Tox- Cast endpoints. In particular, we consider 52 hormone receptor activity assays (e.g., estrogen or androgen receptor), each having roughly 500-9,000 tested compounds. To characterize the compounds, we compare transformer-based embeddings, structural fingerprints (Morgan and MACCS), and physicochemical descriptors. For these features, we apply different DR techniques (i.e., principal component analysis, variance-based selection, mutual information, and minimum-redundancy-maximum-relevance). Finally, we train different ML models (i.e., random forest (RF), multi-layer perceptron (MLP), CatBoost, support vector machine (SVM), and TabPFN). Using a nested 5×4 cross-validation setup, we evaluate 98 model–feature–DR combinations per assay. Performance is assessed using the Matthews correlation coefficient (MCC), one of the most appropriate metrics for imbalanced data. Despite the fact that the winning algorithm of the toxicity prediction challenge Tox21 was a deep neural network (NN), our results show that RF trained on physicochemical properties without DR is the combination that is most often ranked as top 1 (in 12/52 assays). Moreover, it ranks among the top 5 out of 98 combinations in almost half (25/52) of the tested assays. When ranking models per assay, RFs significantly outperform all other models, achieving an average rank of 1.4 and an MCC of 0.207. SVM, Cat- Boost, follow with ranks around 3 and MCC values near 0.15, while MLP and TabPFN show the worst performance with ranks around 4 and average MCC values around 0.13. Given that tree-based methods also show a short runtime and high interpretability, we recommend employing them for toxicity prediction. Scientific Contribution: This study presents the first comprehensive benchmark of ML models, molecular representations, and DR methods for hormone receptor activity assays in ToxCast. In contrast to prior work, we systematically evaluate complete modelling pipelines with rigorous hyperparameter tuning and chemically-informed data splitting to avoid similarity bias, supported by extensive statistical analysis.</p>

Show More

Keywords

assays compounds toxicity models different

Related Articles

PORE

About

Connect