Back to Search View Original Cite This Article

Abstract

<jats:p>Purpose: Federated learning (FL) is an established method to preserve privacy of training data in machine learning. Chemical data, often in a pharmaceutical context, have also been used for FL. So far, this has concentrated on per-molecule data. In this paper, we use per-atom data, specifically nuclear magnetic resonance (NMR) data, and examine the quality of the results obtained. Furthermore, we study how the number of overall data used for training influences the results. Methods: We use a deep-learning based neural network model to predict NMR shifts. We run this centralized and in a federated learning setting using FedN with random and skewed data distribution. For comparison, we also use a smaller neural network and HOSE codes. Results: We show that with low amounts of data, FL achieved better performance than centralized learning for the smallest datasets considered, and for large amounts of data, is almost as good as centralized learning. On the other 1 hand, we find constantly high variance for FL results, even with the full amount of data available. Conclusion: Federated learning is generally suitable for atom-wise data in chemistry. For low amounts of data, it can even significantly improve results. On the other hand, high variances in results show that there can be negative effects of the use of federated learning as well.</jats:p>

Show More

Keywords

data learning results federated centralized

Related Articles

PORE

About

Connect