Abstract
<title>Abstract</title> <p>STIs (Sexually Transmitted Infections) and HIV (Human Immunodeficiency Virus) remain significant health issues globally because of their widespread occurrence, often asymptomatic development, and social and medical implications. In this paper, we present an innovative approach for predicting STI and HIV risk with user privacy achieved, using structured tabular datasets from the Demographic and Health Surveys Program based on the distributed CatBoost algorithm. The presented dataset contains 168,459 observations collected across eight different countries, with variables related to demographics, behavior, and attitudes toward STI and HIV status. Instead of the classical centralized approach, where all training is done on the entire dataset, our framework utilizes multiple independent training sessions of CatBoost models on different subsets of data with subsequent aggregation of results with a reliability-weighted method. For increased robustness, the framework also employs local ensemble modeling, validation-based threshold adjustment, and prediction confidence calculation. Evaluation of our model is performed using a Leave-One-Country-Out (LOCO) methodology. As observed from the experimental outcomes, the proposed approach shows high predictability in terms of classification of both HIV and STIs, yet with privacy-preservation achieved. For the prediction of HIV cases, the proposed framework obtained an average AUC score of 0.9850 and an average accuracy score of 0.9537. When it comes to STI prediction, the average AUC and accuracy scores obtained were 0.9396 and 0.9117, respectively. Further investigation through t-SNE analysis and feature correlation analysis highlighted heterogeneity in different countries and confirmed that the difference in performance at the level of each country is reasonable to be interpreted. In summary, the proposed framework can be considered as a viable and robust solution to the problem of prediction of HIV and STIs in a distributed environment, and thus a substitute for centralized ML models in the applications where user privacy is a concern.</p>