Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p> Genomic selection can accelerate rice breeding by predicting agronomic performance directly from genomic markers, but its adoption is limited by the black-box nature of many machine-learning models and the lack of tools that translate predictions into actionable parent-cross recommendations. We developed an explainable machine-learning framework using the publicly available 1k-RiCA SNP panel (353 <italic>indica</italic> rice accessions genotyped at 965 loci) to (i) predict flowering time (FLW) and plant height (PH) from SNP genotypes, (ii) classify accessions into subspecies subgroups, and (iii) rank all pairwise parent crosses by predicted offspring performance. Ridge regression, random forest, and XGBoost were evaluated for regression, while random forest and XGBoost were compared for subgroup classification. XGBoost produced the best FLW predictions (test R² = 0.69, RMSE = 2.52 days). For PH, a random forest model incorporating FLW as a covariate with univariate feature selection achieved the best performance (test R² = 0.55, RMSE = 5.94 cm). Random forest also achieved the highest subgroup-classification performance (76.7% accuracy; macro-AUC = 0.82). Grain yield was excluded because correlation analysis showed a negligible association between SNP markers and yield (r = 0.03), indicating strong environmental influence. SHAP analysis identified markers driving each prediction. All 62,128 pairwise parent crosses were scored and ranked to generate a prioritized breeding shortlist. The pipeline was deployed as an interactive web application supporting dataset upload, automated preprocessing, SHAP-based explanations, and ranked cross export, providing transparent, statistically robust decision support for rice breeding using mid-density SNP panels without whole-genome sequencing. </p>

Show More

Keywords

performance random forest rice breeding

Related Articles

PORE

About

Connect