Abstract
<jats:p> Perovskite oxides (ABO <jats:sub>3</jats:sub> ) are widely investigated for energy-related applications due to their compositional flexibility and tunable physicochemical properties. However, identifying formable and cubic perovskite oxides from the vast compositional space remains challenging using conventional approaches, especially for doped perovskite oxides. In this work, a two-stage machine learning framework was developed to accelerate the discovery of formable and cubic perovskite oxides in doped ABO <jats:sub>3</jats:sub> compositions. A dataset comprising 1,502 ABO <jats:sub>3</jats:sub> compounds was compiled from published literature. Compositional features derived from elemental properties were used to describe them, and feature selection was employed to identify the most relevant descriptors. It was found that perovskite formability and cubic structure were greatly determined by radius-based and volume-related descriptors, respectively. Compared with other algorithms, three ensemble learning algorithms, including Random Forest (RF), Extreme Gradient Boosting (XGB), and Light Gradient Boosting Machine (LGBM), displayed better performance. For formability prediction, XGB yielded the best results with an accuracy of 0.910 and an F1 score of 0.937, while for cubic structure prediction, it achieved an accuracy of 0.924 and an F1 score of 0.889. The trained models were subsequently applied to screen 710,269 virtual ABO3 compositions, identifying 505,698 and 217,596 candidates predicted to be formable and cubic, respectively, with 189,247 satisfying both criteria. This study provides an efficient framework for high-throughput screening and rational design of advanced perovskite oxides. </jats:p>