Comparative Evaluation of Ensemble Machine Learning Models for Predicting Band Gaps of Double Perovskites

Abstract

Accurate prediction of band gaps in double perovskites is crucial for accelerating the discovery of advanced materials for energy and optoelectronic applications. Traditional Density Functional Theory (DFT) approaches, while reliable, are computationally intensive and impractical for large-scale screening. To overcome these limitations, this study develops a Gradient Boosting Regression (GBR)–based machine learning framework that integrates polynomial feature expansion, standardized preprocessing, and rigorous hyperparameter optimization. The dataset consists of 4,121 double perovskite compounds with 39 composition-derived descriptors, including electronegativity, ionic radii, oxidation states, and orbital energies. The data were split into training and test sets using an 80:20 random split, and model performance was evaluated using multiple metrics. The GBR model achieved a coefficient of determination (R²) of 0.9573, adjusted R² of 0.9556, mean absolute error (MAE) of 0.0988 eV, root mean square error (RMSE) of 0.2180 eV, explained variance (EV) of 0.9574, and residual predictive deviation (RPD) of 4.8393, demonstrating high predictive accuracy and reliability. Compared to previous studies using Support Vector Regression, Random Forest, and XGBoost, the proposed framework provides superior performance while maintaining physical interpretability. These results establish a reproducible and efficient data-driven methodology capable of near-DFT-level predictions, facilitating high-throughput screening and rational design of double perovskite materials.

Loading publication timeline...

WhatsApp Chat