Developing an Explainable Artificial Intelligence–Based Model for Early Software Defect Prediction And Software Quality Improvement
DOI:
https://doi.org/10.51483/IJAIML.6.11s.2026.1313-1333Keywords:
explainable artificial intelligence; XAI; software defect prediction; software quality; machine learning; SHAP; XGBoost; software engineering.Abstract
Background: Software defects remain a major source of rework, maintenance cost, operational risk, and deterioration in software quality. Machine-learning-based software defect prediction can prioritize modules that are more likely to contain defects, but high-performing models are often difficult for practitioners to interpret. Explainable artificial intelligence (XAI) offers a mechanism for making predictions more transparent and actionable.
Aim: This study develops an integrated XAI-based framework for early software defect prediction and software quality improvement and evaluates software professionals’ perceptions of usefulness, explainability, trust, quality improvement, and adoption intention.
Methods: A two-component design is proposed. The technical component uses static and change-oriented software metrics to train Logistic Regression, Random Forest, and XGBoost classifiers. Data preprocessing includes missing-value treatment, duplicate control, stratified partitioning, class-imbalance management within training folds, hyperparameter tuning, and held-out testing. Performance is evaluated using precision, recall, F1-score, ROC-AUC, PR-AUC, balanced accuracy, and calibration. SHAP is used for global and local post-hoc explanations. The human-evaluation component uses a 34-item five-point Likert questionnaire plus eight demographic/professional items. For the worked research model, a synthetic sample of 300 software professionals is analyzed.
Results: XGBoost achieved the strongest held-out performance (accuracy 0.943, precision 0.936, recall 0.928, F1 0.932, ROC-AUC 0.970). SHAP ranked code churn, cyclomatic complexity, prior defects, coupling, and lines of code among the most influential features. In the survey, all multi-item scales demonstrated acceptable-to-excellent internal consistency (alpha = .82-.91). Explainability significantly predicted trust, and trust significantly predicted adoption intention. Early defect prediction and perceived usefulness significantly predicted perceived software quality improvement. The indirect effect of explainability on adoption through trust was positive.
Conclusion: An XAI-oriented defect-prediction workflow can combine predictive performance with interpretable evidence that supports developer review and quality decisions. The framework should be validated with real repository data and real practitioner responses before empirical claims are made.





