Reassessing Maternal Health Risk Prediction: Duplicate-Aware Validation, Label Ambiguity, and Explainable Machine Learning

Authors

  • Ankit Garg
  • Khushboo Tripathi
  • Ashima Rani
  • Nishu Sethi
  • Anshu Malhotra

DOI:

https://doi.org/10.51483/IJAIML.6.3.2026.1011-1022

Keywords:

maternal health risk; machine learning; data leakage; duplicate-aware cross-validation; XGBoost; explainable AI; SHAP; reproducibility

Abstract

Maternal health risk classification is increasingly studied with machine learning, but benchmark performance can be misleading when repeated or dependent records are allowed to occur in both training and test folds. This study performs a duplicate-aware re-evaluation of the public UCI Maternal Health Risk benchmark using seven classifiers and explainable artificial intelligence. The analyzed snapshot contained 1,014 records with six physiological or demographic predictors and three risk classes. A data-quality audit identified 562 exact duplicate rows (55.42%), only 416 unique feature profiles, and 35 profiles with contradictory class labels; 215 records (21.20%) belonged to such label-ambiguous profiles. Conventional stratified cross-validation was therefore compared with profile-grouped stratified cross-validation that forced identical six-feature profiles into the same fold. Under repeated 3 × 5-fold group-aware validation, XGBoost achieved 0.662 ± 0.047 accuracy, 0.645 ± 0.053 macro-F1, 0.498 ± 0.071 Matthews correlation coefficient, and 0.844 ± 0.037 macro one-vs-rest AUC. In contrast, conventional cross-validation produced markedly higher macro-F1 for Extra Trees (0.857 vs 0.610), Random Forest (0.855 vs 0.634), and XGBoost (0.779 vs 0.645). The mean macro-F1 inflation was 0.247, 0.221, and 0.134, respectively. A Friedman test showed significant model differences under group-aware validation (p = 6.55 × 10^-4); XGBoost significantly exceeded Extra Trees and logistic regression after Holm correction but not Random Forest. Out-of-fold analysis confirmed that the mid-risk class remained the most difficult (F1 = 0.404). SHAP and permutation analyses consistently ranked blood sugar and systolic blood pressure as the dominant predictors. Training-fold-only SMOTE produced only a small macro-F1 increase (~0.007), indicating that duplicate dependence and label ambiguity, rather than class imbalance alone, are central limitations of this benchmark. The findings demonstrate that leakage-resistant validation and explicit data auditing are essential before claiming clinically relevant performance on maternal-health benchmark data.

Downloads

Published

2026-09-24

How to Cite

Garg, A., Tripathi, K., Rani, A., Sethi, N., & Malhotra, A. (2026). Reassessing Maternal Health Risk Prediction: Duplicate-Aware Validation, Label Ambiguity, and Explainable Machine Learning. International Journal of Artificial Intelligence and Machine Learning, 6(3), 1011–1022. https://doi.org/10.51483/IJAIML.6.3.2026.1011-1022