Explainable Hybrid CNN–Vision Transformer Framework For Robust Moringa Leaf Disease Diagnosis Using U-Net Segmentation And Yolov8 Localization
DOI:
https://doi.org/10.51483/IJAIML.6.2.2026.79-110Keywords:
Plants, Agriculture, Result, Model, Disease, Model.Abstract
Moringa oleifera is a plant species that is economically, nutritively, and medicinally relevant, where its productivity is adversely impacted by leaf infections caused by bacteria and fungi. Early detection of such disease infection in plants is crucial for ensuring better yield and supporting precision agriculture; but the CNN models used for detecting these diseases have limited ability in capturing global context relationships and are not explainable in addition to showing poor performance in a variety of environmental settings. In this study, we propose a hybrid deep learning model for moringa leaves classification that makes use of semantic segmentation using U-Net, ResNet50, and Vision Transformer (ViT). The proposed model involves the process of image processing and semantic segmentation of disease-infected regions, followed by hybrid feature extraction that takes advantage of discriminative local spatial features extracted from the image by ResNet50 and long-range context features extracted by ViT. Moreover, YOLOv8 is used for detecting the diseased parts in the moringa leaves and the Grad-CAM based Explainable Artificial Intelligence (XAI) technique provides the visualization support to make the model more transparent and reliable. Experimental results show that the proposed model achieves 99.68% classification accuracy, 99.54% precision, 99.47% recall, 99.50% F1-score, 99.81% specificity and an AUC of 0.998 which are superior to ResNet50, VGG16, DenseNet121 and Vision Transformer alone. The visualization results of Grad-CAM have shown that the model is able to focus on the biologically meaningful regions in the moringa leaves affected by diseases.





