Explainable AI In Deep Learning: A Comparative Taxonomy And Evaluation Framework For Post-Hoc Interpretability Methods
Keywords:
explainable AI; interpretability; deep learning; SHAP; LIME; Grad-CAM; feature attribution; model transparencyAbstract
The issue of opacity in the decision making process is getting more relevant nowadays with the rapid deployment of deep learning models in the domains like healthcare, cybersecurity, and education. Explaining the decisions made by artificial intelligence models becomes more challenging with the increasing complexity of the tasks that need to be solved. Explainable Artificial Intelligence (XAI) emerged as a multidisciplinary research field that aims to create human-friendly explanations for predictions of sophisticated models without substantial loss in prediction performance. This paper presents an integrative literature review and a comparative framework for post-hoc XAI methods for deep learning. Using the taxonomy of XAI approaches that is based on differentiation into intrinsic and post-hoc explanations, local and global explanations, and model-specific and model-agnostic explanations, the paper highlights three main groups of methods: perturbation-based methods (e.g., LIME), axiomatic game-theoretic methods (e.g., SHAP), and gradient/attribution-based methods (e.g., Saliency Maps, Grad-CAM, Integrated Gradients, DeepLIFT, and Layer-wise Relevance Propagation). Comparative taxonomy chart, explanation pipeline for the three families of methods, and multi-dimensional qualitative comparison of selected methods are provided as a way of visualization of trade-offs regarding model-agnosticism, computational efficiency, theoretical background, stability of the explanation, and human-comprehensibility. Literature review also involves discussion of the applied domains of the methods, such as malware classification, medical imaging, and educational text classification. Open challenges in the area of XAI research are presented as follows: lack of a commonly agreed set of evaluation metrics, sensitivity of the explanations to small changes in the input data, and trade-off between fidelity and interpretability of the explanations.





