Beyond Post-Hoc Interpretability: Active SHAP-Guided Dynamic Ensemble Routing For PDF Malware Detection

Authors

  • Deepika Sharma
  • Manoj Devare

Keywords:

PDF malware detection, ensemble learning, SHAP explainability, feature subspace decomposition, dynamic weighted fusion, machine learning security.

Abstract

Rogue PDF attacks are considered one of the biggest threats to organizations since adversaries use structural misdirection or payload dispersal to bypass detection. The identification of sophisticated evasive behavior cannot rely on traditional machine learning due to their static reliance on weighted ensembles. Furthermore, the binary classifier that is built on such frameworks is vulnerable to the semantic gap. This paper presents the Explainability-Guided Feature-Group Aware Weighted Ensemble Learning (EG-FGA-WEL) framework which incorporates SHapley Additive exPlanations (SHAP) at the response with inference instead of post-hoc interpretability. With the EG-FGA-WEL, the features of PDFs are separated into the behavioral and statistical subspace, which also addresses the issue of multicollinearity. In addition to that, SHAP pruning is used to eliminate noise in the dimensions of the subspace. With regard to this current research, when local structural anomalies or indicators of evasion are present, instance-specific SHAP values alter the predictive trust of the behavioral and statistical subspace estimators. The results of the framework evaluation on a real-world corpus of 7,976 PDFs demonstrated a held-out test accuracy of 96.68%, a ROC-AUC of 0.9949, and a false positive rate of 0.0025 (2 false positives out of 797 benign test samples). The dynamic routing effectively reduced false positives compared to a monolithic Random Forest baseline (4 false positives) while maintaining high malicious precision (99.73%).

Downloads

Published

2026-06-24

How to Cite

Sharma, D., & Devare, M. (2026). Beyond Post-Hoc Interpretability: Active SHAP-Guided Dynamic Ensemble Routing For PDF Malware Detection. International Journal of Artificial Intelligence and Machine Learning, 6(6s), 668–679. Retrieved from https://svedbergopen.com/index.php/ijaiml/article/view/740