An Imbalance-Aware Feature-Optimized Heterogeneous Ensemble for Robust Network Intrusion Detection
DOI:
https://doi.org/10.51483/IJAIML.6.11s.2026.1564-1571Keywords:
Network Intrusion Detection, CSE-CIC-IDS2018, Ensemble Learning, Feature Selection, RFE, Class Imbalance, LightGBM, XGBoost, FT-Transformer, Deep Learning.Abstract
Accurately classifying types of cyberattacks continues to be a major challenge for network intrusion detection systems. This is especially true when dealing with dimensional traffic data, imbalanced classes and attacks that change over time. To tackle these issues this study introduces an ensemble framework that is aware of class imbalance and optimized for feature selection. The framework is built to handle intrusion detection and uses five base models. LightGBM, XGBoost, Random Forest, a Deep Neural Network (DNN) and an FT-Transformer. These models work together through voting to make the final prediction. The experiments use the CSE-CIC-IDS2018 dataset. Attack instances are grouped into four categories: BENIGN, DoS, DDoS and BruteForce. To ensure the results are reliable and to avoid information leakage several preprocessing steps are applied. These include chronological data splitting removing features using correlation-based filtering and Recursive Feature Elimination (RFE). The framework also uses train- standardization and applies custom methods to handle class imbalance for each model. As a result the original 78 traffic features are reduced to 30 that are most useful. The ensemble is tested using stratified five-fold out-of-fold (OOF) validation and an unseen chronological test set. This setup allows evaluation of both prediction consistency and how well the model performs on time-based data. The ensemble achieves an OOF accuracy of 95.38% and a Macro-F1 score of 95.22% ± 0.02%. On the chronological test set it reaches 95.23% accuracy and 95.09% Macro-F1. Performance is very strong for BENIGN and DDoS traffic with F1-scores of 99.98% and 99.99% respectively. For BruteForce the recall is 95.70% while DoS has a recall at 87.14%. Most of the mistakes occur between DoS and BruteForce showing the difficulty in distinguishing these two. The small difference between OOF and test set results shows that the framework is stable and can generalize over time. Overall the results show that this approach is both reliable and effective, for network intrusion detection.





