Mitigating Identity-Leaking Features in SDN Traffic Classification through Behavioral Learning
Keywords:
Network Traffic Classification; Feature Selection; Behavioral Robustness; Machine Learning Integrity; Encrypted Traffic Identi-fication.Abstract
Recent growth in network traffic classification has been associated with near-perfect accuracy, with many studies reporting results above 99%. However, such results may create a false impression of model robustness. These high accuracy levels are often achieved through implicit dependence on identity-leaking attributes, such as port numbers and Layer-7 protocol labels. Such features introduce serious methodological vulnerabilities and make models fragile in modern encrypted, tunneled, and dynamically routed network environments. This paper critiques the existing accuracy-centered paradigm by proposing a domain-oriented behavioral machine learning framework that explicitly removes identity-based shortcuts and prioritizes methodological integrity over nominal performance. A highly restrictive filtering pipeline is presented to isolate protocol-neutral behavioral signatures that describe the inherent spatio-temporal characteristics of application traffic. A compact hybrid feature set of 30 behavioral attributes, derived through One-vs-Rest binary importance analysis, combines domain-protected flow measurements with application-specific local signatures. A large-scale experimental assessment using a comprehensive SDN dataset shows that while eliminating identity-leaking features recalibrates the inflated nominal accuracy of 99% to a robust baseline of 80%, it significantly improves generalization stability and cross-session reliability. To confirm that the observed performance results are due to feature quality rather than algorithmic bias, the proposed feature set is evaluated across several machine learning architectures, including Random Forest, XGBoost, Decision Tree, and LightGBM. These models show a consistent convergence in performance. The findings demonstrate that behavioral robustness, rather than inflated accuracy, is the key requirement for deployable network traffic classification systems. This work establishes a practical behavioral standard for encrypted traffic identification and offers a conceptual framework for building reflective, identity-free machine learning models in modern network infrastructures.





