An Imbalance-Sensitive Multi-Sampling Ensemble Framework with Validation-Driven Threshold Optimization for Banking Transaction Fraud Detection
Keywords:
Banking transaction fraud detection, class imbalance, data leakage, decision-threshold optimization, ensemble learning, PR-AUC, resampling, SMOTE.Abstract
Banking transaction fraud detection remains difficult because fraudulent events are rare, behaviourally heterogeneous and continuously evolving. Two methodological weaknesses recur in the literature: preprocessing or resampling statistics estimated on the full dataset, which leaks information into the evaluation partitions, and a fixed decision threshold of 0.50, which is rarely optimal under extreme skew. This paper presents an imbalance-sensitive multi-sampling ensemble framework with validation-driven threshold optimization that addresses both. A corpus of 240,000 banking transactions containing 4,782 fraudulent records (1.99%) is partitioned by stratified sampling into 70% training, 15% validation and 15% test subsets. Every imputation, encoding and scaling parameter is estimated on the training subset alone and applied unchanged to the remaining partitions, and all resampling is confined to the training subset. Seven resampling operators are crossed with five ensemble classifiers to yield 35 controlled configurations. For each configuration a decision threshold is selected by maximizing the validation F1-score over τ ∈ {0.01, 0.02, …, 0.99} and is then applied unchanged to the held-out test subset. Performance is reported through precision, recall, F1-score, ROC-AUC, PR-AUC and confusion-matrix analysis. Random Forest and Extra Trees are the most robust learners across all seven operators; the ADASYN–Extra Trees configuration recovers all 717 fraudulent test transactions with no false positives at a selected threshold of 0.6983. The framework offers a reproducible and auditable protocol for highly imbalanced financial fraud screening. The implementation is available at https://github.com/neha27upadhyay-cmyk/Fraud-Detection-using-ML





