An expanding window temporal validation framework for global AQI and PM2.5 forecasting using machine learning and deep learning models
DOI:
https://doi.org/10.51483/IJAIML.6.8s.2026.1090-1100Keywords:
Air Quality Index, PM2.5, expanding window validation, machine learning, time-series forecasting, global air pollution.Abstract
Accurate forecasting of air quality is a critical public health priority, requiring predictive models that generalize across diverse spatiotemporal conditions. While numerous studies have applied machine learning to this problem, many rely on static train-test splits or single-city datasets, limiting the assessment of model stability and generalizability. This study addresses these limitations by developing and evaluating an expanding window temporal validation framework for multi-step forecasting of Air Quality Index (AQI) and PM2.5 concentrations using a comprehensive global dataset spanning 2020-2025. A leakage-resistant methodology was implemented where models were trained on an expanding historical window (2020-2023) and validated on subsequent out-of-sample years (2024-2025), consistent with recent best practices in temporal validation. We systematically compared five models Linear Regression, Random Forest, XGBoost, Support Vector Regression, and Long Short-Term Memory networks using a consistent feature set of lagged variables, rolling statistics, and seasonal indicators. Results demonstrate that the expanding window framework effectively simulates real-world deployment conditions and prevents over-optimistic performance estimates. Linear Regression consistently achieved the highest R² (0.7217) and lowest error metrics on hold-out test data, outperforming more complex architectures. The LSTM model underperformed, likely due to the structured monthly data and limited training sample size. The study concludes that for continental-scale monthly air quality forecasting, robust, interpretable models like Linear Regression or XGBoost can be more reliable than complex deep learning approaches. These findings provide a validated, reproducible baseline for future research and underscore the importance of selecting model architectures aligned with the data's spatiotemporal characteristics and forecasting task.





