AutoML-Enabled Data Analytics for Scalable and Adaptive Predictive Modeling in Large-Scale Datasets

Authors

  • Rajitha Gentyala

DOI:

https://doi.org/10.51483/IJAIML.6.8s.2026.327-334

Keywords:

automated machine learning; AutoML; scalable predictive modeling; big data analytics; hyperparameter optimization; ensemble learning; adaptive data pipelines.

Abstract

As large, diverse datasets proliferate across domains such as industry, healthcare, finance, and science, the need for scalable, adaptable predictive modeling pipelines is growing. To address this need, Automated Machine Learning (AutoML) has become a practical solution that integrates data preprocessing, feature engineering, model selection, hyperparameter tuning, and ensembling into an automated, unified pipeline. This study presents a framework for large-scale data analytics that leverages AutoML to achieve scalable and adaptive predictive modeling in data-rich contexts. The framework adopts distributed ensembling, model selection with meta-learning, and Bayesian hyperparameter optimization to ensure predictive capacity and computational economy while handling large volumes of data. The framework was tested on synthetic and benchmark datasets ranging from 10K to 50M, using five benchmark learners, namely Random Forest, XGBoost, LightGBM, a feedforward neural network, and a weighted ensemble meta-learner. These results show that the AutoML-enabled pipeline consistently achieved higher predictive accuracy than any single base learner, with a mean (across validation folds) cross-validation accuracy of 89.4% for the AutoML search space; meanwhile, the weighted ensemble configuration had an order of magnitude lower development time for large data volumes than non-AutoML, manually tuned pipelines. The framework worked well in large-scale deployments, as scalability tests were performed using 64 compute nodes, showing near-linear performance increases across the nodes. These conclusions are placed in the context of previous work on the automation of machine learning and hyperparameter optimization, and big data analytics, while noting constraints on generalizability to other domains, interpretability, and computational requirements. Overall, the study shows that the use of AutoML-based analytics offers a plausible route toward democratizing predictive modeling for large-scale data in the future, provided that adaptive resource management and explanation mechanisms are included in future implementations.

Downloads

Published

2026-08-01

How to Cite

Gentyala, R. (2026). AutoML-Enabled Data Analytics for Scalable and Adaptive Predictive Modeling in Large-Scale Datasets. International Journal of Artificial Intelligence and Machine Learning, 6(8s), 327–334. https://doi.org/10.51483/IJAIML.6.8s.2026.327-334