A Hybrid Nlp And Big Data Analytics Framework For Privacy-Aware Information Extraction In Intelligent Transportation Systems
DOI:
https://doi.org/10.51483/IJAIML.6.2.2026.307-319Keywords:
Intelligent Transportation Systems; natural language processing; big data analytics; privacy preservation; information extraction; named entity recognition; traffic incident detection.Abstract
Intelligent Transportation Systems (ITS) now can produce a continuous mixture of structured, semi-structured, and unstructured data from vehicles, road-side units, GPS devices, smart cards, traffic controllers, ticketing platforms, passenger feedback portals, and control-room messages. The analytics of these available data is clear, but the direct storage of raw transport records has two practical problems that are underestimated. It increases the processing cost, and it keeps personally identifying details inside the analytical pipeline longer than necessary. This paper proposes a Hybrid Natural Language Processing and Big Data Analytics Framework for Privacy-Aware Information Extraction in ITS. The framework first converts the heterogeneous transport text into normalized tokens, named entities, and event records using rule assisted NLP, gazetteer matching, domain patterns, and event clustering. It then masks or generalizes the sensitive identifiers before data is retained for downstream analytics. Instead of treating raw messages as the main analytical object, the framework stores compact event level records containing incident type, route, time, generalized location, and anonymized operational identifiers. A transparent simulation using a synthetic ITS message corpus was conducted to examine the extraction accuracy, storage reduction, processing time, and privacy masking coverage. Results show that the proposed pipeline improves macro F1-score over a simple keyword baseline, it reduces retained storage in the synthetic setting, and it masks all explicitly generated passenger and vehicle identifiers. The obtained results should not be read as field deployment evidence; rather, they indicate that privacy-aware extraction before long-term storage is a plausible design for the direction of smart transportation analytics. This paper also discusses its limitations relating to city specific language, multilingual data, model drift, and the privacy-utility trade-off that appears whenever exact mobility traces are generalized.





