Distributed Content Classification Architecture: Privacy-Preserving Federated Learning For Heterogeneous Edge Environments
Keywords:
distributed content classification; federated learning; edge computing; privacy-preserving machine learning; heterogeneous environments.Abstract
One of the main issues in categorizing digital content is that not only data is scattered across various devices, the systems are extremely diverse, and at the same time, user privacy should be respected.
To address these issues, DCCA (Distributed Content Classification with Aggregation) is introduced as a federated learning framework based on privacy preservation, which enables different devices to collaboratively train a model without sharing raw data with each other.
The suggested system maintains the input of two feature extraction engines combining TF, IDF and Word2Vec representations and a weighted federated aggregation strategy that can manage non, IID data distributions.Model training takes place entirely on local devices, and only encrypted model parameters are shared with a central coordinator for secure aggregation.
Extensive experiments on 18, 500 annotated content samples spread over heterogeneous edge nodes have demonstrated that DCCA can attain classification accuracy and F1, score of 90.5% and 91.6%, respectively. Moreover, the system reduces communication overhead by 25.6% and achieves convergence after 19 federated rounds. Under realistic non, IID scenarios, the model accurately classifies contents into three categories Instructional (38.9%), Recreational (48.1%), and Utility (13.0%).
Additionally, performance analysis highlights stable convergence and dependable functioning even on diverse devices. Overall, the results suggest that federated learning can drastically lessen the gap in performance between centralized models (with 94.1% accuracy) and decentralized privacy, conscious systems.DCCA offers a practical and scalable way to privacy, aware content classification at present, day edge computing environments.





