Quality-Aware Spatial–Frequency–Temporal Fusion for Robust Deepfake Detection
DOI:
https://doi.org/10.51483/IJAIML.6.8s.2026.926-937Keywords:
Deepfake Detection, Media Forensics, Multimedia Security, Synthetic Media, Attention Mechanisms.Abstract
The growth in availability and sophistication of generative AI has accelerated the creation of deepfakes, which are synthetic and modified visual media that present challenges to the integrity of information, security, and privacy. This paper provides a measured and thorough solution to the detection and analysis of deep fakes in multimedia content using advanced Deep Learning frameworks. The proposed system is holistic, analyzing spatio-temporal and semantic consistency, and includes convolutional neural networks for analysis in the spatial domain, feature extraction in the frequency domain, and verification of temporal consistency. A new hybrid architecture is being proposed, whereby XceptionNet is being combined with frequency-aware components and attention-based mechanisms for deepfake detection. The proposed system is thoroughly evaluated using the FaceForensics++, Celeb-DF, DFDC benchmark datasets and demonstrates cross-dataset generalization. Further, the proposed framework is robust to manipulation, and preserves interpretability through attention visualization and a deepfake detection architecture reaching an accuracy of 94.7%. A comprehensive evaluation of the framework on different datasets demonstrating a 12% increase in generalization showed the supremacy of model.





