1 School of Computer Science and Engineering, Anhui University of Science and Technology, Huainan, China.
2 School of artificial intelligence, Anhui University of Science and Technology, Huainan, China.
International Journal of Science and Research Archive, 2026, 19(03), 540-563
Article DOI: 10.30574/ijsra.2026.19.3.1312
Received on 03 May 2026; revised on 10 June 2026; accepted on 13 June 2026
The advent of deepfakes, driven by developments in Generative Adversarial Networks (GANs), represents a fundamental threat to the validity of digital content. Although early detection entailed mostly reacting to visual anomalies, the increasing sophistication of deepfakes now demands approaches that react to both visuals and sound. This survey provides a comprehensive assessment of the current state of multimodal deepfake detection, with particular emphasis on GAN-based generation and detection approaches. We categorize existing approaches into three general classes: early fusion, late fusion, and hybrid fusion models, and contrast their performance on widely used benchmarking datasets. We also explore the use of cutting-edge architectures such as Transformers and diffusion models to improve detection accuracy and resilience. The survey also introduces key challenges such as generalization challenges across tasks, class imbalance, adversarial attacks, and the gap between the audio and vision streams. Finally, we introduce some possible directions of future work, including designing zero-shot detection systems, leveraging explainable AI techniques, and striving for real-time detection on edge devices. The objective of this study is to provide insights to allow researchers and practitioners to develop more effective, dynamic, and explainable multimodal detection systems to combat the constantly evolving threat that deepfakes represent.
Deepfake Detection; Generative Adversarial Networks; Multimodal Learning; Audio-Visual Fusion; GANs; Forgery Detection; Synthetic Media; Audio-Visual Inconsistency; Adversarial AI; Media Forensics
Preview Article PDF
Maryam Tariq, Fazal Shah, Sahar Ali and Abdulaziz Hani Aldali. Beyond the Face: Multimodal Deepfake Detection Using GANs and Audio-Visual Cues. International Journal of Science and Research Archive, 2026, 19(03), 540-563. Article DOI: https://doi.org/10.30574/ijsra.2026.19.3.1312.






