International Journal For Multidisciplinary Research

E-ISSN: 2582-2160     Impact Factor: 9.24

A Widely Indexed Open Access Peer Reviewed Multidisciplinary Bi-monthly Scholarly International Journal

Call for Paper Volume 8, Issue 4 (July-August 2026) Submit your research before last 3 days of August to publish your research paper in the issue of July-August.

A Deep Learning Framework for Real-Time Detection of Deepfake Images, Audio and Videos

Author(s) Mr. Srijib Samanta
Country India
Abstract ABSTRACT:
Deepfakes have progressed from visibly defective face swaps to high-quality synthetic images, cloned speech and audio–visual manipulations that can survive social-media compression. A detector intended for real-time use must therefore do more than classify isolated frames: it must capture spatial traces, temporal inconsistency, acoustic artifacts and cross-modal synchronisation while operating within strict latency and memory limits. This paper develops a unified deep learning framework for detecting deepfake images, audio and videos. The study follows a structured research review and design-science methodology; it proposes an architecture and evaluation protocol rather than claiming fabricated experimental results. The framework contains a lightweight visual branch, a self-supervised acoustic branch, a temporal video branch and an audio–visual consistency branch. Reliability-aware gated fusion combines available modalities without forcing a missing or corrupted stream to contribute equally. A streaming inference layer uses face tracking, voice-activity detection, adaptive frame sampling, sliding windows, calibration and early exit to support real-time operation. The proposed assessment emphasises cross-dataset and cross-generator generalisation, area under the receiver operating characteristic curve, equal error rate, calibration, attack-wise recall, latency, throughput and memory. FaceForensics++, Celeb-DF v2, DFDC, ASVspoof 2021 and FakeAVCeleb are identified as complementary data sources, with identity-disjoint splitting and compression-aware augmentation required to prevent leakage. The analysis argues that multimodal fusion improves evidential coverage but does not automatically guarantee robustness: dataset bias, unseen synthesis methods, adversarial post-processing and demographic imbalance remain central threats. The paper concludes that operationally credible deepfake detection requires open-set evaluation, uncertainty-aware decisions, interpretable evidence and secure human review rather than a single opaque probability score.
Keywords Keywords: deepfake detection; multimodal learning; audio forensics; video forensics; real-time inference; cross-modal fusion; deep learning; digital authenticity.
Field Engineering
Published In Volume 8, Issue 4, July-August 2026
Published On 2026-08-25

Share this