AV-DeepFake-Net: Attention-Guided and Uncertainty-Aware Network for Audiovisual DeepFake Detection

dc.contributor.affiliationHainan University
dc.contributor.affiliationHainan University
dc.contributor.affiliationHainan University
dc.contributor.affiliationEdinburgh Napier University
dc.contributor.affiliationTaif University; Taif University
dc.contributor.affiliationKing Khalid University
dc.contributor.affiliationPrincess Nourah bint Abdulrahman University
dc.contributor.authorNasir Saleem; Hainan University
dc.contributor.authorAhmad Ali; Hainan University
dc.contributor.authorZhuoqi Zeng; Hainan University
dc.contributor.authorAdeel Hussain; Edinburgh Napier University
dc.contributor.authorSami Bourouis; Taif University; Taif University
dc.contributor.authorSami Dhahbi; King Khalid University
dc.contributor.authorAfef Dhahbi; Princess Nourah bint Abdulrahman University
dc.contributor.orcid
dc.contributor.orcid
dc.contributor.orcid
dc.contributor.orcid
dc.contributor.orcid
dc.contributor.orcid
dc.contributor.orcid
dc.contributor.rorhttps://ror.org/03q648j11
dc.contributor.rorhttps://ror.org/03q648j11
dc.contributor.rorhttps://ror.org/03q648j11
dc.contributor.rorhttps://ror.org/03zjvnn91
dc.contributor.rorhttps://ror.org/014g1a453
dc.contributor.rorhttps://ror.org/052kwzs30
dc.contributor.rorhttps://ror.org/05b0cyh02
dc.date.accessioned2026-09-07T14:07:08Z
dc.date.issued2026-08-28
dc.date.updated2026-09-07T14:07:08Z
dc.description.abstractDetecting audio-visual DeepFake (AVDeepFake) is becoming increasingly important as synthetic media tools become widely accessible and spread across consumer devices. In this study, we present an Detecting audio-visual DeepFakes (AV-DeepFakes) has become increasingly critical with the rapid proliferation of accessible synthetic media generation tools across consumer platforms. In this work, we propose a highperformance, deployment-efficient AV-DeepFake detection framework tailored for real-world consumer devices. The proposed model integrates a 3D convolutional visual encoder with a 2D convolutional audio encoder to learn synchronized multimodal representations, effectively capturing spatial, spectral, and prosodic inconsistencies inherent in manipulated content. To detect temporal forgeries, we introduce a bidirectional complementary boundary module that precisely localizes manipulation onsets and offsets. A cross-modal attention fusion mechanism aggregates modality-specific cues, while an uncertainty-aware gating strategy suppresses unreliable signals to improve robustness. Furthermore, a cross-modal discrepancy minimization loss encourages alignment for genuine samples while maximizing divergence for forged content, strengthening multimodal consistency learning. Extensive evaluations on FaceForensics++ and LAV-DF demonstrate the effectiveness of the proposed approach, achieving 97.1% AUC for clip-level detection and an 81.2% F1-score for temporal boundary localization, while reducing inference time by 5× compared with transformer-based methods.
dc.description.endingpagee6431
dc.description.startingpagee6431
dc.identifier.urihttps://doi.org/10.9781/ijimai.2026.6431
dc.identifier.urihttps://reunir.unir.net/handle/123456789/20561
dc.publisherUniversidad Internacional de La Rioja
dc.relation.ispartof10
dc.relation.ispartofvolume1
dc.rightsopenAccess
dc.rights.uriopenAccess
dc.subjectAV DeepFake Detection
dc.subjectCross-Modal Fusion
dc.subjectTemporal Boundary Localization
dc.titleAV-DeepFake-Net: Attention-Guided and Uncertainty-Aware Network for Audiovisual DeepFake Detection

Archivos

Bloque original

Mostrando 1 - 1 de 1
Cargando...
Nombre:
ijimai10_1_2026_6431.pdf
Tamaño:
1.05 MB
Formato:
Adobe Portable Document Format

Bloque de licencias

Mostrando 1 - 1 de 1
Cargando...
Nombre:
license.txt
Tamaño:
1.65 KB
Formato:
Item-specific license agreed upon to submission
Descripción: