AV-DeepFake-Net: Attention-Guided and Uncertainty-Aware Network for Audiovisual DeepFake Detection
| dc.contributor.affiliation | Hainan University | |
| dc.contributor.affiliation | Hainan University | |
| dc.contributor.affiliation | Hainan University | |
| dc.contributor.affiliation | Edinburgh Napier University | |
| dc.contributor.affiliation | Taif University; Taif University | |
| dc.contributor.affiliation | King Khalid University | |
| dc.contributor.affiliation | Princess Nourah bint Abdulrahman University | |
| dc.contributor.author | Nasir Saleem; Hainan University | |
| dc.contributor.author | Ahmad Ali; Hainan University | |
| dc.contributor.author | Zhuoqi Zeng; Hainan University | |
| dc.contributor.author | Adeel Hussain; Edinburgh Napier University | |
| dc.contributor.author | Sami Bourouis; Taif University; Taif University | |
| dc.contributor.author | Sami Dhahbi; King Khalid University | |
| dc.contributor.author | Afef Dhahbi; Princess Nourah bint Abdulrahman University | |
| dc.contributor.orcid | ||
| dc.contributor.orcid | ||
| dc.contributor.orcid | ||
| dc.contributor.orcid | ||
| dc.contributor.orcid | ||
| dc.contributor.orcid | ||
| dc.contributor.orcid | ||
| dc.contributor.ror | https://ror.org/03q648j11 | |
| dc.contributor.ror | https://ror.org/03q648j11 | |
| dc.contributor.ror | https://ror.org/03q648j11 | |
| dc.contributor.ror | https://ror.org/03zjvnn91 | |
| dc.contributor.ror | https://ror.org/014g1a453 | |
| dc.contributor.ror | https://ror.org/052kwzs30 | |
| dc.contributor.ror | https://ror.org/05b0cyh02 | |
| dc.date.accessioned | 2026-09-07T14:07:08Z | |
| dc.date.issued | 2026-08-28 | |
| dc.date.updated | 2026-09-07T14:07:08Z | |
| dc.description.abstract | Detecting audio-visual DeepFake (AVDeepFake) is becoming increasingly important as synthetic media tools become widely accessible and spread across consumer devices. In this study, we present an Detecting audio-visual DeepFakes (AV-DeepFakes) has become increasingly critical with the rapid proliferation of accessible synthetic media generation tools across consumer platforms. In this work, we propose a highperformance, deployment-efficient AV-DeepFake detection framework tailored for real-world consumer devices. The proposed model integrates a 3D convolutional visual encoder with a 2D convolutional audio encoder to learn synchronized multimodal representations, effectively capturing spatial, spectral, and prosodic inconsistencies inherent in manipulated content. To detect temporal forgeries, we introduce a bidirectional complementary boundary module that precisely localizes manipulation onsets and offsets. A cross-modal attention fusion mechanism aggregates modality-specific cues, while an uncertainty-aware gating strategy suppresses unreliable signals to improve robustness. Furthermore, a cross-modal discrepancy minimization loss encourages alignment for genuine samples while maximizing divergence for forged content, strengthening multimodal consistency learning. Extensive evaluations on FaceForensics++ and LAV-DF demonstrate the effectiveness of the proposed approach, achieving 97.1% AUC for clip-level detection and an 81.2% F1-score for temporal boundary localization, while reducing inference time by 5× compared with transformer-based methods. | |
| dc.description.endingpage | e6431 | |
| dc.description.startingpage | e6431 | |
| dc.identifier.uri | https://doi.org/10.9781/ijimai.2026.6431 | |
| dc.identifier.uri | https://reunir.unir.net/handle/123456789/20561 | |
| dc.publisher | Universidad Internacional de La Rioja | |
| dc.relation.ispartof | 10 | |
| dc.relation.ispartofvolume | 1 | |
| dc.rights | openAccess | |
| dc.rights.uri | openAccess | |
| dc.subject | AV DeepFake Detection | |
| dc.subject | Cross-Modal Fusion | |
| dc.subject | Temporal Boundary Localization | |
| dc.title | AV-DeepFake-Net: Attention-Guided and Uncertainty-Aware Network for Audiovisual DeepFake Detection |


