Skip to main navigation Skip to search Skip to main content

FlashF5-MaViT: a fast five frequency Mamba with CDC–FastViT architecture for deepfake detection

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

The exponential evolution of generative AI has intensified the need for deepfake detection methods with strong cross-dataset generalization. Existing approaches often attempt to improve robustness by scaling network parameters, however, such strategies frequently result in unstable outcomes and excessive computational cost. To address these limitations, we propose FlashF5-MaViT, a lightweight 4.21M parameters deepfake detector that integrates multi-frequency analysis with spatial modeling.

FlashF5-MaViT unifies three complementary components: (i) a Quad-Frequency Pathway incorporating the Fourier, Cosine, Azimuthally averaged, and Cepstrum representations, globally modeled via a lightweight Mamba module, (ii) a Wavelet Scattering Transform branch with lightweight Central Difference Convolution for extracting local texture invariants, (iii) a FastViT hybrid spatial backbone for capturing visual level representations. A lightweight attention mechanism fuses frequency derived features with spatial embeddings, producing a balanced representation of global semantics and local spectral artifacts. Trained on FF++ (HQ) and evaluated across cross datasets: Celeb-DF, WildDeepfake, and DeepFake Detection, FlashF5-MaViT consistently outperforms 13 of 15 higher parameter baselines, including CNN and Vision Transformer models up to 86M parameters, while maintaining uniform AUC results across datasets. Remarkably, it achieves state-of-the-art performance on the challenging WildDeepfake dataset, where larger models fail to generalize effectively with an AUC of 81.09% leading to an average improvement of 1% over baseline performance. These results show that combining multi-frequency modeling with FastViT creates a concise but strong framework for deepfake detection, avoiding the need for large models and moving closer to practical and reliable use. FlashF5-MaViT source code and is publicly accessible at https://github.com/noureldinalaa/FlashF5-MaViT
Original languageEnglish
Title of host publication2025 9th International Conference on Vision, Image and Signal Processing (ICVISP)
Place of PublicationPiscataway, U.S.
PublisherInstitute of Electrical and Electronics Engineers
Number of pages8
ISBN (Electronic)9798331556822
ISBN (Print)9798331556839
DOIs
Publication statusPublished - 31 Mar 2026
Event2025 9th International Conference on Vision, Image and Signal Processing (ICVISP) - Xi'an, China
Duration: 28 Nov 202530 Nov 2025

Publication series

NameInternational Conference on Vision, Image and Signal Processing (ICVISP)
PublisherInstitute of Electrical and Electronics Engineers, Inc.

Conference

Conference2025 9th International Conference on Vision, Image and Signal Processing (ICVISP)
Period28/11/2530/11/25

Keywords

  • Central Difference Convolution (CDC)
  • Cross-dataset generalization
  • Deepfake detection
  • FastViT
  • Lightweight models
  • Mamba architecture
  • Multi-frequency modeling

Fingerprint

Dive into the research topics of 'FlashF5-MaViT: a fast five frequency Mamba with CDC–FastViT architecture for deepfake detection'. Together they form a unique fingerprint.

Cite this