Projects per year
Abstract
The exponential evolution of generative AI has intensified the need for deepfake detection methods with strong cross-dataset generalization. Existing approaches often attempt to improve robustness by scaling network parameters, however, such strategies frequently result in unstable outcomes and excessive computational cost. To address these limitations, we propose FlashF5-MaViT, a lightweight 4.21M parameters deepfake detector that integrates multi-frequency analysis with spatial modeling.
FlashF5-MaViT unifies three complementary components: (i) a Quad-Frequency Pathway incorporating the Fourier, Cosine, Azimuthally averaged, and Cepstrum representations, globally modeled via a lightweight Mamba module, (ii) a Wavelet Scattering Transform branch with lightweight Central Difference Convolution for extracting local texture invariants, (iii) a FastViT hybrid spatial backbone for capturing visual level representations. A lightweight attention mechanism fuses frequency derived features with spatial embeddings, producing a balanced representation of global semantics and local spectral artifacts. Trained on FF++ (HQ) and evaluated across cross datasets: Celeb-DF, WildDeepfake, and DeepFake Detection, FlashF5-MaViT consistently outperforms 13 of 15 higher parameter baselines, including CNN and Vision Transformer models up to 86M parameters, while maintaining uniform AUC results across datasets. Remarkably, it achieves state-of-the-art performance on the challenging WildDeepfake dataset, where larger models fail to generalize effectively with an AUC of 81.09% leading to an average improvement of 1% over baseline performance. These results show that combining multi-frequency modeling with FastViT creates a concise but strong framework for deepfake detection, avoiding the need for large models and moving closer to practical and reliable use. FlashF5-MaViT source code and is publicly accessible at https://github.com/noureldinalaa/FlashF5-MaViT
FlashF5-MaViT unifies three complementary components: (i) a Quad-Frequency Pathway incorporating the Fourier, Cosine, Azimuthally averaged, and Cepstrum representations, globally modeled via a lightweight Mamba module, (ii) a Wavelet Scattering Transform branch with lightweight Central Difference Convolution for extracting local texture invariants, (iii) a FastViT hybrid spatial backbone for capturing visual level representations. A lightweight attention mechanism fuses frequency derived features with spatial embeddings, producing a balanced representation of global semantics and local spectral artifacts. Trained on FF++ (HQ) and evaluated across cross datasets: Celeb-DF, WildDeepfake, and DeepFake Detection, FlashF5-MaViT consistently outperforms 13 of 15 higher parameter baselines, including CNN and Vision Transformer models up to 86M parameters, while maintaining uniform AUC results across datasets. Remarkably, it achieves state-of-the-art performance on the challenging WildDeepfake dataset, where larger models fail to generalize effectively with an AUC of 81.09% leading to an average improvement of 1% over baseline performance. These results show that combining multi-frequency modeling with FastViT creates a concise but strong framework for deepfake detection, avoiding the need for large models and moving closer to practical and reliable use. FlashF5-MaViT source code and is publicly accessible at https://github.com/noureldinalaa/FlashF5-MaViT
| Original language | English |
|---|---|
| Title of host publication | 2025 9th International Conference on Vision, Image and Signal Processing (ICVISP) |
| Place of Publication | Piscataway, U.S. |
| Publisher | Institute of Electrical and Electronics Engineers |
| Number of pages | 8 |
| ISBN (Electronic) | 9798331556822 |
| ISBN (Print) | 9798331556839 |
| DOIs | |
| Publication status | Published - 31 Mar 2026 |
| Event | 2025 9th International Conference on Vision, Image and Signal Processing (ICVISP) - Xi'an, China Duration: 28 Nov 2025 → 30 Nov 2025 |
Publication series
| Name | International Conference on Vision, Image and Signal Processing (ICVISP) |
|---|---|
| Publisher | Institute of Electrical and Electronics Engineers, Inc. |
Conference
| Conference | 2025 9th International Conference on Vision, Image and Signal Processing (ICVISP) |
|---|---|
| Period | 28/11/25 → 30/11/25 |
Keywords
- Central Difference Convolution (CDC)
- Cross-dataset generalization
- Deepfake detection
- FastViT
- Lightweight models
- Mamba architecture
- Multi-frequency modeling
Fingerprint
Dive into the research topics of 'FlashF5-MaViT: a fast five frequency Mamba with CDC–FastViT architecture for deepfake detection'. Together they form a unique fingerprint.Projects
- 1 Active
-
Deepfake Detection
Liang, X. (PI), Nebel, J.-C. (CoI), Greenhill, D. (CoI), Edwards, P. (Researcher) & Badr, N. (Researcher)
6/03/23 → …
Project: Research
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver