Instructions to use deepsafe/deepsafe-services with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use deepsafe/deepsafe-services with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("deepsafe/deepsafe-services", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
File size: 11,284 Bytes
505afd3 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 | # [ICCV2025] [FakeSTormer] Vulnerability-Aware Spatio-Temporal Learning for Generalizable Deepfake Video Detection

This is an official implementation of FakeSTormer! [[πPaper](https://openaccess.thecvf.com/content/ICCV2025/papers/Nguyen_Vulnerability-Aware_Spatio-Temporal_Learning_for_Generalizable_Deepfake_Video_Detection_ICCV_2025_paper.pdf)]
## Updates
- [x] 26/11/2025:*Official release of code (v1) and pretrained weights π.*
- [x] 08/07/2025: *First version pre-released for this open source code π±.*
- [x] 26/06/2025: *FakeSTormer has been accepted to ICCV2025 π.*
## Abstract
Detecting deepfake videos is highly challenging given the complexity of characterizing spatio-temporal artifacts. Most existing methods rely on binary classifiers trained using real and fake image sequences, therefore hindering their generalization capabilities to unseen generation methods. Moreover, with the constant progress in generative Artificial Intelligence (AI), deepfake artifacts are becoming imperceptible at both the spatial and the temporal levels, making them extremely difficult to capture. To address these issues, we propose a fine-grained deepfake video detection approach called FakeSTormer that enforces the modeling of subtle spatio-temporal inconsistencies while avoiding overfitting. Specifically, we introduce a multi-task learning framework that incorporates two auxiliary branches for explicitly attending artifact-prone spatial and temporal regions. Additionally, we propose a video-level data synthesis strategy that generates pseudo-fake videos with subtle spatio-temporal artifacts, providing high-quality samples and hand-free annotations for our additional branches. Extensive experiments on several challenging benchmarks demonstrate the superiority of our approach compared to recent state-of-the-art methods.
## Main Results
Results on 6 datasets ([CDF2](https://github.com/yuezunli/celeb-deepfakeforensics), [DFW](https://github.com/deepfakeinthewild/deepfake-in-the-wild), [DFD](https://blog.research.google/2019/09/contributing-data-to-deepfake-detection.html), [DFDC, DFDCP](https://ai.meta.com/datasets/dfdc/), and [DiffSwap](https://openaccess.thecvf.com/content/CVPR2023/papers/Zhao_DiffSwap_High-Fidelity_and_Controllable_Face_Swapping_via_3D-Aware_Masked_Diffusion_CVPR_2023_paper.pdf)) under cross-dataset evaluation setting reported by AUC (%) at video-level.
| | CDF2 | DFW | DFD | DFDC | DFDCP | DiffSwap |
|--|--------|------------|------------|------------|---------|-----------|
|<table><thead><tr><th>Compression</th></tr></thead><tbody><tr><td>c23</td></tr></tbody><tbody><tr><td>c0</td></tr></tbody></table>|<table><thead><tr><th>AUC</th></tr></thead><tbody><tr><td>92.4</td></tr></tbody><tbody><tr><td>96.5</td></tr></tbody></table>|<table><thead><tr><th>AUC</th></tr></thead><tbody><tr><td>74.2</td></tr></tbody><tbody><tr><td>76.3</td></tr></tbody></table>|<table><thead><tr><th>AUC</th></tr></thead><tbody><tr><td>98.5</td></tr></tbody><tbody><tr><td>98.9</td></tr></tbody></table>|<table><thead><tr><th>AUC</th></tr></thead><tbody><tr><td>74.6</td></tr></tbody><tbody><tr><td>77.6</td></tr></tbody></table>|<table><thead><tr><th>AUC</th></tr></thead><tbody><tr><td>90.0</td></tr></tbody><tbody><tr><td>94.1</td></tr></tbody></table>|<table><thead><tr><th>AUC</th></tr></thead><tbody><tr><td>96.9</td></tr></tbody><tbody><tr><td>97.7</td></tr></tbody></table>
## Recommended Environment
*For experimental purposes, we encourage the installation of the following libraries. Both Conda or Python virtual env should work.*
* CUDA: 11.4
* [Python](https://www.python.org/): >= 3.8.x
* [PyTorch](https://pytorch.org/get-started/previous-versions/): 1.8.0
* [TensorboardX](https://github.com/lanpa/tensorboardX): 2.5.1
* [ImgAug](https://github.com/aleju/imgaug): 0.4.0
* [Scikit-image](https://scikit-image.org/): 0.17.2
* [Torchvision](https://pytorch.org/vision/stable/index.html): 0.9.0
* [Albumentations](https://albumentations.ai/): 1.1.0
* [mmcv](https://github.com/open-mmlab/mmcv): 1.6.1
* [natsort](https://pypi.org/project/natsort/): 8.4.0
## Pre-trained Models
* π *The pre-trained weights of FakeSTormer can be found [here](https://www.dropbox.com/scl/fo/elk2szqf0du4l6zm5job9/AAdVmNH--6ywHBZGNQJlR5o?rlkey=j8xesf2fu4ahxdw99w5ndrkb2&st=fe6drzpx&dl=0)*
## Docker Build (Optional)
*We further provide an optional Docker file that can be used to build a working env with Docker. More detailed steps can be found [here](dockerfiles/README.md).*
1. Install docker to the system (skip the step if docker has already been installed):
```shell
sudo apt install docker
```
2. To start your docker environment, please go to the folder **dockerfiles**:
```shell
cd dockerfiles
```
3. Create a docker image (you can put any name you want):
```shell
docker build --tag 'fakestormer' .
```
## Quickstart
1. **Preparation**
1. ***Prepare environment***
Installing main packages as the recommended environment. *Note that we recommend building mmcv from source as below.*
> git clone https://github.com/open-mmlab/mmcv.git \
cd mmcv \
git checkout v1.6.1 \
MMCV_WITH_OPS=1 pip install -e .
2. ***Prepare dataset***
1. Downloading [FF++](https://github.com/ondyari/FaceForensics) *Original* dataset for training data preparation. Following the original split convention, it is firstly used to randomly extract frames and facial crops:
```
python package_utils/images_crop.py -d {dataset} \
-c {compression} \
-n {num_frames} \
-t {task}
```
(*This script can also be utilized for cropping faces in other datasets such as [CDF2](https://github.com/yuezunli/celeb-deepfakeforensics), [DFD](https://blog.research.google/2019/09/contributing-data-to-deepfake-detection.html), [DFDCP, DFDC](https://ai.meta.com/datasets/dfdc/) for cross-evaluation test. You do not need to run crop for [DFW](https://github.com/deepfakeinthewild/deepfake-in-the-wild) as the data is already preprocessed*).
| Parameter | Value | Definition |
| --- | --- | --- |
| -d | Subfolder in each dataset. For example: *['Face2Face','Deepfakes','FaceSwap','NeuralTextures', ...]*| You can use one of those datasets.|
| -c | *['raw','c23','c40']*| You can use one of those compressions|
| -n | *256* | Number of frames (*default* 32 for val/test and 256 for train) |
| -t | *['train', 'val', 'test']* | Default train|
These faces cropped are saved for online pseudo-fake generation in the training process, following the data structure below:
```
ROOT = '/data/deepfake_cluster/datasets_df'
βββ Celeb-DFv2
βββ...
βββ FF++
βββ c0
βββ c23
βββ test
βΒ Β βββ videos
βΒ Β βββ Deepfakes
| βββ 000_003
| βββ 044_945
| βββ 138_142
| βββ ...
βΒ Β βββ Face2Face
βΒ Β βββ FaceSwap
βΒ Β βββ NeuralTextures
βΒ Β βββ original
| βββ frames
βββ train
βΒ Β βββ videos
βΒ Β βββ aligned
| βββ 001
| βββ 002
| βββ ...
βΒ Β βββ original
| βββ 001
| βββ 002
| βββ ...
| βββ frames
βββ val
βββ videos
βββ aligned
βββ original
βββ frames
βββ c40
```
2. Downloading **Dlib** [[81]](https://github.com/codeniko/shape_predictor_81_face_landmarks) facial landmarks detector pretrained and place into ```/pretrained/``` for *SBI* synthesis.
3. Landmarks detection. After completing the following script running, a file that stores metadata information of the data is saved at ```processed_data/c23/{SPLIT}_FaceForensics_videos_<n_landmarks>.json```.
```
python package_utils/geo_landmarks_extraction.py \
--config configs/data_preprocessing_c23.yaml \
--extract_landmarks
```
2. **Training script**
We offer a number of config files for different compression levels of training data. For *c23*, opening ```configs/temporal/FakeSTormer_base_c23.yaml```, please make sure you set ```TRAIN: True``` and ```FROM_FILE: True``` and run:
```
.scripts/fakestormer_sbi.sh
```
Otherwise, with *[c0, c40]*, the config file is ```configs/temporal/FakeSTormer_base_[c0, c40].yaml```. You can also find other configs for other network architectures in the ```configs/``` folder.
3. **Testing script**
Opening ```configs/temporal/FakeSTormer_base_c23.yaml```, with ```subtask: eval``` in the *test* section, we support evaluation mode, please turn off ```TRAIN: False``` and ```FROM_FILE: False``` and run:
```
.scripts/test_fakestormer.sh
```
For others (.e.g., data compression levels, network architectures), please change the path of the corresponding config file.
> β οΈ *Please make sure you set the correct path to your downloaded pre-trained weights in the config files.*
> βΉοΈ *Flip test can be used by setting ```flip_test: True```*
> βΉοΈ *The mode for single video inference is also provided, please set ```sub_task: test_vid``` and pass a video path as an argument in test.py*
## Contact
Please contact dat.nguyen@uni.lu. Any questions or discussions are welcomed!
## License
This software is Β© University of Luxembourg and is licensed under the snt academic license. See [LICENSE](LICENSE)
## Acknowledge
We acknowledge the excellent implementation from [OpenMMLab](https://github.com/open-mmlab) ([mmengine](https://github.com/open-mmlab/mmengine), [mmcv](https://github.com/open-mmlab/mmcv)), [SBI](https://github.com/mapooon/SelfBlendedImages), and [LAA-Net](https://github.com/10Ring/LAA-Net).
## Citation
Please kindly consider citing our papers in your publications.
```
@inproceedings{nguyen2025vulnerability,
title={Vulnerability-Aware Spatio-Temporal Learning for Generalizable Deepfake Video Detection},
author={Nguyen, Dat and Astrid, Marcella and Kacem, Anis and Ghorbel, Enjie and Aouada, Djamila},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
pages={10786--10796},
year={2025}
}
```
|