DeepSafe Ensemble Artifacts
The meta-learners that turn 19 individual detector scores into one calibrated verdict, for DeepSafe.
Unlike the model code and weights in the other DeepSafe repositories, these are first-party: trained by us, licensed PolyForm Noncommercial 1.0.0, same as the project.
Contents
| Modality | Meta-learner | Held-out AUC |
|---|---|---|
| Image | LightGBM over 7 models | 0.9466 |
| Audio | Random Forest over 3 models | 0.8290 |
| Video | XGBoost over 9 models | 0.6694 |
Each modality ships four files:
<modality>_meta_learner.pklโ the trained model<modality>_scaler.pklโ feature scaling<modality>_calibrator.pklโ Platt calibration, so scores read as probabilities<modality>_config.jsonโ feature order and fallback weights
Trained on the 15,499-sample medium evaluation tier. Without these, the inference server produces per-model scores but no ensemble verdict.
Usage
setup.sh fetches these automatically. Manually:
from huggingface_hub import snapshot_download
snapshot_download("deepsafe/ensemble", local_dir="models/ensemble/artifacts")
The numbers are the point
Held-out video AUC is 0.6694. That is barely above chance on generators the models were not trained for, and it is lower than the cross-validated figure produced during training. We publish the held-out number because the gap between the two is the finding. See BENCHMARK.md.
Security note
These are Python pickles, which execute code on load. Only load them from a
source you trust. DeepSafe's loader refuses any pickle outside its configured
artifacts directory, but that is a guardrail, not a guarantee. If you are
security-sensitive, retrain your own with deepsafe fit --tier 1.