Video Classification
Transformers
Safetensors
vjepa21
feature-extraction
video
vjepa
vjepa2
v-jepa-2.1
self-supervised
world-model
custom_code
Instructions to use apiantonio/vjepa2.1-vit-gigantic-384 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use apiantonio/vjepa2.1-vit-gigantic-384 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("video-classification", model="apiantonio/vjepa2.1-vit-gigantic-384", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("apiantonio/vjepa2.1-vit-gigantic-384", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Fix transformers 4.x/5.x compat, implement output_hidden_states/attentions and out_layers, fix hierarchical predictor input, add video processor
4bec704 verified | """V-JEPA 2.1 — HuggingFace port.""" | |
| from .configuration_vjepa21 import VJEPA21Config | |
| from .modeling_vjepa21 import ( | |
| VJEPA21ForVideoClassification, | |
| VJEPA21Model, | |
| VJEPA21PreTrainedModel, | |
| ) | |
| __all__ = [ | |
| "VJEPA21Config", | |
| "VJEPA21Model", | |
| "VJEPA21PreTrainedModel", | |
| "VJEPA21ForVideoClassification", | |
| ] | |
| # `BaseVideoProcessor` needs torchvision. Import it lazily so that a runtime | |
| # without torchvision can still load the model; `AutoVideoProcessor` resolves the | |
| # class through `auto_map` and does not go through this file. | |
| try: # pragma: no cover - depends on the environment | |
| from .video_processing_vjepa21 import VJEPA21VideoProcessor | |
| __all__.append("VJEPA21VideoProcessor") | |
| except ImportError: # pragma: no cover | |
| pass | |