Feature Extraction
Transformers
Safetensors
Algerian Arabic
dzair
algerian-darija
arabizi
arabic
encoder
replaced-token-detection
low-resource
custom_code
Eval Results (legacy)
Instructions to use algerian-nlp/DZAIR with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use algerian-nlp/DZAIR with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="algerian-nlp/DZAIR", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("algerian-nlp/DZAIR", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download tokenizer_config.json from algerian-nlp/DZAIR: direct link, hf CLI and curl.
- Browser
- Download file 563 Bytes
-
https://huggingface.co/algerian-nlp/DZAIR/resolve/main/tokenizer_config.json
- Command line
-
hf download hf://algerian-nlp/DZAIR/tokenizer_config.json
-
curl -L -o tokenizer_config.json https://huggingface.co/algerian-nlp/DZAIR/resolve/main/tokenizer_config.json
563 Bytes
| { | |
| "tokenizer_class": "DebertaV2Tokenizer", | |
| "model_input_names": [ | |
| "input_ids", | |
| "attention_mask" | |
| ], | |
| "unk_token": "[UNK]", | |
| "pad_token": "[PAD]", | |
| "cls_token": "[CLS]", | |
| "sep_token": "[SEP]", | |
| "mask_token": "[MASK]", | |
| "do_lower_case": false, | |
| "keep_accents": true, | |
| "split_by_punct": false, | |
| "padding_side": "right", | |
| "truncation_side": "right", | |
| "model_max_length": 512, | |
| "latin_lowercase_before_encode": true, | |
| "note": "Lowercase Latin spans before encoding; the tokenizer then wraps input as [CLS] chunk [SEP]. See the model card." | |
| } | |