Text Generation
MLX
Safetensors
English
llama
optiq
quantized
mixed-precision
4bit
8bit
apple-silicon
edit-prediction
next-edit-suggestion
code
autocomplete
fim
4-bit precision
Instructions to use bouroo/zeta-2.1-OptiQ-4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use bouroo/zeta-2.1-OptiQ-4 with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("bouroo/zeta-2.1-OptiQ-4") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use bouroo/zeta-2.1-OptiQ-4 with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "bouroo/zeta-2.1-OptiQ-4" --prompt "Once upon a time"
- Atomic Chat
Upload generation_config.json with huggingface_hub
Browse files- generation_config.json +10 -0
generation_config.json
ADDED
|
@@ -0,0 +1,10 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"bos_token_id": 0,
|
| 3 |
+
"eos_token_id": 2,
|
| 4 |
+
"pad_token_id": 1,
|
| 5 |
+
"do_sample": true,
|
| 6 |
+
"temperature": 0.2,
|
| 7 |
+
"top_p": 0.95,
|
| 8 |
+
"max_new_tokens": 128,
|
| 9 |
+
"repetition_penalty": 1.0
|
| 10 |
+
}
|