WaveCut commited on
Commit
96a21c6
·
verified ·
1 Parent(s): f0bcd17

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +48 -21
README.md CHANGED
@@ -1,39 +1,66 @@
1
  ---
 
 
 
2
  license: apache-2.0
3
- base_model: WaveCut/Qwythos-9B-v2-Heretic
 
 
 
4
  tags:
5
- - heretic
6
- - uncensored
7
- - abliteration
8
- - mlx
 
 
9
  pipeline_tag: text-generation
10
- library_name: mlx
11
  ---
12
 
13
- # WaveCut/Qwythos-9B-v2-Heretic-MLX-8bit
 
 
 
 
14
 
15
- This model [WaveCut/Qwythos-9B-v2-Heretic-MLX-8bit](https://huggingface.co/WaveCut/Qwythos-9B-v2-Heretic-MLX-8bit) was
16
- converted to MLX format from [WaveCut/Qwythos-9B-v2-Heretic](https://huggingface.co/WaveCut/Qwythos-9B-v2-Heretic)
17
- using mlx-lm version **0.31.3**.
 
 
18
 
19
- ## Use with mlx
 
 
 
 
 
20
 
21
  ```bash
22
- pip install mlx-lm
 
 
 
23
  ```
24
 
 
 
25
  ```python
26
  from mlx_lm import load, generate
27
-
28
  model, tokenizer = load("WaveCut/Qwythos-9B-v2-Heretic-MLX-8bit")
 
 
 
 
 
 
 
 
29
 
30
- prompt = "hello"
31
 
32
- if tokenizer.chat_template is not None:
33
- messages = [{"role": "user", "content": prompt}]
34
- prompt = tokenizer.apply_chat_template(
35
- messages, add_generation_prompt=True, return_dict=False,
36
- )
37
 
38
- response = generate(model, tokenizer, prompt=prompt, verbose=True)
39
- ```
 
 
1
  ---
2
+ language:
3
+ - en
4
+ - ru
5
  license: apache-2.0
6
+ library_name: mlx
7
+ base_model:
8
+ - WaveCut/Qwythos-9B-v2-Heretic
9
+ base_model_relation: quantized
10
  tags:
11
+ - heretic
12
+ - uncensored
13
+ - abliteration
14
+ - mlx
15
+ - apple-silicon
16
+ - qwen3.5
17
  pipeline_tag: text-generation
 
18
  ---
19
 
20
+ # Qwythos-9B-v2-Heretic-MLX-8bit
21
+
22
+ **MLX 8-bit quantization** of [`WaveCut/Qwythos-9B-v2-Heretic`](https://huggingface.co/WaveCut/Qwythos-9B-v2-Heretic) — the Heretic-decensored version of [`empero-ai/Qwythos-9B-v2`](https://huggingface.co/empero-ai/Qwythos-9B-v2). Built for **Apple Silicon** (M1/M2/M3/M4).
23
+
24
+ ## Specs
25
 
26
+ | Field | Value |
27
+ |---|---|
28
+ | Bits/weight | **8.501** |
29
+ | File size | ~9.5 GB |
30
+ | Minimum RAM | ~11 GB unified memory |
31
 
32
+ ## Quantization
33
+
34
+ | Step | Tool | Version |
35
+ |---|---|---|
36
+ | Convert + quantize | `mlx_lm.convert` | **mlx-lm 0.31.3** (mlx 0.31.x) |
37
+ | Quant mode | `affine` (default) | `-q --q-bits 8` |
38
 
39
  ```bash
40
+ python -m mlx_lm.convert \
41
+ --hf-path WaveCut/Qwythos-9B-v2-Heretic \
42
+ -q --q-bits 8 \
43
+ --upload-repo WaveCut/Qwythos-9B-v2-Heretic-MLX-8bit
44
  ```
45
 
46
+ ## Usage
47
+
48
  ```python
49
  from mlx_lm import load, generate
 
50
  model, tokenizer = load("WaveCut/Qwythos-9B-v2-Heretic-MLX-8bit")
51
+ response = generate(model, tokenizer, prompt="Hello", max_tokens=256)
52
+ print(response)
53
+ ```
54
+
55
+ ```bash
56
+ # CLI
57
+ mlx_lm.generate --model WaveCut/Qwythos-9B-v2-Heretic-MLX-8bit --prompt "Hello"
58
+ ```
59
 
60
+ ## Architecture
61
 
62
+ Qwen3.5 **hybrid** — 32 blocks mixing attention and SSM (Mamba-style) layers. Supported in mlx-lm ≥ 0.31.0.
 
 
 
 
63
 
64
+ ## Disclaimer
65
+
66
+ Uncensored (safety alignment removed via Heretic). The original `empero-ai/Qwythos-9B-v2` maintainers are not affiliated with this derivative. Use responsibly.