spicyneuron commited on
Commit
027348d
·
verified ·
1 Parent(s): dd47df0

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +14 -13
README.md CHANGED
@@ -11,8 +11,9 @@ tags:
11
  ---
12
 
13
  [GLM 5.1](https://huggingface.co/zai-org/GLM-5.1) optimized to run _comfortably_
14
- on a Mac Studio M3 512. This is the quality-first version. Smaller, compact
15
- version [here](https://huggingface.co/spicyneuron/GLM-5.1-MLX-2.9bit).
 
16
 
17
  - A mixed-precision quant that balances speed, memory, and accuracy.
18
  - 3-bit baseline with important layers at 4, 8 and BF16.
@@ -30,17 +31,17 @@ uvx --from mlx-lm mlx_lm.server \
30
 
31
  # Benchmarks
32
 
33
- metric | baa-ai/GLM-5.1-RAM-270GB-MLX | 2.9 bit | 3.6 bit (this model)
34
- --- | --- | --- | ---
35
- bpw | 3.110 | 2.906 | 3.645
36
- base memory | 269.303 | 251.702 | 315.648
37
- peak memory (1024/512) | 291.257 | 272.358 | 341.020
38
- prompt tok/s (1024) | 194.958 ± 0.075 | 194.216 ± 0.167 | 190.508 ± 0.880
39
- gen tok/s (512) | 21.381 ± 0.050 | 19.527 ± 0.035 | 17.873 ± 0.156
40
- kl mean | 0.686 ± 0.054 | 0.268 ± 0.009 | 0.117 ± 0.004
41
- kl p95 | 1.478 ± 0.054 | 0.537 ± 0.009 | 0.236 ± 0.004
42
- perplexity | 4.780 ± 0.020 | 4.118 ± 0.016 | 3.945 ± 0.016
43
- piqa | 0.776 ± 0.010 | 0.794 ± 0.009 | 0.820 ± 0.017
44
 
45
  Tested on a Mac Studio M3 Ultra with:
46
 
 
11
  ---
12
 
13
  [GLM 5.1](https://huggingface.co/zai-org/GLM-5.1) optimized to run _comfortably_
14
+ on a Mac Studio M3 512. This is the balanced version. Alternatives:
15
+ [speed-first](https://huggingface.co/spicyneuron/GLM-5.1-MLX-2.9bit),
16
+ [quality-first](https://huggingface.co/spicyneuron/GLM-5.1-MLX-4.5bit)
17
 
18
  - A mixed-precision quant that balances speed, memory, and accuracy.
19
  - 3-bit baseline with important layers at 4, 8 and BF16.
 
31
 
32
  # Benchmarks
33
 
34
+ | metric | baa-ai/GLM-5.1-RAM-270GB-MLX | 2.9 bit | 3.6 bit (this model) | 4.5 bit |
35
+ |---|---|---|---|---|
36
+ | bpw | 3.110 | 2.906 | 3.645 | 4.538 |
37
+ | base memory | 269.303 | 251.702 | 315.648 | 392.992 |
38
+ | peak memory (1024/512) | 291.257 | 272.358 | 341.020 | 424.067 |
39
+ | prompt tok/s (1024) | 194.958 ± 0.075 | 194.216 ± 0.167 | 190.508 ± 0.880 | 193.563 ± 0.094 |
40
+ | gen tok/s (512) | 21.381 ± 0.050 | 19.527 ± 0.035 | 17.873 ± 0.156 | 17.259 ± 0.032 |
41
+ | kl mean | 0.686 ± 0.054 | 0.268 ± 0.009 | 0.117 ± 0.004 | 0.048 ± 0.002 |
42
+ | kl p95 | 1.478 ± 0.054 | 0.537 ± 0.009 | 0.236 ± 0.004 | 0.097 ± 0.002 |
43
+ | perplexity | 4.780 ± 0.020 | 4.118 ± 0.016 | 3.945 ± 0.016 | 3.920 ± 0.016 |
44
+ | piqa | 0.776 ± 0.010 | 0.794 ± 0.009 | 0.820 ± 0.017 | 0.814 ± 0.017 |
45
 
46
  Tested on a Mac Studio M3 Ultra with:
47