spicyneuron commited on
Commit
7f27410
·
verified ·
1 Parent(s): 027348d

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +6 -8
README.md CHANGED
@@ -38,11 +38,13 @@ uvx --from mlx-lm mlx_lm.server \
38
  | peak memory (1024/512) | 291.257 | 272.358 | 341.020 | 424.067 |
39
  | prompt tok/s (1024) | 194.958 ± 0.075 | 194.216 ± 0.167 | 190.508 ± 0.880 | 193.563 ± 0.094 |
40
  | gen tok/s (512) | 21.381 ± 0.050 | 19.527 ± 0.035 | 17.873 ± 0.156 | 17.259 ± 0.032 |
41
- | kl mean | 0.686 ± 0.054 | 0.268 ± 0.009 | 0.117 ± 0.004 | 0.048 ± 0.002 |
42
- | kl p95 | 1.478 ± 0.054 | 0.537 ± 0.009 | 0.236 ± 0.004 | 0.097 ± 0.002 |
43
  | perplexity | 4.780 ± 0.020 | 4.118 ± 0.016 | 3.945 ± 0.016 | 3.920 ± 0.016 |
44
  | piqa | 0.776 ± 0.010 | 0.794 ± 0.009 | 0.820 ± 0.017 | 0.814 ± 0.017 |
45
 
 
 
46
  Tested on a Mac Studio M3 Ultra with:
47
 
48
  ```
@@ -52,10 +54,7 @@ mlx_lm.benchmark --prompt-tokens 1024 --generation-tokens 512 --num-trials 5
52
  mlx_lm.evaluate --tasks piqa --seed 123 --num-shots 0 --limit 500
53
  ```
54
 
55
- Note:
56
-
57
- - `mlx_lm.kld` is approximate, based on `top_k` not full logits. Here's the [code](https://github.com/ml-explore/mlx-lm/pull/1146).
58
- - GLM 5.1 KL divergence calculated against the largest quant I could run locally (~495 GB), so real KL is higher.
59
 
60
  # Methodology
61
 
@@ -65,5 +64,4 @@ MLX quantization options differ from llama.cpp, but the principles are the
65
  same:
66
 
67
  - Sensitive layers like MoE routing, attention, and output embeddings get higher precision
68
- - More tolerant layers like MoE experts get lower precision
69
-
 
38
  | peak memory (1024/512) | 291.257 | 272.358 | 341.020 | 424.067 |
39
  | prompt tok/s (1024) | 194.958 ± 0.075 | 194.216 ± 0.167 | 190.508 ± 0.880 | 193.563 ± 0.094 |
40
  | gen tok/s (512) | 21.381 ± 0.050 | 19.527 ± 0.035 | 17.873 ± 0.156 | 17.259 ± 0.032 |
41
+ | kl mean\* | 0.686 ± 0.054 | 0.268 ± 0.009 | 0.117 ± 0.004 | 0.048 ± 0.002 |
42
+ | kl p95\* | 1.478 ± 0.054 | 0.537 ± 0.009 | 0.236 ± 0.004 | 0.097 ± 0.002 |
43
  | perplexity | 4.780 ± 0.020 | 4.118 ± 0.016 | 3.945 ± 0.016 | 3.920 ± 0.016 |
44
  | piqa | 0.776 ± 0.010 | 0.794 ± 0.009 | 0.820 ± 0.017 | 0.814 ± 0.017 |
45
 
46
+ \* GLM 5.1 KL divergence calculated against the largest quant I could run locally (~495 GB), so real KL is higher.
47
+
48
  Tested on a Mac Studio M3 Ultra with:
49
 
50
  ```
 
54
  mlx_lm.evaluate --tasks piqa --seed 123 --num-shots 0 --limit 500
55
  ```
56
 
57
+ `mlx_lm.kld` is approximate, based on `top_k` not full logits. Here's the [code](https://github.com/ml-explore/mlx-lm/pull/1146).
 
 
 
58
 
59
  # Methodology
60
 
 
64
  same:
65
 
66
  - Sensitive layers like MoE routing, attention, and output embeddings get higher precision
67
+ - More tolerant layers like MoE experts get lower precision