ubergarm commited on
Commit
75dda6d
·
1 Parent(s): 4c458bf

uploading IQ5_K

Browse files
Files changed (1) hide show
  1. README.md +63 -3
README.md CHANGED
@@ -22,8 +22,9 @@ Currently cooking this now!
22
  - [x] calculate imatrix and upload to HF first so others can use as desired
23
  - [x] cook Q8_0 and test perplexity of BF16 and Q8_0 for baseline data
24
  - [x] adjust MTP nextn tensors to full q8_0 (won't effect RAM+VRAM usage otherwise)
25
- - [ ] cook IQ5_K with full q8_0 attn/shexp/first 3 dense layers and test
26
- - [ ] upload IQ5_K if all looking good
 
27
  - [ ] continue with smaller quants
28
  - [ ] check if any folks open discussions with desired RAM/VRAM breakpoints
29
 
@@ -52,7 +53,8 @@ These first two are just test quants for baseline perplexity comparison:
52
  * `Q8_0` 354.794 GiB (8.505 BPW)
53
  - Final estimate: PPL over 565 chunks for n_ctx=512 = 3.9320 +/- 0.02428
54
 
55
- ## IQ5_K TODO
 
56
 
57
  <details>
58
 
@@ -111,6 +113,64 @@ numactl -N ${SOCKET} -m ${SOCKET} \
111
 
112
  </details>
113
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
114
  ## Quick Start
115
  ```bash
116
  # Clone and checkout
 
22
  - [x] calculate imatrix and upload to HF first so others can use as desired
23
  - [x] cook Q8_0 and test perplexity of BF16 and Q8_0 for baseline data
24
  - [x] adjust MTP nextn tensors to full q8_0 (won't effect RAM+VRAM usage otherwise)
25
+ - [x] cook IQ5_K with full q8_0 attn/shexp/first 3 dense layers and test
26
+ - [x] upload IQ5_K if all looking good
27
+ - [ ] upload smol-IQ4_KSS if all looking good
28
  - [ ] continue with smaller quants
29
  - [ ] check if any folks open discussions with desired RAM/VRAM breakpoints
30
 
 
53
  * `Q8_0` 354.794 GiB (8.505 BPW)
54
  - Final estimate: PPL over 565 chunks for n_ctx=512 = 3.9320 +/- 0.02428
55
 
56
+ ## IQ5_K 250.635 GiB (6.008 BPW)
57
+ Final estimate: PPL over 565 chunks for n_ctx=512 = 3.9445 +/- 0.02439
58
 
59
  <details>
60
 
 
113
 
114
  </details>
115
 
116
+ ## smol-IQ4_KSS TODO
117
+ Final estimate: PPL over 565 chunks for n_ctx=512 = TODO
118
+
119
+ <details>
120
+
121
+ <summary>👈 Secret Recipe</summary>
122
+
123
+ ```bash
124
+ #!/usr/bin/env bash
125
+
126
+ custom="
127
+ # 93 Repeating Layers [0-92]
128
+
129
+ # Attention
130
+ blk\..*\.attn_q.*=q8_0
131
+ blk\..*\.attn_k.*=q8_0
132
+ blk\..*\.attn_v.*=q8_0
133
+ blk\..*\.attn_output.*=q8_0
134
+
135
+ # First 3 Dense Layers [0-2]
136
+ blk\..*\.ffn_down\.weight=q8_0
137
+ blk\..*\.ffn_(gate|up)\.weight=q8_0
138
+
139
+ # Shared Expert Layers [3-92]
140
+ blk\..*\.ffn_down_shexp\.weight=q8_0
141
+ blk\..*\.ffn_(gate|up)_shexp\.weight=q8_0
142
+
143
+ # Routed Experts Layers [3-92]
144
+ blk\..*\.ffn_down_exps\.weight=iq4_kss
145
+ blk\..*\.ffn_(gate|up)_exps\.weight=iq4_kss
146
+
147
+ # NextN MTP Layer [92]
148
+ blk\..*\.nextn\.embed_tokens\.weight=q8_0
149
+ blk\..*\.nextn\.shared_head_head\.weight=q8_0
150
+ blk\..*\.nextn\.eh_proj\.weight=q8_0
151
+
152
+ # Non-Repeating Layers
153
+ token_embd\.weight=iq4_k
154
+ output\.weight=iq6_k
155
+ "
156
+
157
+ custom=$(
158
+ echo "$custom" | grep -v '^#' | \
159
+ sed -Ez 's:\n+:,:g;s:,$::;s:^,::'
160
+ )
161
+
162
+ numactl -N ${SOCKET} -m ${SOCKET} \
163
+ ./build/bin/llama-quantize \
164
+ --custom-q "$custom" \
165
+ --imatrix /mnt/data/models/ubergarm/GLM-4.7-GGUF/imatrix-GLM-4.7-BF16.dat \
166
+ /mnt/data/models/ubergarm/GLM-4.7-GGUF/GLM-160x21B-4.7-BF16-00001-of-00015.gguf \
167
+ /mnt/data/models/ubergarm/GLM-4.7-GGUF/GLM-4.7-smol-IQ4_KSS.gguf \
168
+ IQ4_KSS \
169
+ 128
170
+ ```
171
+
172
+ </details>
173
+
174
  ## Quick Start
175
  ```bash
176
  # Clone and checkout