ubergarm commited on
Commit
e9891e6
·
1 Parent(s): 98fa7cb

add new recipe big-IQ2_KS

Browse files

Kinda strange mix, but those gate/up seem very happy at iq3_ks and not
so much iq4_kss...

Routed Experts Layers [3-92]
blk\..*\.ffn_down_exps\.weight=iq5_ks
blk\..*\.ffn_(gate|up)_exps\.weight=iq3_ks

Files changed (2) hide show
  1. README.md +58 -0
  2. images/perplexity.png +2 -2
README.md CHANGED
@@ -161,6 +161,64 @@ numactl -N ${SOCKET} -m ${SOCKET} \
161
 
162
  </details>
163
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
164
  ## IQ3_KS 155.219 GiB (3.721 BPW)
165
  Final estimate: PPL over 565 chunks for n_ctx=512 = 4.1330 +/- 0.02573
166
 
 
161
 
162
  </details>
163
 
164
+ ## big-IQ3_KS 171.890 GiB (4.120 BPW)
165
+ Final estimate: PPL over 565 chunks for n_ctx=512 = 4.0410 +/- 0.02501
166
+
167
+ <details>
168
+
169
+ <summary>👈 Secret Recipe</summary>
170
+
171
+ ```bash
172
+ #!/usr/bin/env bash
173
+
174
+ custom="
175
+ # 93 Repeating Layers [0-92]
176
+
177
+ # Attention
178
+ blk\..*\.attn_q.*=q8_0
179
+ blk\..*\.attn_k.*=q8_0
180
+ blk\..*\.attn_v.*=q8_0
181
+ blk\..*\.attn_output.*=q8_0
182
+
183
+ # First 3 Dense Layers [0-2]
184
+ blk\..*\.ffn_down\.weight=q8_0
185
+ blk\..*\.ffn_(gate|up)\.weight=q8_0
186
+
187
+ # Shared Expert Layers [3-92]
188
+ blk\..*\.ffn_down_shexp\.weight=q8_0
189
+ blk\..*\.ffn_(gate|up)_shexp\.weight=q8_0
190
+
191
+ # Routed Experts Layers [3-92]
192
+ blk\..*\.ffn_down_exps\.weight=iq5_ks
193
+ blk\..*\.ffn_(gate|up)_exps\.weight=iq3_ks
194
+
195
+ # NextN MTP Layer [92]
196
+ blk\..*\.nextn\.embed_tokens\.weight=q8_0
197
+ blk\..*\.nextn\.shared_head_head\.weight=q8_0
198
+ blk\..*\.nextn\.eh_proj\.weight=q8_0
199
+
200
+ # Non-Repeating Layers
201
+ token_embd\.weight=iq6_k
202
+ output\.weight=iq6_k
203
+ "
204
+
205
+ custom=$(
206
+ echo "$custom" | grep -v '^#' | \
207
+ sed -Ez 's:\n+:,:g;s:,$::;s:^,::'
208
+ )
209
+
210
+ numactl -N ${SOCKET} -m ${SOCKET} \
211
+ ./build/bin/llama-quantize \
212
+ --custom-q "$custom" \
213
+ --imatrix /mnt/data/models/ubergarm/GLM-4.7-GGUF/imatrix-GLM-4.7-BF16.dat \
214
+ /mnt/data/models/ubergarm/GLM-4.7-GGUF/GLM-160x21B-4.7-BF16-00001-of-00015.gguf \
215
+ /mnt/data/models/ubergarm/GLM-4.7-GGUF/GLM-4.7-big-IQ3_KS.gguf \
216
+ IQ3_KS \
217
+ 128
218
+ ```
219
+
220
+ </details>
221
+
222
  ## IQ3_KS 155.219 GiB (3.721 BPW)
223
  Final estimate: PPL over 565 chunks for n_ctx=512 = 4.1330 +/- 0.02573
224
 
images/perplexity.png CHANGED

Git LFS Details

  • SHA256: 3ffce9bc62d30bbb91ef58e65fcc59b96a910f4520604b82f045f3735ba7af48
  • Pointer size: 131 Bytes
  • Size of remote file: 152 kB

Git LFS Details

  • SHA256: 30a2cbca94c54d57f52db8b92615d4ecb1518a5b024f749012fe1fb5b6f5772e
  • Pointer size: 131 Bytes
  • Size of remote file: 160 kB