Qwen3.8-27B-IQ1_M-GGUF / QUANTIZATION_LOG.txt
MarxistLeninist's picture
Add files using upload-large-folder tool
5e5d861 verified
Raw
History Blame Contribute Delete
132 kB
llama_print_build_info: build = 1 (62061f910)
llama_print_build_info: built with GNU 13.3.0 for Linux x86_64
llama_quantize: quantizing '/root/qwen38_bf16/BF16/Qwen3.8-27B-BF16-00001-of-00002.gguf' to '/mnt/geth-vol1/qwen38_out/Qwen3.8-27B-IQ1_M.gguf.part' as IQ1_M using 16 threads
llama_model_loader: additional 1 GGUFs metadata loaded.
llama_model_loader: loaded meta data with 50 key-value pairs and 866 tensors from /root/qwen38_bf16/BF16/Qwen3.8-27B-BF16-00001-of-00002.gguf (version GGUF V3 (latest))
llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
llama_model_loader: - kv 0: general.architecture str = qwen35
llama_model_loader: - kv 1: general.type str = model
llama_model_loader: - kv 2: general.sampling.top_k i32 = 20
llama_model_loader: - kv 3: general.sampling.top_p f32 = 0.950000
llama_model_loader: - kv 4: general.sampling.temp f32 = 1.000000
llama_model_loader: - kv 5: general.name str = Qwen3.8-27B
llama_model_loader: - kv 6: general.basename str = Qwen3.8-27B
llama_model_loader: - kv 7: general.description str = Renewal of the beloved Qwen model, de...
llama_model_loader: - kv 8: general.quantized_by str = Unsloth
llama_model_loader: - kv 9: general.size_label str = 27B
llama_model_loader: - kv 10: general.license str = apache-2.0
llama_model_loader: - kv 11: general.repo_url str = https://huggingface.co/unsloth
llama_model_loader: - kv 12: general.base_model.count u32 = 1
llama_model_loader: - kv 13: general.base_model.0.name str = Qwen3.8 27B
llama_model_loader: - kv 14: general.base_model.0.organization str = Qwen
llama_model_loader: - kv 15: general.base_model.0.repo_url str = https://huggingface.co/Qwen/Qwen3.8-27B
llama_model_loader: - kv 16: general.tags arr[str,1] = ["unsloth"]
llama_model_loader: - kv 17: qwen35.block_count u32 = 65
llama_model_loader: - kv 18: qwen35.context_length u32 = 262144
llama_model_loader: - kv 19: qwen35.embedding_length u32 = 5120
llama_model_loader: - kv 20: qwen35.feed_forward_length u32 = 17408
llama_model_loader: - kv 21: qwen35.attention.head_count u32 = 24
llama_model_loader: - kv 22: qwen35.attention.head_count_kv u32 = 4
llama_model_loader: - kv 23: qwen35.rope.dimension_sections arr[i32,4] = [11, 11, 10, 0]
llama_model_loader: - kv 24: qwen35.rope.freq_base f32 = 10000000.000000
llama_model_loader: - kv 25: qwen35.attention.layer_norm_rms_epsilon f32 = 0.000001
llama_model_loader: - kv 26: qwen35.attention.key_length u32 = 256
llama_model_loader: - kv 27: qwen35.attention.value_length u32 = 256
llama_model_loader: - kv 28: general.file_type u32 = 32
llama_model_loader: - kv 29: qwen35.nextn_predict_layers u32 = 1
llama_model_loader: - kv 30: qwen35.ssm.conv_kernel u32 = 4
llama_model_loader: - kv 31: qwen35.ssm.state_size u32 = 128
llama_model_loader: - kv 32: qwen35.ssm.group_count u32 = 16
llama_model_loader: - kv 33: qwen35.ssm.time_step_rank u32 = 48
llama_model_loader: - kv 34: qwen35.ssm.inner_size u32 = 6144
llama_model_loader: - kv 35: qwen35.full_attention_interval u32 = 4
llama_model_loader: - kv 36: qwen35.rope.dimension_count u32 = 64
llama_model_loader: - kv 37: general.quantization_version u32 = 2
llama_model_loader: - kv 38: tokenizer.ggml.model str = gpt2
llama_model_loader: - kv 39: tokenizer.ggml.pre str = qwen35
llama_model_loader: - kv 40: tokenizer.ggml.tokens arr[str,248320] = ["!", "\"", "#", "$", "%", "&", "'", ...
llama_model_loader: - kv 41: tokenizer.ggml.token_type arr[i32,248320] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
llama_model_loader: - kv 42: tokenizer.ggml.merges arr[str,247587] = ["Ġ Ġ", "ĠĠ ĠĠ", "i n", "Ġ t",...
llama_model_loader: - kv 43: tokenizer.ggml.eos_token_id u32 = 248046
llama_model_loader: - kv 44: tokenizer.ggml.padding_token_id u32 = 248055
llama_model_loader: - kv 45: tokenizer.ggml.bos_token_id u32 = 248044
llama_model_loader: - kv 46: split.no u16 = 0
llama_model_loader: - kv 47: split.count u16 = 2
llama_model_loader: - kv 48: split.tensors.count i32 = 866
llama_model_loader: - kv 49: tokenizer.chat_template str = {%- set image_count = namespace(value...
llama_model_loader: - type f32: 360 tensors
llama_model_loader: - type bf16: 506 tensors
llama_model_quantize_impl: have importance matrix data with 496 entries
llama_tensor_get_type: blk.64.attn_k.weight - applying manual override: iq1_m -> q4_K
llama_tensor_get_type: blk.64.attn_output.weight - applying manual override: iq1_m -> q4_K
llama_tensor_get_type: blk.64.attn_q.weight - applying manual override: iq1_m -> q4_K
llama_tensor_get_type: blk.64.attn_v.weight - applying manual override: iq1_m -> q4_K
llama_tensor_get_type: blk.64.ffn_down.weight - applying manual override: iq1_m -> q4_K
llama_tensor_get_type: blk.64.ffn_gate.weight - applying manual override: iq1_m -> q4_K
llama_tensor_get_type: blk.64.ffn_up.weight - applying manual override: iq1_m -> q4_K
llama_tensor_get_type: blk.64.nextn.eh_proj.weight - applying manual override: iq1_m -> q4_K
[ 1/ 866] output.weight - [ 5120, 248320, 1, 1], type = bf16,
====== llama_model_quantize_impl: did not find weights for output.weight
converting to q5_K .. load_imatrix: imatrix datasets=['/Volumes/Lexar/aq_ds4flash/imatrix_agentic.txt']
load_imatrix: loaded 496 importance matrix entries from /mnt/geth-vol1/qwen38_assets/Qwen3.8-27B-agentic.imatrix computed on 1580 chunks
prepare_imatrix: have 496 importance matrix entries
size = 2425.00 MiB -> 833.59 MiB
[ 2/ 866] output_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 3/ 866] token_embd.weight - [ 5120, 248320, 1, 1], type = bf16,
====== llama_model_quantize_impl: did not find weights for token_embd.weight
converting to q2_K .. size = 2425.00 MiB -> 397.85 MiB
[ 4/ 866] blk.0.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 5/ 866] blk.0.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 6/ 866] blk.0.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 7/ 866] blk.0.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q2_K .. size = 170.00 MiB -> 27.89 MiB
[ 8/ 866] blk.0.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 9/ 866] blk.0.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 10/ 866] blk.0.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 11/ 866] blk.0.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 12/ 866] blk.0.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 13/ 866] blk.0.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 14/ 866] blk.0.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 15/ 866] blk.0.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 16/ 866] blk.0.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 17/ 866] blk.0.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 18/ 866] blk.1.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 19/ 866] blk.1.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 20/ 866] blk.1.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 21/ 866] blk.1.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q2_K .. size = 170.00 MiB -> 27.89 MiB
[ 22/ 866] blk.1.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 23/ 866] blk.1.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 24/ 866] blk.1.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 25/ 866] blk.1.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 26/ 866] blk.1.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 27/ 866] blk.1.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 28/ 866] blk.1.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 29/ 866] blk.1.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 30/ 866] blk.1.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 31/ 866] blk.1.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 32/ 866] blk.2.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 33/ 866] blk.2.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 34/ 866] blk.2.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 35/ 866] blk.2.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q2_K .. size = 170.00 MiB -> 27.89 MiB
[ 36/ 866] blk.2.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 37/ 866] blk.2.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 38/ 866] blk.2.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 39/ 866] blk.2.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 40/ 866] blk.2.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 41/ 866] blk.2.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 42/ 866] blk.2.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 43/ 866] blk.2.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 44/ 866] blk.2.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 45/ 866] blk.2.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 46/ 866] blk.3.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to iq1_m .. size = 10.00 MiB -> 1.09 MiB
[ 47/ 866] blk.3.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 48/ 866] blk.3.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 49/ 866] blk.3.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq2_xxs .. size = 60.00 MiB -> 7.73 MiB
[ 50/ 866] blk.3.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to iq1_m .. size = 120.00 MiB -> 13.12 MiB
[ 51/ 866] blk.3.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 52/ 866] blk.3.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q4_K .. size = 10.00 MiB -> 2.81 MiB
[ 53/ 866] blk.3.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q2_K .. size = 170.00 MiB -> 27.89 MiB
[ 54/ 866] blk.3.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 55/ 866] blk.3.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 56/ 866] blk.3.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 57/ 866] blk.4.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 58/ 866] blk.4.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 59/ 866] blk.4.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 60/ 866] blk.4.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q2_K .. size = 170.00 MiB -> 27.89 MiB
[ 61/ 866] blk.4.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 62/ 866] blk.4.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 63/ 866] blk.4.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 64/ 866] blk.4.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 65/ 866] blk.4.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 66/ 866] blk.4.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 67/ 866] blk.4.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 68/ 866] blk.4.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 69/ 866] blk.4.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 70/ 866] blk.4.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 71/ 866] blk.5.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 72/ 866] blk.5.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 73/ 866] blk.5.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 74/ 866] blk.5.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q2_K .. size = 170.00 MiB -> 27.89 MiB
[ 75/ 866] blk.5.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 76/ 866] blk.5.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 77/ 866] blk.5.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 78/ 866] blk.5.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 79/ 866] blk.5.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 80/ 866] blk.5.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 81/ 866] blk.5.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 82/ 866] blk.5.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 83/ 866] blk.5.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 84/ 866] blk.5.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 85/ 866] blk.6.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 86/ 866] blk.6.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 87/ 866] blk.6.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 88/ 866] blk.6.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q2_K .. size = 170.00 MiB -> 27.89 MiB
[ 89/ 866] blk.6.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 90/ 866] blk.6.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 91/ 866] blk.6.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 92/ 866] blk.6.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 93/ 866] blk.6.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 94/ 866] blk.6.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 95/ 866] blk.6.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 96/ 866] blk.6.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 97/ 866] blk.6.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 98/ 866] blk.6.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 99/ 866] blk.7.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to iq1_m .. size = 10.00 MiB -> 1.09 MiB
[ 100/ 866] blk.7.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 101/ 866] blk.7.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 102/ 866] blk.7.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq2_xxs .. size = 60.00 MiB -> 7.73 MiB
[ 103/ 866] blk.7.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to iq1_m .. size = 120.00 MiB -> 13.12 MiB
[ 104/ 866] blk.7.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 105/ 866] blk.7.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q4_K .. size = 10.00 MiB -> 2.81 MiB
[ 106/ 866] blk.7.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q2_K .. size = 170.00 MiB -> 27.89 MiB
[ 107/ 866] blk.7.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 108/ 866] blk.7.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 109/ 866] blk.7.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 110/ 866] blk.8.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 111/ 866] blk.8.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 112/ 866] blk.8.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 113/ 866] blk.8.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 114/ 866] blk.8.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 115/ 866] blk.8.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 116/ 866] blk.8.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 117/ 866] blk.8.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 118/ 866] blk.8.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 119/ 866] blk.8.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 120/ 866] blk.8.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 121/ 866] blk.8.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 122/ 866] blk.8.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 123/ 866] blk.8.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 124/ 866] blk.9.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 125/ 866] blk.9.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 126/ 866] blk.9.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 127/ 866] blk.9.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 128/ 866] blk.9.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 129/ 866] blk.9.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 130/ 866] blk.9.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 131/ 866] blk.9.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 132/ 866] blk.9.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 133/ 866] blk.9.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 134/ 866] blk.9.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 135/ 866] blk.9.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 136/ 866] blk.9.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 137/ 866] blk.9.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 138/ 866] blk.10.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 139/ 866] blk.10.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 140/ 866] blk.10.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 141/ 866] blk.10.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 142/ 866] blk.10.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 143/ 866] blk.10.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 144/ 866] blk.10.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 145/ 866] blk.10.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 146/ 866] blk.10.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 147/ 866] blk.10.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 148/ 866] blk.10.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 149/ 866] blk.10.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 150/ 866] blk.10.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 151/ 866] blk.10.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 152/ 866] blk.11.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to iq1_m .. size = 10.00 MiB -> 1.09 MiB
[ 153/ 866] blk.11.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 154/ 866] blk.11.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 155/ 866] blk.11.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq2_xxs .. size = 60.00 MiB -> 7.73 MiB
[ 156/ 866] blk.11.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to iq1_m .. size = 120.00 MiB -> 13.12 MiB
[ 157/ 866] blk.11.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 158/ 866] blk.11.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q4_K .. size = 10.00 MiB -> 2.81 MiB
[ 159/ 866] blk.11.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 160/ 866] blk.11.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 161/ 866] blk.11.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 162/ 866] blk.11.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 163/ 866] blk.12.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 164/ 866] blk.12.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 165/ 866] blk.12.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 166/ 866] blk.12.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 167/ 866] blk.12.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 168/ 866] blk.12.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 169/ 866] blk.12.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 170/ 866] blk.12.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 171/ 866] blk.12.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 172/ 866] blk.12.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 173/ 866] blk.12.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 174/ 866] blk.12.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 175/ 866] blk.12.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 176/ 866] blk.12.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 177/ 866] blk.13.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 178/ 866] blk.13.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 179/ 866] blk.13.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 180/ 866] blk.13.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 181/ 866] blk.13.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 182/ 866] blk.13.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 183/ 866] blk.13.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 184/ 866] blk.13.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 185/ 866] blk.13.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 186/ 866] blk.13.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 187/ 866] blk.13.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 188/ 866] blk.13.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 189/ 866] blk.13.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 190/ 866] blk.13.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 191/ 866] blk.14.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 192/ 866] blk.14.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 193/ 866] blk.14.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 194/ 866] blk.14.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 195/ 866] blk.14.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 196/ 866] blk.14.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 197/ 866] blk.14.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 198/ 866] blk.14.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 199/ 866] blk.14.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 200/ 866] blk.14.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 201/ 866] blk.14.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 202/ 866] blk.14.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 203/ 866] blk.14.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 204/ 866] blk.14.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 205/ 866] blk.15.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to iq1_m .. size = 10.00 MiB -> 1.09 MiB
[ 206/ 866] blk.15.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 207/ 866] blk.15.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 208/ 866] blk.15.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq2_xxs .. size = 60.00 MiB -> 7.73 MiB
[ 209/ 866] blk.15.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to iq1_m .. size = 120.00 MiB -> 13.12 MiB
[ 210/ 866] blk.15.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 211/ 866] blk.15.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q4_K .. size = 10.00 MiB -> 2.81 MiB
[ 212/ 866] blk.15.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 213/ 866] blk.15.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 214/ 866] blk.15.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 215/ 866] blk.15.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 216/ 866] blk.16.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 217/ 866] blk.16.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 218/ 866] blk.16.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 219/ 866] blk.16.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 220/ 866] blk.16.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 221/ 866] blk.16.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 222/ 866] blk.16.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 223/ 866] blk.16.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 224/ 866] blk.16.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 225/ 866] blk.16.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 226/ 866] blk.16.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 227/ 866] blk.16.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 228/ 866] blk.16.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 229/ 866] blk.16.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 230/ 866] blk.17.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 231/ 866] blk.17.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 232/ 866] blk.17.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 233/ 866] blk.17.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 234/ 866] blk.17.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 235/ 866] blk.17.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 236/ 866] blk.17.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 237/ 866] blk.17.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 238/ 866] blk.17.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 239/ 866] blk.17.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 240/ 866] blk.17.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 241/ 866] blk.17.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 242/ 866] blk.17.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 243/ 866] blk.17.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 244/ 866] blk.18.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 245/ 866] blk.18.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 246/ 866] blk.18.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 247/ 866] blk.18.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 248/ 866] blk.18.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 249/ 866] blk.18.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 250/ 866] blk.18.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 251/ 866] blk.18.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 252/ 866] blk.18.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 253/ 866] blk.18.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 254/ 866] blk.18.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 255/ 866] blk.18.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 256/ 866] blk.18.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 257/ 866] blk.18.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 258/ 866] blk.19.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to iq1_m .. size = 10.00 MiB -> 1.09 MiB
[ 259/ 866] blk.19.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 260/ 866] blk.19.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 261/ 866] blk.19.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq2_xxs .. size = 60.00 MiB -> 7.73 MiB
[ 262/ 866] blk.19.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to iq1_m .. size = 120.00 MiB -> 13.12 MiB
[ 263/ 866] blk.19.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 264/ 866] blk.19.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q4_K .. size = 10.00 MiB -> 2.81 MiB
[ 265/ 866] blk.19.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 266/ 866] blk.19.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 267/ 866] blk.19.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 268/ 866] blk.19.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 269/ 866] blk.20.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 270/ 866] blk.20.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 271/ 866] blk.20.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 272/ 866] blk.20.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 273/ 866] blk.20.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 274/ 866] blk.20.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 275/ 866] blk.20.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 276/ 866] blk.20.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 277/ 866] blk.20.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 278/ 866] blk.20.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 279/ 866] blk.20.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 280/ 866] blk.20.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 281/ 866] blk.20.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 282/ 866] blk.20.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 283/ 866] blk.21.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 284/ 866] blk.21.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 285/ 866] blk.21.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 286/ 866] blk.21.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 287/ 866] blk.21.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 288/ 866] blk.21.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 289/ 866] blk.21.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 290/ 866] blk.21.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 291/ 866] blk.21.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 292/ 866] blk.21.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 293/ 866] blk.21.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 294/ 866] blk.21.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 295/ 866] blk.21.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 296/ 866] blk.21.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 297/ 866] blk.22.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 298/ 866] blk.22.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 299/ 866] blk.22.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 300/ 866] blk.22.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 301/ 866] blk.22.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 302/ 866] blk.22.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 303/ 866] blk.22.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 304/ 866] blk.22.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 305/ 866] blk.22.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 306/ 866] blk.22.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 307/ 866] blk.22.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 308/ 866] blk.22.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 309/ 866] blk.22.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 310/ 866] blk.22.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 311/ 866] blk.23.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to iq1_m .. size = 10.00 MiB -> 1.09 MiB
[ 312/ 866] blk.23.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 313/ 866] blk.23.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 314/ 866] blk.23.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq2_xxs .. size = 60.00 MiB -> 7.73 MiB
[ 315/ 866] blk.23.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to iq1_m .. size = 120.00 MiB -> 13.12 MiB
[ 316/ 866] blk.23.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 317/ 866] blk.23.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q4_K .. size = 10.00 MiB -> 2.81 MiB
[ 318/ 866] blk.23.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 319/ 866] blk.23.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 320/ 866] blk.23.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 321/ 866] blk.23.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 322/ 866] blk.24.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 323/ 866] blk.24.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 324/ 866] blk.24.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 325/ 866] blk.24.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 326/ 866] blk.24.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 327/ 866] blk.24.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 328/ 866] blk.24.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 329/ 866] blk.24.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 330/ 866] blk.24.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 331/ 866] blk.24.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 332/ 866] blk.24.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 333/ 866] blk.24.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 334/ 866] blk.24.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 335/ 866] blk.24.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 336/ 866] blk.25.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 337/ 866] blk.25.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 338/ 866] blk.25.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 339/ 866] blk.25.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 340/ 866] blk.25.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 341/ 866] blk.25.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 342/ 866] blk.25.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 343/ 866] blk.25.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 344/ 866] blk.25.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 345/ 866] blk.25.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 346/ 866] blk.25.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 347/ 866] blk.25.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 348/ 866] blk.25.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 349/ 866] blk.25.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 350/ 866] blk.26.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 351/ 866] blk.26.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 352/ 866] blk.26.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 353/ 866] blk.26.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 354/ 866] blk.26.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 355/ 866] blk.26.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 356/ 866] blk.26.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 357/ 866] blk.26.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 358/ 866] blk.26.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 359/ 866] blk.26.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 360/ 866] blk.26.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 361/ 866] blk.26.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 362/ 866] blk.26.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 363/ 866] blk.26.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 364/ 866] blk.27.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to iq1_m .. size = 10.00 MiB -> 1.09 MiB
[ 365/ 866] blk.27.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 366/ 866] blk.27.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 367/ 866] blk.27.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq2_xxs .. size = 60.00 MiB -> 7.73 MiB
[ 368/ 866] blk.27.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to iq1_m .. size = 120.00 MiB -> 13.12 MiB
[ 369/ 866] blk.27.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 370/ 866] blk.27.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q4_K .. size = 10.00 MiB -> 2.81 MiB
[ 371/ 866] blk.27.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 372/ 866] blk.27.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 373/ 866] blk.27.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 374/ 866] blk.27.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 375/ 866] blk.28.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 376/ 866] blk.28.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 377/ 866] blk.28.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 378/ 866] blk.28.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 379/ 866] blk.28.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 380/ 866] blk.28.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 381/ 866] blk.28.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 382/ 866] blk.28.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 383/ 866] blk.28.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 384/ 866] blk.28.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 385/ 866] blk.28.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 386/ 866] blk.28.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 387/ 866] blk.28.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 388/ 866] blk.28.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 389/ 866] blk.29.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 390/ 866] blk.29.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 391/ 866] blk.29.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 392/ 866] blk.29.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 393/ 866] blk.29.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 394/ 866] blk.29.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 395/ 866] blk.29.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 396/ 866] blk.29.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 397/ 866] blk.29.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 398/ 866] blk.29.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 399/ 866] blk.29.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 400/ 866] blk.29.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 401/ 866] blk.29.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 402/ 866] blk.29.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 403/ 866] blk.30.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 404/ 866] blk.30.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 405/ 866] blk.30.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 406/ 866] blk.30.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 407/ 866] blk.30.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 408/ 866] blk.30.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 409/ 866] blk.30.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 410/ 866] blk.30.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 411/ 866] blk.30.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 412/ 866] blk.30.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 413/ 866] blk.30.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 414/ 866] blk.30.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 415/ 866] blk.30.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 416/ 866] blk.30.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 417/ 866] blk.31.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to iq1_m .. size = 10.00 MiB -> 1.09 MiB
[ 418/ 866] blk.31.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 419/ 866] blk.31.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 420/ 866] blk.31.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq2_xxs .. size = 60.00 MiB -> 7.73 MiB
[ 421/ 866] blk.31.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to iq1_m .. size = 120.00 MiB -> 13.12 MiB
[ 422/ 866] blk.31.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 423/ 866] blk.31.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q4_K .. size = 10.00 MiB -> 2.81 MiB
[ 424/ 866] blk.31.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 425/ 866] blk.31.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 426/ 866] blk.31.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 427/ 866] blk.31.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 428/ 866] blk.32.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 429/ 866] blk.32.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 430/ 866] blk.32.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 431/ 866] blk.32.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 432/ 866] blk.32.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 433/ 866] blk.32.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 434/ 866] blk.32.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 435/ 866] blk.32.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 436/ 866] blk.32.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 437/ 866] blk.32.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 438/ 866] blk.32.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 439/ 866] blk.32.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 440/ 866] blk.32.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 441/ 866] blk.32.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 442/ 866] blk.33.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 443/ 866] blk.33.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 444/ 866] blk.33.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 445/ 866] blk.33.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 446/ 866] blk.33.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 447/ 866] blk.33.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 448/ 866] blk.33.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 449/ 866] blk.33.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 450/ 866] blk.33.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 451/ 866] blk.33.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 452/ 866] blk.33.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 453/ 866] blk.33.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 454/ 866] blk.33.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 455/ 866] blk.33.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 456/ 866] blk.34.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 457/ 866] blk.34.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 458/ 866] blk.34.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 459/ 866] blk.34.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 460/ 866] blk.34.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 461/ 866] blk.34.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 462/ 866] blk.34.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 463/ 866] blk.34.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 464/ 866] blk.34.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 465/ 866] blk.34.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 466/ 866] blk.34.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 467/ 866] blk.34.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 468/ 866] blk.34.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 469/ 866] blk.34.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 470/ 866] blk.35.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to iq1_m .. size = 10.00 MiB -> 1.09 MiB
[ 471/ 866] blk.35.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 472/ 866] blk.35.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 473/ 866] blk.35.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq2_xxs .. size = 60.00 MiB -> 7.73 MiB
[ 474/ 866] blk.35.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to iq1_m .. size = 120.00 MiB -> 13.12 MiB
[ 475/ 866] blk.35.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 476/ 866] blk.35.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q4_K .. size = 10.00 MiB -> 2.81 MiB
[ 477/ 866] blk.35.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 478/ 866] blk.35.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 479/ 866] blk.35.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 480/ 866] blk.35.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 481/ 866] blk.36.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 482/ 866] blk.36.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 483/ 866] blk.36.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 484/ 866] blk.36.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 485/ 866] blk.36.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 486/ 866] blk.36.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 487/ 866] blk.36.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 488/ 866] blk.36.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 489/ 866] blk.36.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 490/ 866] blk.36.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 491/ 866] blk.36.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 492/ 866] blk.36.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 493/ 866] blk.36.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 494/ 866] blk.36.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 495/ 866] blk.37.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 496/ 866] blk.37.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 497/ 866] blk.37.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 498/ 866] blk.37.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 499/ 866] blk.37.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 500/ 866] blk.37.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 501/ 866] blk.37.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 502/ 866] blk.37.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 503/ 866] blk.37.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 504/ 866] blk.37.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 505/ 866] blk.37.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 506/ 866] blk.37.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 507/ 866] blk.37.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 508/ 866] blk.37.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 509/ 866] blk.38.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 510/ 866] blk.38.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 511/ 866] blk.38.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 512/ 866] blk.38.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 513/ 866] blk.38.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 514/ 866] blk.38.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 515/ 866] blk.38.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 516/ 866] blk.38.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 517/ 866] blk.38.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 518/ 866] blk.38.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 519/ 866] blk.38.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 520/ 866] blk.38.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 521/ 866] blk.38.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 522/ 866] blk.38.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 523/ 866] blk.39.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to iq1_m .. size = 10.00 MiB -> 1.09 MiB
[ 524/ 866] blk.39.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 525/ 866] blk.39.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 526/ 866] blk.39.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq2_xxs .. size = 60.00 MiB -> 7.73 MiB
[ 527/ 866] blk.39.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to iq1_m .. size = 120.00 MiB -> 13.12 MiB
[ 528/ 866] blk.39.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 529/ 866] blk.39.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q4_K .. size = 10.00 MiB -> 2.81 MiB
[ 530/ 866] blk.39.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 531/ 866] blk.39.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 532/ 866] blk.39.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 533/ 866] blk.39.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 534/ 866] blk.40.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 535/ 866] blk.40.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 536/ 866] blk.40.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 537/ 866] blk.40.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 538/ 866] blk.40.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 539/ 866] blk.40.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 540/ 866] blk.40.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 541/ 866] blk.40.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 542/ 866] blk.40.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 543/ 866] blk.40.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 544/ 866] blk.40.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 545/ 866] blk.40.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 546/ 866] blk.40.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 547/ 866] blk.40.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 548/ 866] blk.41.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 549/ 866] blk.41.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 550/ 866] blk.41.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 551/ 866] blk.41.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 552/ 866] blk.41.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 553/ 866] blk.41.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 554/ 866] blk.41.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 555/ 866] blk.41.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 556/ 866] blk.41.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 557/ 866] blk.41.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 558/ 866] blk.41.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 559/ 866] blk.41.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 560/ 866] blk.41.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 561/ 866] blk.41.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 562/ 866] blk.42.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 563/ 866] blk.42.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 564/ 866] blk.42.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 565/ 866] blk.42.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 566/ 866] blk.42.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 567/ 866] blk.42.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 568/ 866] blk.42.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 569/ 866] blk.42.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 570/ 866] blk.42.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 571/ 866] blk.42.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 572/ 866] blk.42.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 573/ 866] blk.42.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 574/ 866] blk.42.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 575/ 866] blk.42.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 576/ 866] blk.43.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to iq1_m .. size = 10.00 MiB -> 1.09 MiB
[ 577/ 866] blk.43.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 578/ 866] blk.43.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 579/ 866] blk.43.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq2_xxs .. size = 60.00 MiB -> 7.73 MiB
[ 580/ 866] blk.43.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to iq1_m .. size = 120.00 MiB -> 13.12 MiB
[ 581/ 866] blk.43.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 582/ 866] blk.43.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q4_K .. size = 10.00 MiB -> 2.81 MiB
[ 583/ 866] blk.43.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 584/ 866] blk.43.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 585/ 866] blk.43.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 586/ 866] blk.43.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 587/ 866] blk.44.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 588/ 866] blk.44.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 589/ 866] blk.44.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 590/ 866] blk.44.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 591/ 866] blk.44.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 592/ 866] blk.44.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 593/ 866] blk.44.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 594/ 866] blk.44.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 595/ 866] blk.44.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 596/ 866] blk.44.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 597/ 866] blk.44.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 598/ 866] blk.44.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 599/ 866] blk.44.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 600/ 866] blk.44.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 601/ 866] blk.45.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 602/ 866] blk.45.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 603/ 866] blk.45.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 604/ 866] blk.45.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 605/ 866] blk.45.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 606/ 866] blk.45.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 607/ 866] blk.45.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 608/ 866] blk.45.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 609/ 866] blk.45.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 610/ 866] blk.45.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 611/ 866] blk.45.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 612/ 866] blk.45.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 613/ 866] blk.45.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 614/ 866] blk.45.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 615/ 866] blk.46.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 616/ 866] blk.46.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 617/ 866] blk.46.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 618/ 866] blk.46.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 619/ 866] blk.46.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 620/ 866] blk.46.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 621/ 866] blk.46.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 622/ 866] blk.46.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 623/ 866] blk.46.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 624/ 866] blk.46.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 625/ 866] blk.46.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 626/ 866] blk.46.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 627/ 866] blk.46.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 628/ 866] blk.46.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 629/ 866] blk.47.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to iq1_m .. size = 10.00 MiB -> 1.09 MiB
[ 630/ 866] blk.47.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 631/ 866] blk.47.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 632/ 866] blk.47.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq2_xxs .. size = 60.00 MiB -> 7.73 MiB
[ 633/ 866] blk.47.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to iq1_m .. size = 120.00 MiB -> 13.12 MiB
[ 634/ 866] blk.47.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 635/ 866] blk.47.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q4_K .. size = 10.00 MiB -> 2.81 MiB
[ 636/ 866] blk.47.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 637/ 866] blk.47.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 638/ 866] blk.47.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 639/ 866] blk.47.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 640/ 866] blk.48.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 641/ 866] blk.48.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 642/ 866] blk.48.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 643/ 866] blk.48.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 644/ 866] blk.48.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 645/ 866] blk.48.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 646/ 866] blk.48.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 647/ 866] blk.48.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 648/ 866] blk.48.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 649/ 866] blk.48.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 650/ 866] blk.48.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 651/ 866] blk.48.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 652/ 866] blk.48.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 653/ 866] blk.48.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 654/ 866] blk.49.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 655/ 866] blk.49.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 656/ 866] blk.49.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 657/ 866] blk.49.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 658/ 866] blk.49.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 659/ 866] blk.49.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 660/ 866] blk.49.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 661/ 866] blk.49.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 662/ 866] blk.49.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 663/ 866] blk.49.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 664/ 866] blk.49.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 665/ 866] blk.49.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 666/ 866] blk.49.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 667/ 866] blk.49.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 668/ 866] blk.50.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 669/ 866] blk.50.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 670/ 866] blk.50.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 671/ 866] blk.50.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 672/ 866] blk.50.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 673/ 866] blk.50.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 674/ 866] blk.50.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 675/ 866] blk.50.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 676/ 866] blk.50.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 677/ 866] blk.50.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 678/ 866] blk.50.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 679/ 866] blk.50.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 680/ 866] blk.50.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 681/ 866] blk.50.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 682/ 866] blk.51.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to iq1_m .. size = 10.00 MiB -> 1.09 MiB
[ 683/ 866] blk.51.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 684/ 866] blk.51.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 685/ 866] blk.51.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq2_xxs .. size = 60.00 MiB -> 7.73 MiB
[ 686/ 866] blk.51.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to iq1_m .. size = 120.00 MiB -> 13.12 MiB
[ 687/ 866] blk.51.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 688/ 866] blk.51.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q4_K .. size = 10.00 MiB -> 2.81 MiB
[ 689/ 866] blk.51.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 690/ 866] blk.51.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 691/ 866] blk.51.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 692/ 866] blk.51.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 693/ 866] blk.52.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 694/ 866] blk.52.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 695/ 866] blk.52.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 696/ 866] blk.52.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 697/ 866] blk.52.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 698/ 866] blk.52.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 699/ 866] blk.52.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 700/ 866] blk.52.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 701/ 866] blk.52.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 702/ 866] blk.52.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 703/ 866] blk.52.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 704/ 866] blk.52.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 705/ 866] blk.52.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 706/ 866] blk.52.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 707/ 866] blk.53.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 708/ 866] blk.53.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 709/ 866] blk.53.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 710/ 866] blk.53.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 711/ 866] blk.53.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 712/ 866] blk.53.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 713/ 866] blk.53.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 714/ 866] blk.53.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 715/ 866] blk.53.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 716/ 866] blk.53.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 717/ 866] blk.53.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 718/ 866] blk.53.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 719/ 866] blk.53.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 720/ 866] blk.53.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 721/ 866] blk.54.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 722/ 866] blk.54.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 723/ 866] blk.54.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 724/ 866] blk.54.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 725/ 866] blk.54.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 726/ 866] blk.54.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 727/ 866] blk.54.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 728/ 866] blk.54.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 729/ 866] blk.54.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 730/ 866] blk.54.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 731/ 866] blk.54.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 732/ 866] blk.54.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 733/ 866] blk.54.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 734/ 866] blk.54.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 735/ 866] blk.55.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to iq1_m .. size = 10.00 MiB -> 1.09 MiB
[ 736/ 866] blk.55.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 737/ 866] blk.55.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 738/ 866] blk.55.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq2_xxs .. size = 60.00 MiB -> 7.73 MiB
[ 739/ 866] blk.55.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to iq1_m .. size = 120.00 MiB -> 13.12 MiB
[ 740/ 866] blk.55.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 741/ 866] blk.55.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q4_K .. size = 10.00 MiB -> 2.81 MiB
[ 742/ 866] blk.55.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 743/ 866] blk.55.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 744/ 866] blk.55.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 745/ 866] blk.55.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 746/ 866] blk.56.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 747/ 866] blk.56.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 748/ 866] blk.56.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 749/ 866] blk.56.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 750/ 866] blk.56.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 751/ 866] blk.56.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 752/ 866] blk.56.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 753/ 866] blk.56.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 754/ 866] blk.56.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 755/ 866] blk.56.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 756/ 866] blk.56.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 757/ 866] blk.56.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 758/ 866] blk.56.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 759/ 866] blk.56.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 760/ 866] blk.57.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 761/ 866] blk.57.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 762/ 866] blk.57.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 763/ 866] blk.57.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 764/ 866] blk.57.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 765/ 866] blk.57.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 766/ 866] blk.57.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 767/ 866] blk.57.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 768/ 866] blk.57.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 769/ 866] blk.57.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 770/ 866] blk.57.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 771/ 866] blk.57.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 772/ 866] blk.57.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 773/ 866] blk.57.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 774/ 866] blk.58.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 775/ 866] blk.58.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 776/ 866] blk.58.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 777/ 866] blk.58.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 778/ 866] blk.58.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 779/ 866] blk.58.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 780/ 866] blk.58.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 781/ 866] blk.58.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 782/ 866] blk.58.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 783/ 866] blk.58.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 784/ 866] blk.58.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 785/ 866] blk.58.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 786/ 866] blk.58.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 787/ 866] blk.58.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 788/ 866] blk.59.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to iq1_m .. size = 10.00 MiB -> 1.09 MiB
[ 789/ 866] blk.59.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 790/ 866] blk.59.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 791/ 866] blk.59.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq2_xxs .. size = 60.00 MiB -> 7.73 MiB
[ 792/ 866] blk.59.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to iq1_m .. size = 120.00 MiB -> 13.12 MiB
[ 793/ 866] blk.59.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 794/ 866] blk.59.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q4_K .. size = 10.00 MiB -> 2.81 MiB
[ 795/ 866] blk.59.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 796/ 866] blk.59.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 797/ 866] blk.59.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 798/ 866] blk.59.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 799/ 866] blk.60.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 800/ 866] blk.60.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 801/ 866] blk.60.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 802/ 866] blk.60.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 803/ 866] blk.60.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 804/ 866] blk.60.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 805/ 866] blk.60.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 806/ 866] blk.60.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 807/ 866] blk.60.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 808/ 866] blk.60.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 809/ 866] blk.60.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 810/ 866] blk.60.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 811/ 866] blk.60.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 812/ 866] blk.60.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 813/ 866] blk.61.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 814/ 866] blk.61.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 815/ 866] blk.61.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 816/ 866] blk.61.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 817/ 866] blk.61.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 818/ 866] blk.61.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 819/ 866] blk.61.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 820/ 866] blk.61.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 821/ 866] blk.61.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 822/ 866] blk.61.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 823/ 866] blk.61.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 824/ 866] blk.61.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 825/ 866] blk.61.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 826/ 866] blk.61.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 827/ 866] blk.62.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 828/ 866] blk.62.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 829/ 866] blk.62.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 830/ 866] blk.62.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 831/ 866] blk.62.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 832/ 866] blk.62.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 833/ 866] blk.62.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 834/ 866] blk.62.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 835/ 866] blk.62.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 836/ 866] blk.62.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to iq1_m .. size = 0.47 MiB -> 0.05 MiB
[ 837/ 866] blk.62.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 838/ 866] blk.62.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 839/ 866] blk.62.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 840/ 866] blk.62.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 60.00 MiB -> 6.56 MiB
[ 841/ 866] blk.63.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to iq1_m .. size = 10.00 MiB -> 1.09 MiB
[ 842/ 866] blk.63.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 843/ 866] blk.63.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 844/ 866] blk.63.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to iq2_xxs .. size = 60.00 MiB -> 7.73 MiB
[ 845/ 866] blk.63.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to iq1_m .. size = 120.00 MiB -> 13.12 MiB
[ 846/ 866] blk.63.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 847/ 866] blk.63.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q4_K .. size = 10.00 MiB -> 2.81 MiB
[ 848/ 866] blk.63.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 849/ 866] blk.63.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 850/ 866] blk.63.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to iq1_m .. size = 170.00 MiB -> 18.59 MiB
[ 851/ 866] blk.63.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 852/ 866] blk.64.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16,
====== llama_model_quantize_impl: did not find weights for blk.64.attn_k.weight
converting to q4_K .. size = 10.00 MiB -> 2.81 MiB
[ 853/ 866] blk.64.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 854/ 866] blk.64.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 855/ 866] blk.64.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16,
====== llama_model_quantize_impl: did not find weights for blk.64.attn_output.weight
converting to q4_K .. size = 60.00 MiB -> 16.88 MiB
[ 856/ 866] blk.64.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16,
====== llama_model_quantize_impl: did not find weights for blk.64.attn_q.weight
converting to q4_K .. size = 120.00 MiB -> 33.75 MiB
[ 857/ 866] blk.64.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 858/ 866] blk.64.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16,
====== llama_model_quantize_impl: did not find weights for blk.64.attn_v.weight
converting to q4_K .. size = 10.00 MiB -> 2.81 MiB
[ 859/ 866] blk.64.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16,
====== llama_model_quantize_impl: did not find weights for blk.64.ffn_down.weight
converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 860/ 866] blk.64.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16,
====== llama_model_quantize_impl: did not find weights for blk.64.ffn_gate.weight
converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 861/ 866] blk.64.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16,
====== llama_model_quantize_impl: did not find weights for blk.64.ffn_up.weight
converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 862/ 866] blk.64.nextn.eh_proj.weight - [ 10240, 5120, 1, 1], type = bf16,
====== llama_model_quantize_impl: did not find weights for blk.64.nextn.eh_proj.weight
converting to q4_K .. size = 100.00 MiB -> 28.12 MiB
[ 863/ 866] blk.64.nextn.enorm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 864/ 866] blk.64.nextn.hnorm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 865/ 866] blk.64.nextn.shared_head_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 866/ 866] blk.64.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
llama_model_quantize_impl: model size = 52115.19 MiB (16.00 BPW)
llama_model_quantize_impl: quant size = 7495.00 MiB (2.30 BPW)
llama_quantize: quantize time = 2766880.41 ms
llama_quantize: total time = 2766880.41 ms