# Local model package This directory is intentionally excluded from the source repository. ## Layout - `hf/Confucius4-T3PO/`: copied Hugging Face checkpoint. - `gguf/Confucius4-T3PO-F16.gguf`: F16 GGUF converted from that checkpoint. - `gguf/Confucius4-T3PO-Q6_K.gguf`: Q6_K quantization of the F16 GGUF. - `gguf/Confucius4-T3PO-Q5_K_M.gguf`: Q5_K_M quantization of the F16 GGUF. ## Conversion - Converter: `ggml-org/llama.cpp` - Converter commit: `5202104b59ada9005db079eea43882a2b7bf5802` - Architecture detected by the converter: `Qwen2ForCausalLM` - GGUF architecture: `qwen2` - GGUF version: `3` - Tensor count: `579` | File | Output type | Size (bytes) | | --- | :---: | ---: | | `Confucius4-T3PO-F16.gguf` | `F16` | 29,547,714,080 | | `Confucius4-T3PO-Q6_K.gguf` | `Q6_K` | 12,124,681,760 | | `Confucius4-T3PO-Q5_K_M.gguf` | `Q5_K_M` | 10,508,871,200 | Equivalent local conversion command: ```bash python3 convert_hf_to_gguf.py \ model/hf/Confucius4-T3PO \ --outfile model/gguf/Confucius4-T3PO-F16.gguf \ --outtype f16 \ --model-name Confucius4-T3PO ``` The low-bit variants are quantized from the F16 GGUF: ```bash llama-quantize model/gguf/Confucius4-T3PO-F16.gguf \ model/gguf/Confucius4-T3PO-Q6_K.gguf Q6_K llama-quantize model/gguf/Confucius4-T3PO-F16.gguf \ model/gguf/Confucius4-T3PO-Q5_K_M.gguf Q5_K_M ``` No GGUF splitting was applied.