--- base_model: openbmb/MiniCPM5-1B-Base license: apache-2.0 library_name: gguf pipeline_tag: text-generation tags: - gguf - quantized - rotorquant - llama.cpp --- > [!TIP] > **KV-cache quantization (upstream, no fork needed):** llama.cpp/Ollama cover > this natively — `-ctk q8_0 -ctv q8_0` (~half KV memory, negligible quality > loss) or `-ctk q4_0 -ctv q4_0` (~quarter memory, small quality cost). In > Ollama: `OLLAMA_KV_CACHE_TYPE=q8_0` with `OLLAMA_FLASH_ATTENTION=1`. # MiniCPM5-1B-Base — RotorQuant GGUF Q5_K_M [`openbmb/MiniCPM5-1B-Base`](https://huggingface.co/openbmb/MiniCPM5-1B-Base) quantized pack, published as `MiniCPM5-1B-Base-RotorQuant-GGUF-Q5_K_M`. ## Method llama.cpp Q5_K_M quantization. ## Release line Released under the **RotorQuant** line. RotorQuant and TurboQuant are this project's release labels for this pack, not distinct quantization algorithms — both brand repos carry byte-identical weights. No brand-specific speedup is claimed or measured. ## Modality `pipeline_tag: text-generation`. This is a llama.cpp GGUF conversion of the text tower; no modality beyond `pipeline_tag` above is claimed or included. ## License This pack is a **derivative** of [`openbmb/MiniCPM5-1B-Base`](https://huggingface.co/openbmb/MiniCPM5-1B-Base); all credit for the original model, training, and weights belongs to the upstream authors. This repo republishes a quantized conversion of those weights only. Governed by the [apache-2.0](https://www.apache.org/licenses/LICENSE-2.0). See the upstream repo and the linked license for the full terms — no license text is reproduced here.