--- license: apache-2.0 base_model: ProCreations/grug-27b tags: - grug - gguf - llama.cpp - reasoning - token-efficient language: - en --- # grug-27b-gguf grug brain squeezed into small rock. run on your cave computer with llama.cpp. this GGUF of [grug-27b](https://huggingface.co/ProCreations/grug-27b): Qwen3.6-27B that think in dense grug-speak inside ``, answer in normal english. same reasoning depth, way fewer think token. full story on main model card. ## rock sizes | file | quant | size | grug opinion | |---|---|---|---| | grug-27b-Q8_0.gguf | Q8_0 | 28.6 GB | basically bf16. big rock. | | grug-27b-Q6_K.gguf | Q6_K | 22.1 GB | very good rock | | grug-27b-Q5_K_M.gguf | Q5_K_M | 19.2 GB | good rock | | grug-27b-Q4_K_M.gguf | Q4_K_M | 16.5 GB | best size/smart trade. grug pick this. | | grug-27b-Q3_K_M.gguf | Q3_K_M | 13.3 GB | small rock. smart mostly survive. | | mmproj-grug-27b-f16.gguf | mmproj f16 | see repo | eye rock. give grug vision back. | every rock load-tested with llama.cpp before upload. no missing-tensor sickness (grug check twice now, learn from 9b). ## Q4 person? special rock exist grug make QAT version of Q4_K_M: weights trained while feeling 4-bit rounding rock before final squish. better Q4 quality, same grug brain: [grug-27b-qat-q4-gguf](https://huggingface.co/ProCreations/grug-27b-qat-q4-gguf). rocks here best for Q8/Q6/Q5 people. ## how run need recent llama.cpp (qwen3_5 arch support). ```bash llama-server -m grug-27b-Q4_K_M.gguf -c 16384 --temp 0.6 --top-p 0.95 --top-k 20 ``` - vision NOW work: pair any quant with `mmproj-grug-27b-f16.gguf` (`llama-server -m grug-27b-Q4_K_M.gguf --mmproj mmproj-grug-27b-f16.gguf`). MTP still not included. - context: base support 262144, pick what your RAM allow - thinking on by default, reasoning arrive inside `...` - for agent frameworks (OpenCode etc): works with think-stripped history, grug trained for exactly that world grug made by ProCreations. base brain by Qwen team.