--- license: other library_name: llama.cpp base_model: - poolside/Laguna-XS-2.1 base_model_relation: quantized pipeline_tag: text-generation quantized_by: MagicQuant language: - en tags: - gguf - quantized - magicquant - pruned - reap - moe --- ## What this is A **50% expert-pruned** [poolside/Laguna-XS-2.1](https://huggingface.co/poolside/Laguna-XS-2.1) (33.4B, 256-expert MoE) reduced to **17.7B / 128 experts** with [REAP](https://github.com/CerebrasResearch/reap) router-weighted expert pruning, then quantized with MagicQuant's measured evolutionary search. **63 GB of BF16 weights become a 9.4 GB GGUF.** ### Read this before using it Expert selection was calibrated on **code** (`theblackcat102/evol-codealpaca-v1`). REAP therefore kept the experts that matter for code and dropped ones that did not — the model is **more specialized**, not uniformly degraded. Measured perplexity (100 chunks, ctx 512, identical settings; code = held-out evol-codealpaca): | Model | Size | code PPL | wikitext PPL | |---|---|---|---| | Laguna-XS-2.1 (unpruned, BF16) | 63 GB | 3.1169 | 12.2918 | | REAP-50% pruned (BF16) | 33 GB | 3.4864 (+11.9%) | 34.8363 (+183%) | | **This file — pruned + MagicQuant Q4** | **9.4 GB** | **3.5703 (+14.5%)** | 35.9179 (+192%) | **Use it for code.** General-English ability is substantially reduced — that is the direct, expected consequence of pruning experts by code activations, and it is disclosed here rather than buried. Quantization itself costs only +2.4% on code; almost all of the delta is the pruning. ### Notes - Requires a `laguna`-aware llama.cpp build (arch support landed July 2026; built and validated here against `e9fa078`). - No ROCmFPX sibling repo: the ROCmFPX fork does not yet carry laguna support. - No MTP/speculative-decoding tensors in this architecture. - License follows the base model (openmdw-1.1). All credit for the model to poolside. # Laguna-XS-2.1-REAP50-MagicQuant-GGUF Derivative of [Laguna-XS-2.1](https://huggingface.co/poolside/Laguna-XS-2.1), pruned with [REAP](https://github.com/CerebrasResearch/reap) (Router-weighted Expert Activation Pruning) and quantized using MagicQuant hybrid evolutionary per-tensor search. ## Base Model This is a derivative of [Laguna-XS-2.1](https://huggingface.co/poolside/Laguna-XS-2.1). All credit for the base model architecture and weights goes to the original authors. The base model's license applies to this derivative. ## Expert Pruning (REAP) This is a Mixture-of-Experts model pruned using **[REAP](https://github.com/CerebrasResearch/reap)** (Router-weighted Expert Activation Pruning) from Cerebras Research: - A calibration pass records router decisions and expert activations on representative data - Each expert is scored with a saliency metric weighted by router usage - The lowest-ranked experts in each MoE layer are dropped - The router is trimmed accordingly so the remaining experts cover the full routing distribution The result is a smaller MoE model with fewer experts per layer, trading a small amount of quality for reduced parameter count and inference cost. ## Quantization Method Quantized using **[MagicQuant](https://github.com/lucasmcoleman/MagicQuant)** hybrid evolutionary per-tensor quantization, based on the methodology by **[magiccodingman](https://github.com/magiccodingman/MagicQuant-Wiki)**: - Tensors are classified into sensitivity groups (Embeddings, Head, Query, Key, Output, FFN Up/Down, MoE Experts, Router) - An evolutionary search finds the optimal quantization type per group, balancing size vs. perplexity - **Q4/Q5/Q6 tier targets** are produced with different size-quality tradeoffs - Small-row tensors and sensitivity-critical layers (embeddings, output head, router) are kept at F32/F16/BF16 - This is NOT a uniform quantization -- each tensor group gets its own optimal type ## GGUF Files | File | Size | Quant | |------|------|-------| | [Laguna-XS-2.1-REAP50-Q4_K_M.gguf](./Laguna-XS-2.1-REAP50-Q4_K_M.gguf) | 10.0 GB | Q4 hybrid | ## Usage ### LM Studio 1. Download the GGUF file of your preferred quantization tier 2. Place it in your LM Studio models directory 3. Load the model in LM Studio -- it will auto-detect the chat template 4. The model supports the base model's full context length ### llama.cpp ```bash # Interactive chat (--jinja uses the model's embedded chat template, not a hardcoded one) llama-cli -m Laguna-XS-2.1-REAP50-Q4_K_M.gguf -c 8192 --jinja -cnv # Single prompt llama-cli -m Laguna-XS-2.1-REAP50-Q4_K_M.gguf -c 8192 -p "Your prompt here" # Server mode llama-server -m Laguna-XS-2.1-REAP50-Q4_K_M.gguf -c 8192 --port 8080 --jinja ``` ### Python (llama-cpp-python) ```python from llama_cpp import Llama llm = Llama(model_path="./Laguna-XS-2.1-REAP50-Q4_K_M.gguf", n_ctx=8192) output = llm.create_chat_completion( messages=[ {"role": "user", "content": "Hello, how are you?"} ] ) print(output["choices"][0]["message"]["content"]) ``` ## Caveats - The base model's license (other) applies to all derivative files - Expert pruning removes a fraction of experts per MoE layer; some task-specific knowledge may be lost - The pruned model has a different number of experts from the original — tooling that hardcodes expert count may need adjustment - Quantization reduces precision -- verify outputs for your specific use case - The hybrid quantization assigns different precision to different tensor groups, which means quality characteristics may differ from uniform quantizations ## Limitations - Pruned experts cannot be recovered; any capabilities concentrated in removed experts are lost - Quantized models may exhibit subtle differences from the full-precision fine-tune - This model inherits any limitations and biases present in the base model --- *Generated with [REAP](https://github.com/CerebrasResearch/reap) + [MagicQuant](https://github.com/lucasmcoleman/MagicQuant)*