What this is

A 50% expert-pruned poolside/Laguna-XS-2.1 (33.4B, 256-expert MoE) reduced to 17.7B / 128 experts with REAP router-weighted expert pruning, then quantized with MagicQuant's measured evolutionary search. 63 GB of BF16 weights become a 9.4 GB GGUF.

Read this before using it

Expert selection was calibrated on code (theblackcat102/evol-codealpaca-v1). REAP therefore kept the experts that matter for code and dropped ones that did not β€” the model is more specialized, not uniformly degraded. Measured perplexity (100 chunks, ctx 512, identical settings; code = held-out evol-codealpaca):

Model Size code PPL wikitext PPL
Laguna-XS-2.1 (unpruned, BF16) 63 GB 3.1169 12.2918
REAP-50% pruned (BF16) 33 GB 3.4864 (+11.9%) 34.8363 (+183%)
This file β€” pruned + MagicQuant Q4 9.4 GB 3.5703 (+14.5%) 35.9179 (+192%)

Use it for code. General-English ability is substantially reduced β€” that is the direct, expected consequence of pruning experts by code activations, and it is disclosed here rather than buried. Quantization itself costs only +2.4% on code; almost all of the delta is the pruning.

Notes

  • Requires a laguna-aware llama.cpp build (arch support landed July 2026; built and validated here against e9fa078).
  • No ROCmFPX sibling repo: the ROCmFPX fork does not yet carry laguna support.
  • No MTP/speculative-decoding tensors in this architecture.
  • License follows the base model (openmdw-1.1). All credit for the model to poolside.

Laguna-XS-2.1-REAP50-MagicQuant-GGUF

Derivative of Laguna-XS-2.1, pruned with REAP (Router-weighted Expert Activation Pruning) and quantized using MagicQuant hybrid evolutionary per-tensor search.

Base Model

This is a derivative of Laguna-XS-2.1. All credit for the base model architecture and weights goes to the original authors. The base model's license applies to this derivative.

Expert Pruning (REAP)

This is a Mixture-of-Experts model pruned using REAP (Router-weighted Expert Activation Pruning) from Cerebras Research:

  • A calibration pass records router decisions and expert activations on representative data
  • Each expert is scored with a saliency metric weighted by router usage
  • The lowest-ranked experts in each MoE layer are dropped
  • The router is trimmed accordingly so the remaining experts cover the full routing distribution

The result is a smaller MoE model with fewer experts per layer, trading a small amount of quality for reduced parameter count and inference cost.

Quantization Method

Quantized using MagicQuant hybrid evolutionary per-tensor quantization, based on the methodology by magiccodingman:

  • Tensors are classified into sensitivity groups (Embeddings, Head, Query, Key, Output, FFN Up/Down, MoE Experts, Router)
  • An evolutionary search finds the optimal quantization type per group, balancing size vs. perplexity
  • Q4/Q5/Q6 tier targets are produced with different size-quality tradeoffs
  • Small-row tensors and sensitivity-critical layers (embeddings, output head, router) are kept at F32/F16/BF16
  • This is NOT a uniform quantization -- each tensor group gets its own optimal type

GGUF Files

File Size Quant
Laguna-XS-2.1-REAP50-Q4_K_M.gguf 10.0 GB Q4 hybrid

Usage

LM Studio

  1. Download the GGUF file of your preferred quantization tier
  2. Place it in your LM Studio models directory
  3. Load the model in LM Studio -- it will auto-detect the chat template
  4. The model supports the base model's full context length

llama.cpp

# Interactive chat (--jinja uses the model's embedded chat template, not a hardcoded one)
llama-cli -m Laguna-XS-2.1-REAP50-Q4_K_M.gguf -c 8192 --jinja -cnv

# Single prompt
llama-cli -m Laguna-XS-2.1-REAP50-Q4_K_M.gguf -c 8192 -p "Your prompt here"

# Server mode
llama-server -m Laguna-XS-2.1-REAP50-Q4_K_M.gguf -c 8192 --port 8080 --jinja

Python (llama-cpp-python)

from llama_cpp import Llama

llm = Llama(model_path="./Laguna-XS-2.1-REAP50-Q4_K_M.gguf", n_ctx=8192)
output = llm.create_chat_completion(
    messages=[
        {"role": "user", "content": "Hello, how are you?"}
    ]
)
print(output["choices"][0]["message"]["content"])

Caveats

  • The base model's license (other) applies to all derivative files
  • Expert pruning removes a fraction of experts per MoE layer; some task-specific knowledge may be lost
  • The pruned model has a different number of experts from the original β€” tooling that hardcodes expert count may need adjustment
  • Quantization reduces precision -- verify outputs for your specific use case
  • The hybrid quantization assigns different precision to different tensor groups, which means quality characteristics may differ from uniform quantizations

Limitations

  • Pruned experts cannot be recovered; any capabilities concentrated in removed experts are lost
  • Quantized models may exhibit subtle differences from the full-precision fine-tune
  • This model inherits any limitations and biases present in the base model

Generated with REAP + MagicQuant

Downloads last month
112
GGUF
Model size
18B params
Architecture
laguna
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for lmcoleman/Laguna-XS-2.1-REAP50-MagicQuant-GGUF

Quantized
(31)
this model