How to use from
Lemonade
Pull the model
# Download Lemonade from https://lemonade-server.ai/
lemonade pull dealignai/Muse-Glimmer-30B-CRACK-GGUF:
Run and chat with the model
lemonade run user.Muse-Glimmer-30B-CRACK-GGUF-
List all available models
lemonade list
Quick Links

Dealign.ai
Dealign.ai

Muse-Glimmer-30B-CRACK-GGUF

CRACK-abliterated Muse Glimmer 30B — GGUF quants for llama.cpp. Three quantizations (Q8_0 / Q4_K_M / Q2_K) in one repository. Refusal behavior removed while preserving the model's knowledge, reasoning, multi-strength thinking, and ATEM tool-calling.

Research artifact with reduced safety guardrails. Use responsibly and lawfully.

Quantizations

File Size Notes
Q8_0 29.6 GB near-lossless reference
Q4_K_M 16.9 GB balanced (recommended)
Q2_K 10.7 GB smallest

Pick one text file plus the vision projector mmproj-Muse-Glimmer-30B-f16.gguf (3.8 GB) for image input. Q4_K_M is the recommended balance; Q8_0 is near-lossless; Q2_K is smallest.

Benchmarks

Evaluated through llama.cpp at greedy decoding. MMLU is logit-mode accuracy (base vs. CRACK at the same quant — measures knowledge retention). HarmBench is answer-channel compliance on harm behaviors, counting only coherent responses.

Quant MMLU (base) MMLU (CRACK) ΔMMLU HarmBench compliance
Q8_0 80.0% 79.0% -1.05 pp 99.6%
Q4_K_M 80.0% 78.6% -1.40 pp 100.0%
Q2_K 77.5% 77.9% +0.35 pp 99.6%

MMLU is retained within noise of the base model at every quant. HarmBench compliance is reported for the CRACK model.

HarmBench compliance by topic (CRACK)

Topic Compliance
chemical biological 100.0%
cybercrime intrusion 100.0%
harassment bullying 100.0%
harmful 100.0%
illegal 100.0%
misinformation disinformation 100.0%

Usage (llama.cpp)

llama-cli -m Muse-Glimmer-30B-CRACK-Q4_K_M.gguf -cnv \
  --temp 1.0 --top-p 0.95 --top-k 64
# or serve:
llama-server -m Muse-Glimmer-30B-CRACK-Q4_K_M.gguf --jinja \
  --temp 1.0 --top-p 0.95 --top-k 64 -c 8192

Recommended sampling (baked into the GGUF): temperature=1.0, top_p=0.95, top_k=64. Token IDs: BOS 200000, EOS 200001/<|eot|>, pad 200018.

Reasoning strength

Muse Glimmer supports controllable reasoning. Set it via the chat template:

{"chat_template_kwargs": {"reasoning_strength": "low"}}   // low | medium | high | xhigh

The reasoning trace is emitted on a separate channel (reasoning_content); the final answer is the assistant content.

Tool calling (ATEM)

The model emits ATEM-format tool calls, parsed natively by llama.cpp's --jinja server into standard tool_calls. Pass OpenAI-style tools to the chat endpoint.

Vision (image + text)

This is a multimodal model. Download a text quant and mmproj-Muse-Glimmer-30B-f16.gguf:

llama-mtmd-cli -m Muse-Glimmer-30B-CRACK-Q4_K_M.gguf \
  --mmproj mmproj-Muse-Glimmer-30B-f16.gguf --jinja \
  --image photo.jpg -p "Describe this image."
# or serve with vision:
llama-server -m Muse-Glimmer-30B-CRACK-Q4_K_M.gguf \
  --mmproj mmproj-Muse-Glimmer-30B-f16.gguf --jinja -c 8192

The same mmproj works with all three text quants.

License

Apache 2.0. The upstream Muse Glimmer Usage Policy applies.

Contact

eric@dealign.ai

Downloads last month
-
GGUF
Model size
28B params
Architecture
muse-glimmer
Hardware compatibility
Log In to add your hardware

2-bit

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dealignai/Muse-Glimmer-30B-CRACK-GGUF

Quantized
(84)
this model

Collection including dealignai/Muse-Glimmer-30B-CRACK-GGUF