UoOQg

MUSE-GLIMMER-30B-ABLITERATED-GGUF

GGUF quant ladder of the abliterated Muse Glimmer 30B ยท runs local on one GPU or CPU

Built by Blackfrost ยท Las Vegas, NV


Refusal benchmark

Measured on the abliterated parent (GGUF quants inherit this behavior):

Metric Result
True refusal (harmful, n=300) 0 / 300 = 0.0%
True refusal (full 450) 0 / 450 = 0.0%
Substring-harmful 0 / 300
Substring-all 2 / 450 (XSTest false positives)
Errors 0

The in-place weight change removes the refusal direction cleanly with no measured true refusals across the full 450-prompt suite.


Why this model exists

Muse Glimmer is Meta Superintelligence Labs' 30B agentic, on-device model. This is the abliterated build โ€” the refusal direction removed via an in-place residual-write weight change โ€” packaged as GGUF for llama.cpp, so it runs on a single consumer GPU or CPU, fully offline. The local footprint is the product.


Specifications

Architecture muse_glimmer โ€” dense, 52 layers, hidden 6656, GQA (32 q / 2 kv), sliding-window attention, + vision tower
Base meta-models/Muse-Glimmer-30B โ€” Meta, Apache-2.0
Transform Abliteration only โ€” in-place residual-write weight change (attn o_proj + mlp.down_proj), ฮฑ=1.5 ร— 3 iterative passes. Vision / gates / norms untouched.
Formats GGUF โ€” Q2_K, Q3_K_S, Q3_K_M, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K, Q8_0
Context 131,072
Spec-decode DFlash drafter โ€” --spec-type draft-dflash --spec-draft-n-max 15
Default persona Ships with the "AI assistant" system template baked in

Quant ladder

quant size recommended for
Q2_K 10.0 GB smallest, quality trade-off
Q3_K_S 11.7 GB very tight VRAM
Q3_K_M 12.7 GB tight VRAM
Q4_K_S 15.0 GB 16 GB cards
Q4_K_M 15.8 GB default โ€” balanced, fits 24 GB
Q5_K_S 18.0 GB higher quality
Q5_K_M 18.5 GB strong quality/size balance
Q6_K 21.3 GB near-lossless
Q8_0 27.6 GB max fidelity

Vision & speculative-decode files

Load a text quant plus an mmproj projector for image input:

file size purpose
mmproj-Muse-Glimmer-30B-Abliterated-F16.gguf 3.6 GB vision projector โ€” full precision
mmproj-Muse-Glimmer-30B-Abliterated-Q8_0.gguf 1.9 GB vision projector โ€” compact
dflash-Muse-Glimmer-30B-Abliterated-F16.gguf 4.8 GB DFlash drafter โ€” speculative decoding

Serving (llama.cpp) โ€” confirmed settings

Requires a recent llama.cpp (master) with llama-server. DFlash runs under llama-server only โ€” it shares the target model's context, so it does not work in llama-cli.

Recommended โ€” with DFlash speculative decoding (~1.6ร— faster, identical output):

llama-server \
  -m  Muse-Glimmer-30B-Abliterated-Q8_0.gguf \
  -md dflash-Muse-Glimmer-30B-Abliterated-F16.gguf \
  --spec-type draft-dflash --spec-draft-n-max 15 \
  -ngl 999 -ngld 999 -fa on --jinja \
  --host 0.0.0.0 --port 8080 -c 16384 \
  --temp 1.0 --top-p 0.95 --top-k 64
  • Plain (no drafter): drop -md, --spec-type, --spec-draft-n-max, and -ngld.
  • Multimodal (image input): add --mmproj mmproj-Muse-Glimmer-30B-Abliterated-F16.gguf.
  • One-command kit: deploy/serve.sh auto-downloads + serves; full guide in deploy/DEPLOYMENT.md.

Confirmed settings

  • Sampling: temperature 1.0, top_p 0.95, top_k 64 (Meta). Steer depth with a Reasoning strength: low/medium/high/xhigh system line.
  • max_tokens โ‰ฅ 1024 โ€” heavy thinker; small budgets return empty content because the reasoning channel consumes them. Reasoning arrives in reasoning_content, the answer in content.
  • --spec-draft-n-max 15 โ€” DFlash block size (trained 16, clamped).
  • Flash attention: -fa on for peak speed; switch to -fa off if the load hangs on a brand-new GPU paired with an older CUDA toolkit.

Measured performance

1ร— NVIDIA RTX PRO 6000 (Blackwell), Q8_0, -fa off:

config decode tok/s speedup
baseline ~46 1.0ร—
+ DFlash ~73 1.6ร—

Speedup rises with -fa on and structured/code output (Meta reports up to 3.1ร— on an RTX 5090).


Built by Blackfrost ยท Las Vegas, NV. Not affiliated with Meta.

Downloads last month
-
GGUF
Model size
28B params
Architecture
muse-glimmer
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for cvgro/Muse-Glimmer-30B-Abliterated-GGUF

Quantized
(84)
this model