How to use from
Hermes Agent
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "sahilchachra/Muse-Glimmer-30B-MXFP8"
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default sahilchachra/Muse-Glimmer-30B-MXFP8
Run Hermes
hermes
Quick Links

Muse-Glimmer-30B — MLX MXFP8

MLX MXFP8 (8-bit microscaling float) quantization of meta-models/Muse-Glimmer-30B, a ~30B dense causal transformer with a ~1.8B perception encoder, built for autonomous agentic tasks on consumer hardware. Runs on Apple Silicon via mlx-vlm. Stays image-text-to-text — the vision tower and projector are kept in bf16.

Precision MXFP8 (E4M3 + E8M0 shared scale, group size 32)
Bits per weight 8.751 bpw
On-disk size 32.6 GB
Quantized language model (incl. lm_head)
Kept in bf16 vision tower + vision adapter/projection
Recommended RAM 32 GB+ unified memory

This is the higher-fidelity build, for 32 GB+ Macs. On a 24 GB machine it exceeds RAM and pages to swap (usable only very slowly); use the MXFP4 build (18.6 GB) there instead.

Verification

Quantized with mlx_lm.quantize_model (mode mxfp8, group 32), keeping the vision path in bf16. The MLX implementation correctly handles this architecture's non-standard pieces (per-layer NoPE on the full-attention layers, final_logit_softcapping, qk_scale_factor, output_multiplier, gated attention, centered RMSNorm).

MXFP8 was validated against the MXFP4 build, which passed 6/6 arithmetic prompts end-to-end with correct answers and coherent reasoning. On a fixed 8-prompt set (arithmetic + open-ended), MXFP8's next-token predictions were captured and compared to MXFP4:

Metric MXFP8 vs MXFP4
top-1 next-token agreement 8/8
logit cosine similarity 0.998 (min 0.997)

Since MXFP8 uses more bits than the behaviorally-verified MXFP4 and agrees with it this closely, it is at least as faithful to the base model. (Full token-by- token generation was not benchmarked here because 32.6 GB exceeds the 24 GB test machine's RAM; on a 32 GB+ Mac it generates at normal speed.)

Usage (mlx-vlm)

pip install -U mlx-vlm   # needs >= 0.6.12 for the muse_glimmer architecture
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template

model, processor = load("sahilchachra/Muse-Glimmer-30B-MXFP8")
config = model.config

messages = [{"role": "user", "content": "What is 84 * 3 / 2?"}]
prompt = apply_chat_template(processor, config, messages, add_generation_prompt=True)
text = generate(model, processor, prompt, max_tokens=256, verbose=True)

For image input, pass an image to apply_chat_template / generate per the mlx-vlm docs — the vision path is preserved in bf16.

Recommended sampling (from the base model card): temperature=1.0, top_p=0.95, top_k=64. Reasoning strength is set via the system prompt (Reasoning strength: low|medium|high|xhigh).

Notes & limitations

Downloads last month
417
Safetensors
Model size
30B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sahilchachra/Muse-Glimmer-30B-MXFP8

Quantized
(161)
this model