How to use from
OpenClaw
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "airagrp/Qwen3.8-27B-MLX-nvfp4-mixed"
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest
# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
  --auth-choice custom-api-key \
  --custom-base-url http://127.0.0.1:8080/v1 \
  --custom-model-id "airagrp/Qwen3.8-27B-MLX-nvfp4-mixed" \
  --custom-provider-id mlx-lm \
  --custom-compatibility openai \
  --custom-text-input \
  --accept-risk \
  --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Links

This repository contains Qwen/Qwen3.8-27B converted to MLX format with a mixed-precision quantization recipe, using mlx-vlm 0.6.17.

Quantization recipe

Module Precision
MLP gate_proj / up_proj / down_proj (64 layers) nvfp4 (group_size=16, bits=4)
Full attention q_proj / k_proj / v_proj / o_proj (16 layers) bfloat16
Linear (GDN) attention in_proj_* / out_proj (48 layers) bfloat16
Token embeddings (embed_tokens) bfloat16
Output head (lm_head) bfloat16
MTP head bfloat16
Vision tower bfloat16
  • Effective size: ~31 GB (8.9 bits per weight), base model is ~54 GB in bfloat16.
  • Quantized modules are stored as packed nvfp4 weights (E2M1 codes, 8 per uint32) with per-16 E4M3 block scales; bfloat16 modules are stored as-is. Per-module precision is detected from the presence of .scales tensors; the global mode is set in config.json (quantization.mode = "nvfp4").

MTP

The native MTP head is merged into this checkpoint as language_model.mtp.* tensors (15 tensors, bfloat16, norms in the MLX +1 convention), stored in mtp.safetensors and referenced from model.safetensors.index.json — it is not a separate drafter model. Use it for speculative decoding (--draft-kind mtp in mlx-vlm) or ignore it; base inference is unaffected.

Use with mlx-vlm

pip install mlx-vlm
import mlx_vlm

model, processor = mlx_vlm.load("airagrp/Qwen3.8-27B-MLX-nvfp4-mixed")
response, _ = mlx_vlm.generate(
    model,
    processor,
    prompts="In one sentence, what is MLX?",
    max_tokens=64,
)
print(response)
mlx_vlm.generate --model airagrp/Qwen3.8-27B-MLX-nvfp4-mixed --prompt "In one sentence, what is MLX?" --max-tokens 64

Use with MLX directly

Load with the standard MLX safetensors layout; weights use the MLX nvfp4 block-scale format (group_size=16, bits=4).

Citations / license

Apache-2.0. Refer to the original model card for architecture details, benchmarks, and usage guidelines.

Downloads last month
58
Safetensors
Model size
28B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for airagrp/Qwen3.8-27B-MLX-nvfp4-mixed

Base model

Qwen/Qwen3.8-27B
Quantized
(977)
this model