How to use from
OpenClaw
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "pipenetwork/Inkling-Small-MLX-3bit"
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest
# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
  --auth-choice custom-api-key \
  --custom-base-url http://127.0.0.1:8080/v1 \
  --custom-model-id "pipenetwork/Inkling-Small-MLX-3bit" \
  --custom-provider-id mlx-lm \
  --custom-compatibility openai \
  --custom-text-input \
  --accept-risk \
  --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Links

Inkling-Small-MLX-3bit

Built with Inkling (Thinking Machines Lab).

MLX (Apple Silicon) conversion of thinkingmachines/Inkling-Small, quantized to 3-bit (affine group quant, group size 64).

Code / loader: github.com/PipeNetwork/inkling-mlx

⚠️ This build is experimental

3-bit measures +20% text perplexity vs 8-bit (6.706 vs 5.569). It answers direct factual and coding questions correctly, but after the answer it tends to fall into repetition loops and emit stray glyphs. At ~116 GB it also only just fits a 128 GB Mac, needing iogpu.wired_limit_mb raised close to the ceiling.

For the same ~112 GB footprint, take REAP25-4bit instead. It keeps 4-bit precision and drops 25% of the routed experts instead, which measured as no perplexity cost (vs this build's +20%), with vision and speech intact. This 3-bit build is kept only for the case where you want the full 256-expert set at that size. With more memory, plain 4-bit shows no measurable loss vs 8-bit.

Inkling Small is a 276B-total / 12B-active sparse-MoE, natively multimodal model (text + image/video + audio → text). This is the full multimodal conversion: all three towers (text backbone, HMLP vision, dMel audio) are ported; the multi-token-prediction head is dropped (inference-irrelevant).

Builds

Variant Size Text ppl Notes
8bit ~280 GB 5.569 near-lossless
6bit ~214 GB 5.569 high quality
4bit ~148 GB 5.452 balanced default
3bit ~115 GB 6.706 ⚠️ experimental — visibly degraded

Perplexity is teacher-forcing over one fixed held-out set (prose / code / reasoning / multilingual) — identical inputs across builds, so the columns compare directly. 4-bit shows no measurable loss vs 8-bit.

There is also a REAP-pruned build: REAP25-4bit keeps 4-bit precision with 192 of 256 routed experts, fitting a 128 GB Mac at ~112 GB for no measurable perplexity cost, with vision and speech intact.

No bf16 build is published. The MLX bf16 conversion is bit-identical to the upstream checkpoint (name-mapping and layout only — the dtype cast is a no-op), so it would carry nothing thinkingmachines/Inkling-Small does not already have, and at ~527 GB it does not fit a 512 GB Mac. If you want it as a requant source, scripts/convert_all.sh regenerates it from the upstream weights in about three minutes.

Quantization scheme: affine int4 (not NVFP4 / MXFP4)

MLX supports FP4 modes and Thinking Machines ships an Inkling-NVFP4 checkpoint — so for the record, we benchmarked round-trip reconstruction error (‖W − Ŵ‖ / ‖W‖ vs bf16) on real Inkling expert weights:

Scheme bits/weight reconstruction error
affine int4 (group 64) 4.50 ~9.1%
nvfp4 (group 16) 4.50 ~10.2%
mxfp4 (group 32) 4.25 ~12.3%

Affine int4 is the most faithful: it is asymmetric (per-group scale and zero-point, 16 uniform levels), which centers on Inkling's near-Gaussian expert weights better than symmetric FP4's fixed non-uniform levels. FP4's real payoff is heavy-tailed activations and native Blackwell FP4 tensor cores — neither helps weight fidelity on Apple Silicon, where MLX would dequantize FP4 anyway. So these builds use affine int4.

⚠️ Loading requires the bundled inkling_mlx loader

The inkling_mm_model architecture is not in stock mlx-lm / mlx-vlm, so this repo bundles a minimal, numerically-validated MLX implementation under inkling_mlx/.

pip install mlx mlx-lm transformers
from inkling_mlx.load import load
from inkling_mlx.generate import greedy_generate
from transformers import AutoTokenizer

model, config = load("/path/to/this/repo")
tok = AutoTokenizer.from_pretrained("/path/to/this/repo", trust_remote_code=True)
ids = tok("The capital of France is")["input_ids"]
print(tok.decode(greedy_generate(model, config, ids, max_new_tokens=64)))

Needs an Apple-Silicon Mac with enough unified memory to hold the weights (≈ the size above).

Status & caveats

  • Text generation works end-to-end via an incremental KV + short-convolution cache.
  • Multimodal is supported end-to-end: the vision/audio towers and their preprocessing (InklingProcessor — image patchify/normalize, audio log-mel→dMel, validated ~1e-7 vs the reference) are included. Pass images/audio via the processor.
  • Quantized: attention / MLP / expert projections, token embed+unembed, and the vision/audio matmuls. Kept in higher precision: the MoE router, RMSNorms, the four short-convolutions per layer, and the relative-position bias.

Conversion is streaming (tensor-by-tensor; the ~527 GB bf16 model never fully loads into RAM) and was validated with fp32 numerical parity against transformers PR #47347. License: Apache-2.0 (inherits the base model).

Downloads last month
43
Safetensors
Model size
33B params
Tensor type
BF16
·
U32
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pipenetwork/Inkling-Small-MLX-3bit

Quantized
(25)
this model