philipjohnbasile's picture
Clarify release scope and connect the model portfolio
9768823 verified
|
Raw
History Blame Contribute Delete
3.2 kB
metadata
license: mit
base_model: zai-org/GLM-5.2-FP8
base_model_relation: quantized
tags:
  - glm
  - mixture-of-experts
  - int4
  - apple-silicon
  - ssd-streaming

GLM-5.2 — int4 container for iliria (full 744B, SSD-streamed)

Explore the model guide · All public work

Release at a glance

This artifact
Purpose The unpruned GLM 5.2 expert set in an int4 container with an int8 MTP head for the iliria engine.
Runtime iliria streams experts from SSD. The documented setup is an Apple Silicon 128 GB machine with roughly 384 GB of weight storage.
Status Recorded runtime; see the evidence and limits below.
Tensor download 383.94 GB (357.57 GiB) of root .safetensors files, including any root sidecars. This is a file-size total, not peak RAM.
Read first The on-disk payload is not resident RAM. These files are not interchangeable with the pruned MLX build; consult iliria for runtime limits.

The full GLM-5.2 (744B total / 39B active) quantized to int4 with an int8 MTP head, packaged as the container that iliria serves directly.

iliria runs this model on a single 128 GB Apple Silicon machine by streaming experts from the SSD rather than holding all 744B parameters in memory. This repo exists so you don't have to run that conversion yourself — otherwise the on-ramp is downloading ~372 GB of FP8 weights and converting them locally for hours.

This is not the pruned build

If you came from GLM-5.2-Demolition-q4a4-soul-MLX, note these are different artifacts and not interchangeable:

this repo the demolition build
model full 744B, unpruned expert-pruned + LoRA-healed
root tensor size ~384 GB (int4) ~105 GB / 98 GiB
format iliria container (out-*.safetensors) MLX
runs on iliria MLX / mlx-lm

Requirements

  • Apple Silicon M-Max / M-Ultra class with 128 GB unified memory
  • ~384 GB free SSD (the weights are streamed from disk, not loaded into RAM)
  • iliria

Use

hf download philipjohnbasile/GLM-5.2-colibri-int4-with-int8-mtp --local-dir ./glm52-int4
# then, from iliria:
./illi serve --model ./glm52-int4 --ram 60

That exposes an OpenAI-compatible endpoint. See the iliria README for the engine's flags, context/RAM tuning, and measured throughput — expect deep-tier speeds, not chat speeds: this is a 744B model reading experts off an SSD for every token.

Contents

164 out-*.safetensors shards (int4 experts + int8 MTP head) plus config.json, generation_config.json, and the tokenizer. Nothing else — no runtime state, no caches.

Provenance

Quantized from zai-org/GLM-5.2-FP8 (MIT). Quantization only — no pruning, no retraining, no distillation; the full expert set is intact. Produced by iliria's tools/convert_fp8_to_int4.py. Released under MIT, matching the base model.