File size: 3,202 Bytes
6819425 9768823 6819425 9768823 6819425 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 | ---
license: mit
base_model: zai-org/GLM-5.2-FP8
base_model_relation: quantized
tags:
- glm
- mixture-of-experts
- int4
- apple-silicon
- ssd-streaming
---
# GLM-5.2 — int4 container for iliria (full 744B, SSD-streamed)
[Explore the model guide](https://huggingface.co/spaces/philipjohnbasile/local-ai-guide) · [All public work](https://huggingface.co/philipjohnbasile)
## Release at a glance
| | This artifact |
|---|---|
| Purpose | The unpruned GLM 5.2 expert set in an int4 container with an int8 MTP head for the iliria engine. |
| Runtime | iliria streams experts from SSD. The documented setup is an Apple Silicon 128 GB machine with roughly 384 GB of weight storage. |
| Status | Recorded runtime; see the evidence and limits below. |
| Tensor download | 383.94 GB (357.57 GiB) of root `.safetensors` files, including any root sidecars. This is a file-size total, not peak RAM. |
| Read first | The on-disk payload is not resident RAM. These files are not interchangeable with the pruned MLX build; consult iliria for runtime limits. |
The **full** GLM-5.2 (744B total / 39B active) quantized to int4 with an int8 MTP head, packaged
as the container that [**iliria**](https://github.com/PhilipJohnBasile/iliria) serves directly.
iliria runs this model on a single 128 GB Apple Silicon machine by **streaming experts from the
SSD** rather than holding all 744B parameters in memory. This repo exists so you don't have to
run that conversion yourself — otherwise the on-ramp is downloading ~372 GB of FP8 weights and
converting them locally for hours.
## This is not the pruned build
If you came from
[GLM-5.2-Demolition-q4a4-soul-MLX](https://huggingface.co/philipjohnbasile/GLM-5.2-Demolition-q4a4-soul-MLX),
note these are different artifacts and **not interchangeable**:
| | this repo | the demolition build |
|---|---|---|
| model | full 744B, unpruned | expert-pruned + LoRA-healed |
| root tensor size | ~384 GB (int4) | ~105 GB / 98 GiB |
| format | iliria container (`out-*.safetensors`) | MLX |
| runs on | iliria | MLX / mlx-lm |
## Requirements
- Apple Silicon **M-Max / M-Ultra class with 128 GB** unified memory
- **~384 GB free SSD** (the weights are streamed from disk, not loaded into RAM)
- [iliria](https://github.com/PhilipJohnBasile/iliria)
## Use
```bash
hf download philipjohnbasile/GLM-5.2-colibri-int4-with-int8-mtp --local-dir ./glm52-int4
# then, from iliria:
./illi serve --model ./glm52-int4 --ram 60
```
That exposes an OpenAI-compatible endpoint. See the iliria README for the engine's flags,
context/RAM tuning, and measured throughput — expect deep-tier speeds, not chat speeds: this is
a 744B model reading experts off an SSD for every token.
## Contents
164 `out-*.safetensors` shards (int4 experts + int8 MTP head) plus `config.json`,
`generation_config.json`, and the tokenizer. Nothing else — no runtime state, no caches.
## Provenance
Quantized from [`zai-org/GLM-5.2-FP8`](https://huggingface.co/zai-org/GLM-5.2-FP8) (MIT).
Quantization only — no pruning, no retraining, no distillation; the full expert set is intact.
Produced by iliria's `tools/convert_fp8_to_int4.py`. Released under MIT, matching the base model.
|