license: mit
library_name: sglang
pipeline_tag: image-text-to-text
tags:
- multimodal
- vision-language
- glm
- sglang
base_model:
- zai-org/GLM-5.2
GLM-5.2-Vision (NVFP4)
GLM-5.2 with sight. A vision-language model that bolts the MoonViT vision encoder from Kimi-K2.6 onto GLM-5.2 through a trained PatchMerger projector.
GLM-5.2 is a strong open reasoning model with no vision input. This checkpoint adds it, without touching a single GLM weight: the text backbone and the vision tower are both frozen and byte-identical to their upstream releases. The only newly-trained parameters are the 49.5M-parameter projector that maps MoonViT's 1152-dim patch embeddings into GLM's 6144-dim token space.
| Component | Detail |
|---|---|
| Text backbone | GLM-5.2 (744B total / A40B active, MoE + MLA + DSA sparse attention) — frozen |
| Vision tower | MoonViT-3d from Kimi-K2.6, 27 layers, 1152-dim — frozen |
| Projector | PatchMerger MLP (pre_norm → linear_1 → GELU → linear_2), 1152→4608→6144 — trained |
| Text weights | NVFP4, from nvidia/GLM-5.2-NVFP4 |
| Size | ~466 GB |
| Hardware | 8×B200, or 4×B200 at 256k context — Blackwell only |
| Image tokens | up to 4096 per image (16384 MoonViT patches, 2×2 merge) |
| Max context | 1048576 (1M tokens) |
Quickstart
SGLang needs a small out-of-tree plugin, because Glm5vForConditionalGeneration is not yet an
upstream architecture. It ships inside this repo, so there is nothing else to clone:
hf download baseten/GLM-5.2-Vision-NVFP4 --include 'plugins/*' --local-dir ./glm5v
pip install ./glm5v/plugins
SGLang
export SGLANG_EXTERNAL_MODEL_PACKAGE=sglang_glm5v
export SGLANG_EXTERNAL_MM_PROCESSOR_PACKAGE=sglang_glm5v
export SGLANG_EXTERNAL_MM_MODEL_ARCH=Glm5vForConditionalGeneration
python -m sglang_glm5v.patch # registers the arch on SGLang's DSA/MLA paths
8×B200 — full 1M context
python -m sglang.launch_server \
--model-path baseten/GLM-5.2-Vision-NVFP4 --trust-remote-code \
--tp-size 8 \
--quantization modelopt_fp4 \
--disable-shared-experts-fusion --disable-flashinfer-autotune \
--attention-backend dsa --mm-attention-backend sdpa \
--kv-cache-dtype fp8_e4m3 --page-size 64 \
--mem-fraction-static 0.85 \
--context-length 1048576 \
--reasoning-parser glm45 --tool-call-parser glm47 \
--served-model-name glm-5.2-vision \
--port 30000
4×B200 — 256k context
Same command with --tp-size 4, a higher memory fraction, and a smaller context:
python -m sglang.launch_server \
--model-path baseten/GLM-5.2-Vision-NVFP4 --trust-remote-code \
--tp-size 4 \
--quantization modelopt_fp4 \
--disable-shared-experts-fusion --disable-flashinfer-autotune \
--attention-backend dsa --mm-attention-backend sdpa \
--kv-cache-dtype fp8_e4m3 --page-size 64 \
--mem-fraction-static 0.90 \
--context-length 262144 \
--reasoning-parser glm45 --tool-call-parser glm47 \
--served-model-name glm-5.2-vision \
--port 30000
Query it
Standard OpenAI multimodal messages — the image is delivered as image_url:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="none")
r = client.chat.completions.create(
model="glm-5.2-vision",
messages=[{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": "https://ultralytics.com/images/bus.jpg"}},
{"type": "text", "text": "Describe this image in detail."},
]}],
temperature=1.0, top_p=0.95, max_tokens=512,
)
print(r.choices[0].message.content)
GLM-5.2 is a reasoning model: with --reasoning-parser glm45 the chain of thought arrives in
message.reasoning_content and the answer in message.content.
Deploy on Baseten
Ready-to-push Truss configs ship in this repo under truss/:
hf download baseten/GLM-5.2-Vision-NVFP4 --include 'truss/*' --local-dir ./glm5v
cd glm5v/truss
| Config | GPUs | Context |
|---|---|---|
config_nvfp4.yaml |
8×B200 | 1M |
config_nvfp4_4gpu.yaml |
4×B200 | 256k |
truss push --config config_nvfp4.yaml --remote <your-remote> # 8 GPU
truss push --config config_nvfp4_4gpu.yaml --remote <your-remote> # 4 GPU
License
MIT, following both parents: GLM-5.2 (MIT) and Kimi-K2.6 (Modified MIT). The projector weights are released under MIT. Redistributed upstream weights remain under their original terms.
Acknowledgements
Built on Z.ai's GLM-5.2 and Moonshot AI's Kimi-K2.6. Neither team was involved in this work; please do not direct issues with this checkpoint to them.