Instructions to use brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw") model = AutoModelForCausalLM.from_pretrained("brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Trellis
How to use brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw with Trellis:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw
- SGLang
How to use brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw with Docker Model Runner:
docker model run hf.co/brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw
MTP draft sparse-index acceptance collapses past ~index_topk under Glm5vForConditionalGeneration on DCP4 (flat glm_moe_dsa target unaffected)
Reporting a runtime behavior in the sparkinfer v31 image referenced by this repo (verdictai/glm52-exl3-sparkinfer:v31-...), since this repo's v31 text benchmarks are the healthy control arm. Cross-posting here because the upstream fork lives on GitHub; happy to move it if there's a better channel.
sparkinfer v31: MTP draft sparse-index acceptance decorrelates with context length under Glm5vForConditionalGeneration (target unaffected)
One-line
On v31, a GLM-5.2 checkpoint wrapped as Glm5vForConditionalGeneration (MoonViT+PatchMerger graft; LLM is GlmMoeDsaForCausalLM at text_config) has its MTP draft acceptance collapse as context grows past ~`index_topk, while the **target** stays coherent to 64K. The identical LLM weights served flat as glm_moe_dsado not collapse. We isolated it to the DSA sparse-indexer/MTP path for the draft under the VLM wrapper, but the causal line is inside the compiledsparse_attn_indexer` op β we could not reach it from Python, and the obvious config-nesting fix does not resolve it (see Ruled out).
Reproduction (minimal)
Serve the same weights two ways on v31, same runtime/KV/knobs, and read vllm:spec_decode_num_accepted_tokens_per_pos_total (position 0) from /metrics on identical text-only prompts:
| prompt tokens | 2K | 4K | 5K | 5.5K | 8K | 32K | 64K |
|---|---|---|---|---|---|---|---|
flat glm_moe_dsa (control) |
β | 0.86 | β | β | 0.86 | 0.81 | 0.87 |
glm5v wrapper |
0.85 | 1.0 | 0.64 | 0.31 | 0.20 | 0.17 | 0.08 |
- Wrapper collapses; flat control is flat-healthy to 64K. Image input not required (pure text triggers it; images make it worse).
- Boundary tracks
index_topk(=2048). On the older v28 base the boundary was ~2048; on v31 (withVLLM_DCP_SHARD_DRAFT=1/TOPK_OWNER_MERGE/GLOBAL_TOPKeffective) it ~doubled to ~4096 β i.e. v31's DCP work helped but did not close it. Smooth decorrelation, not a sharp knee. - Target output stays coherent at all lengths; only draft acceptance falls (accept-len β ~1.1).
Environment (identical across control and wrapper except the wrapper itself)
Runtime verdictai/glm52-exl3-sparkinfer:v31-gg-v20-sic3828fd-vllm0c79e41-cu132-sm120a @sha256:0433ae94665b769b78dd301f952d907508a3ba80bce47a1630ec20ade8812dff (vLLM 0.11.2.dev280+gilded.gnosis.v20.vllm0c79e41.sic3828fd), 4Γ RTX PRO 6000 (SM120), TP4/DCP4, dcp_comm_backend=ag_rs, kv_cache_dtype=nvfp4_ds_mla, dcp_kv_cache_interleave_size=64, --attention-backend B12X_MLA_SPARSE, --moe-backend b12x, index_topk=2048, index_share_for_mtp_iteration=True, spec {"method":"mtp","num_speculative_tokens":3}, VLLM_DCP_SHARD_DRAFT=1, VLLM_DCP_TOPK_OWNER_MERGE=1, VLLM_DCP_GLOBAL_TOPK=1, VLLM_DCP_INDEXER_SHARDS=0.
Control-arm anchor: the flat glm_moe_dsa control is not a bespoke config β it is this repo's own published v31 preset and knobs. brandonmusic's v31 GLM-5.2-EXL3-TR3-3.0bpw release measures the flat text build on identical 4ΓRTX PRO 6000 / TP4/DCP4 / these DCP knobs at decode C1/C4/C8 = 89.7/225.3/293.5 tok/s, GPU KV 1,132,544 tok @0.96 util, Estonia 5/5 pass. Our control arm reproduces that healthy behavior (per-position acceptance flat 0.81β0.87 to 64K); only the Glm5v wrapper differs between the healthy and collapsing arms.
Ruled out (empirically, this rig)
- Runtime version / KV dtype: collapse present on both v28 and v31; matched
nvfp4_ds_mla(KV pool 844K, β the flat control's 869K). - Image tokens: pure-text prompts collapse on the wrapper.
- Budget size /
index_topk: cannot be raised βindex_topk=8192fails at load:SM120 sparse MLA prefill: unsupported ... topk=8192 ... Supported: NVFP4 (GLM-family) topk in {128,512,1024,2048}. And the target holds to 64K ontopk=2048, so 2048 is sufficient; the draft's failure is wrong-selection, not budget. No user-side budget lever exists. - Config nesting (the obvious fix β and it does NOT work): the DCP/indexer spec builders read
num_hidden_layersandmodel_typefrom the top-levelvllm_config.model_config.hf_config. Under the wrapper these areNoneandglm5v; flat they are78andglm_moe_dsa(verified offline viaAutoConfig). We mirrored both onto the top-level config viahf_overrides(num_hidden_layers=78, then alsomodel_type=glm_moe_dsa) β the model loaded and vision still worked, but the boundary did not move (8K stayed0.23, 32K ~0.15). So although the nesting defect is real and statically verifiable, it is not the causal channel for this collapse (or the compiled op does not consume these top-level values). Reads are atL1940β1961).deepseek_v2.py::_indexer_cache_dcp_shard_count/get_kv_cache_specandmla_attention.py( - Mirroring the index cfg onto the DRAFT's
text_config(incl.index_share_for_mtp_iteration): also does NOT work, on both v28 and v31. We propagated{index_topk, index_topk_freq, index_share_for_mtp_iteration=True, index_topk_pattern, use_index_cache}from the top-levelGlm5vConfigonto the draft config inSpeculativeConfig.hf_config_override(propagation confirmed in logs). Same-session A/B on v31 (nvfp4/MTP3, draft-sharding knobs live), per-position pos0 OFFβON: 8K0.32β0.32, 32K0.29β0.20, 64K0.095β0.10β boundary unchanged. So the draft does receive the sparse-index config; the collapse is not a missing-config problem at either the top level or the draft level.
Not intrinsically unsupported (existence proof)
A different engine (stock vLLM + this same MoonViT/PatchMerger graft, DCP4, sparse MLA, NVFP4 KV) sustains ~53% draft acceptance to 45K. So "sparse indexer + DCP + a VLM wrapper is unsupported, use TP-only" is not the explanation β the combination works elsewhere.
Where it must be
Given the above, the wrapper-specific difference reaches the draft's top-2048 selection through a path not visible from the Python config: the compiled sparse_attn_indexer op, and/or how the Glm5v wrapper feeds hidden-states / the shared MTP index (index_share_for_mtp_iteration) to the draft. The symptom β correct while context β€ ~index_topk, smooth decorrelation past it, target unaffected β is consistent with the draft's shared/selected index being computed over the wrong context extent only under the wrapper.
Ask
A one-line log inside the draft's sparse-indexer of (a) the effective merged-index size / selected context extent and (b) the resolved dcp_world_size/dcp_replicated, for the flat glm_moe_dsa control vs the glm5v wrapper, at 4K and 8K. That single comparison should localize whether the draft is selecting over a truncated context under the wrapper. We have a standing minimal repro and can run instrumented builds.
you may want to send your agent back to this, it says " glm_moe_dsa" is part of this. that is neither in the docker compose yaml or server.sh script. the Glm5vForConditionalGeneration wrapper ( i assume you mean of a checkpoint with a vision head strapped to it) is not necessarily compatible with this. This checkpoint and the compose yaml and server.sh script were not written for that. There are other ones who have done it but it's beyond the scope of this, and i'm not sure your agent is accurate about this model, it contains some assumptions that I do not see as true.
Fair enough, and thanks for the quick reply β you're right on the main point. The Glm5vForConditionalGeneration wrapper (MoonViT + PatchMerger, LLM demoted to text_config) is my own graft on top of your weights, not your release. Your checkpoint is flat glm_moe_dsa and the compose/server.sh serve exactly that, so the vision-nesting behavior is genuinely outside your scope. My fault for framing it against your artifacts.
One small note just for the record: glm_moe_dsa isn't something external I assumed β it's the model_type / architectures (GlmMoeDsaForCausalLM) in your own config.json. I only referenced it as the LLM sitting at text_config inside my wrapper, i.e. the same arch as your flat model β not implying the compose or server.sh name it.
Either way it isn't a weights issue: the collapse is image-independent, wrapper-only, and kicks in exactly at index_topk=2048, which points at the sparse-index share for the MTP draft in the sparkinfer runtime (verdictai/glm52-exl3-sparkinfer:v31) under DCP + Glm5v nesting β not your quant. That's the layer I need to instrument, so I'll take it there. Appreciate you drawing the boundary; it actually helps localize it.