Text Generation
Transformers
Safetensors
Trellis
English
Chinese
glm_moe_dsa
glm
exl3
vllm
blackwell
mixture-of-experts
conversational
modelopt
Instructions to use brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw") model = AutoModelForCausalLM.from_pretrained("brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Trellis
How to use brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw with Trellis:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw
- SGLang
How to use brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw with Docker Model Runner:
docker model run hf.co/brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw
Make the packaged server local-only by default
Browse filesBind the published API to loopback and clarify that localhost refers to each downloader own machine.
- README.md +7 -6
- docker-compose.yml +1 -1
- server.sh +2 -1
README.md
CHANGED
|
@@ -82,9 +82,10 @@ chmod +x server.sh
|
|
| 82 |
./server.sh logs
|
| 83 |
```
|
| 84 |
|
| 85 |
-
|
| 86 |
-
`http://localhost:8000/v1`.
|
| 87 |
-
|
|
|
|
| 88 |
|
| 89 |
```bash
|
| 90 |
curl http://localhost:8000/v1/chat/completions \
|
|
@@ -126,6 +127,7 @@ The defaults can be overridden without editing the files:
|
|
| 126 |
| `MODEL_DIR` | Directory containing `server.sh` | Model mount |
|
| 127 |
| `CACHE_DIR` | `~/.cache/glm52-exl3-sparkinfer` | Persistent JIT cache |
|
| 128 |
| `PORT` | `8000` | Host API port |
|
|
|
|
| 129 |
| `CUDA_VISIBLE_DEVICES` | `3,1,2,0` | Physical GPU to TP-rank order |
|
| 130 |
| `GPU_MEMORY_UTILIZATION` | `0.93` | vLLM memory reservation |
|
| 131 |
| `MAX_MODEL_LEN` | `524288` | Per-request context cap |
|
|
@@ -134,9 +136,8 @@ The tested GPU order intentionally keeps physical GPU 3 away from TP rank 3.
|
|
| 134 |
On another host, set `CUDA_VISIBLE_DEVICES=0,1,2,3` or use the order appropriate
|
| 135 |
for that machine.
|
| 136 |
|
| 137 |
-
The Compose file
|
| 138 |
-
|
| 139 |
-
before exposing the API to an untrusted network.
|
| 140 |
|
| 141 |
## Runtime validation
|
| 142 |
|
|
|
|
| 82 |
./server.sh logs
|
| 83 |
```
|
| 84 |
|
| 85 |
+
The OpenAI-compatible endpoint is available locally at
|
| 86 |
+
`http://localhost:8000/v1`. Here, `localhost` always means the machine on
|
| 87 |
+
which the downloader starts this model; it does not refer to the model
|
| 88 |
+
publisher's machine.
|
| 89 |
|
| 90 |
```bash
|
| 91 |
curl http://localhost:8000/v1/chat/completions \
|
|
|
|
| 127 |
| `MODEL_DIR` | Directory containing `server.sh` | Model mount |
|
| 128 |
| `CACHE_DIR` | `~/.cache/glm52-exl3-sparkinfer` | Persistent JIT cache |
|
| 129 |
| `PORT` | `8000` | Host API port |
|
| 130 |
+
| `BIND_ADDRESS` | `127.0.0.1` | Local-only host binding |
|
| 131 |
| `CUDA_VISIBLE_DEVICES` | `3,1,2,0` | Physical GPU to TP-rank order |
|
| 132 |
| `GPU_MEMORY_UTILIZATION` | `0.93` | vLLM memory reservation |
|
| 133 |
| `MAX_MODEL_LEN` | `524288` | Per-request context cap |
|
|
|
|
| 136 |
On another host, set `CUDA_VISIBLE_DEVICES=0,1,2,3` or use the order appropriate
|
| 137 |
for that machine.
|
| 138 |
|
| 139 |
+
The supplied Compose file binds only to loopback by default, so it does not
|
| 140 |
+
publish the API to the LAN or internet.
|
|
|
|
| 141 |
|
| 142 |
## Runtime validation
|
| 143 |
|
docker-compose.yml
CHANGED
|
@@ -3,7 +3,7 @@ services:
|
|
| 3 |
image: ${IMAGE:-verdictai/glm52-exl3-sparkinfer:v1-gg-60c82d972-spi1937274-cu132-sm120a}
|
| 4 |
container_name: glm52-exl3-sparkinfer
|
| 5 |
ports:
|
| 6 |
-
- "
|
| 7 |
gpus: all
|
| 8 |
shm_size: "32g"
|
| 9 |
ipc: host
|
|
|
|
| 3 |
image: ${IMAGE:-verdictai/glm52-exl3-sparkinfer:v1-gg-60c82d972-spi1937274-cu132-sm120a}
|
| 4 |
container_name: glm52-exl3-sparkinfer
|
| 5 |
ports:
|
| 6 |
+
- "${BIND_ADDRESS:-127.0.0.1}:${PORT:-8000}:8000"
|
| 7 |
gpus: all
|
| 8 |
shm_size: "32g"
|
| 9 |
ipc: host
|
server.sh
CHANGED
|
@@ -7,6 +7,7 @@ export IMAGE="${IMAGE:-verdictai/glm52-exl3-sparkinfer:v1-gg-60c82d972-spi193727
|
|
| 7 |
export MODEL_DIR="${MODEL_DIR:-$SCRIPT_DIR}"
|
| 8 |
export CACHE_DIR="${CACHE_DIR:-$HOME/.cache/glm52-exl3-sparkinfer}"
|
| 9 |
export PORT="${PORT:-8000}"
|
|
|
|
| 10 |
export CUDA_VISIBLE_DEVICES="${CUDA_VISIBLE_DEVICES:-3,1,2,0}"
|
| 11 |
export GPU_MEMORY_UTILIZATION="${GPU_MEMORY_UTILIZATION:-0.93}"
|
| 12 |
export MAX_MODEL_LEN="${MAX_MODEL_LEN:-524288}"
|
|
@@ -20,7 +21,7 @@ usage() {
|
|
| 20 |
Usage: ./server.sh [start|stop|restart|logs|status|pull]
|
| 21 |
|
| 22 |
Environment overrides:
|
| 23 |
-
IMAGE, MODEL_DIR, CACHE_DIR, PORT, CUDA_VISIBLE_DEVICES,
|
| 24 |
GPU_MEMORY_UTILIZATION, MAX_MODEL_LEN, COMPOSE_PROJECT_NAME, COMPOSE_FILE
|
| 25 |
EOF
|
| 26 |
}
|
|
|
|
| 7 |
export MODEL_DIR="${MODEL_DIR:-$SCRIPT_DIR}"
|
| 8 |
export CACHE_DIR="${CACHE_DIR:-$HOME/.cache/glm52-exl3-sparkinfer}"
|
| 9 |
export PORT="${PORT:-8000}"
|
| 10 |
+
export BIND_ADDRESS="${BIND_ADDRESS:-127.0.0.1}"
|
| 11 |
export CUDA_VISIBLE_DEVICES="${CUDA_VISIBLE_DEVICES:-3,1,2,0}"
|
| 12 |
export GPU_MEMORY_UTILIZATION="${GPU_MEMORY_UTILIZATION:-0.93}"
|
| 13 |
export MAX_MODEL_LEN="${MAX_MODEL_LEN:-524288}"
|
|
|
|
| 21 |
Usage: ./server.sh [start|stop|restart|logs|status|pull]
|
| 22 |
|
| 23 |
Environment overrides:
|
| 24 |
+
IMAGE, MODEL_DIR, CACHE_DIR, PORT, BIND_ADDRESS, CUDA_VISIBLE_DEVICES,
|
| 25 |
GPU_MEMORY_UTILIZATION, MAX_MODEL_LEN, COMPOSE_PROJECT_NAME, COMPOSE_FILE
|
| 26 |
EOF
|
| 27 |
}
|