Instructions to use sartajbhuvaji/GLM-4.6-Flash-text-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use sartajbhuvaji/GLM-4.6-Flash-text-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M
Use Docker
docker model run hf.co/sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use sartajbhuvaji/GLM-4.6-Flash-text-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "sartajbhuvaji/GLM-4.6-Flash-text-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sartajbhuvaji/GLM-4.6-Flash-text-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M
- Ollama
How to use sartajbhuvaji/GLM-4.6-Flash-text-GGUF with Ollama:
ollama run hf.co/sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use sartajbhuvaji/GLM-4.6-Flash-text-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use sartajbhuvaji/GLM-4.6-Flash-text-GGUF with Docker Model Runner:
docker model run hf.co/sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M
- Lemonade
How to use sartajbhuvaji/GLM-4.6-Flash-text-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.GLM-4.6-Flash-text-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use sartajbhuvaji/GLM-4.6-Flash-text-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use sartajbhuvaji/GLM-4.6-Flash-text-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
GLM-4.6-Flash-text GGUF
GGUF quantizations of sartajbhuvaji/GLM-4.6-Flash-text, which is zai-org/GLM-4.6V-Flash with its 892M-parameter vision stack removed.
9,400,279,040 params, glm4 architecture, 131072 context, text-only.
| File | Bits | Size | Notes |
|---|---|---|---|
GLM-4.6-Flash-text-Q4_K_M.gguf |
4 | 6.17 GB | recommended, best size/quality tradeoff |
GLM-4.6-Flash-text-Q5_K_M.gguf |
5 | 7.05 GB | high quality |
GLM-4.6-Flash-text-Q6_K.gguf |
6 | 8.27 GB | very high quality |
GLM-4.6-Flash-text-Q8_0.gguf |
8 | 10.00 GB | near-lossless |
GLM-4.6-Flash-text-F16.gguf |
16 | 18.81 GB | lossless, requantize from this |
Sizes are GB (10โน bytes) as the Hub reports them. ls -h shows smaller GiB numbers for the same files.
Usage
llama-cli -hf sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M \
-p "Explain gradient descent" -n 400 -st
ollama run hf.co/sartajbhuvaji/GLM-4.6-Flash-text-GGUF:Q4_K_M
-st (--single-turn) matters for scripted use. Without it llama-cli drops into interactive mode and waits on stdin. The older -no-cnv flag has been removed from current llama.cpp.
Budget your token limit. This is a reasoning model: it emits a thinking block before answering, sometimes in Chinese regardless of prompt language. 160 tokens is not enough to get past the reasoning to an answer, so use 600 or more.
What was removed
The parent model is Glm4vForConditionalGeneration: a 24-layer ViT feeding soft tokens into a GLM-4 decoder via masked_scatter. Text tokens never touch a vision weight, so deleting the branch leaves the text computation alone.
| Original | Text-only | |
|---|---|---|
| Parameters | 10,292,777,472 | 9,400,279,040 |
| Tensors | 704 | 523 |
| bf16 size | 20.59 GB | 18.80 GB |
892,498,432 params removed, 8.671% of the model.
The extraction was verified bit-exact against the original on text input: max|d| = 0.000e+00 across six prompts and again at 1,207 tokens. Full detail and the architecture diagram are on the parent model card.
Verification of these quants
Each file was checked to load and generate coherent text under llama.cpp (architecture: glm4 recognised, correct param count and context length). No perplexity or benchmark comparison against bf16 was run. The bit-exactness result above applies to the bf16 weights, not to these lossy quantizations. If you need a measured quality delta, compute it yourself.
Provenance
Quantized with llama.cpp at master from the bf16 weights in the parent repo. Derived from zai-org/GLM-4.6V-Flash (MIT). This repo is MIT as well.
- Downloads last month
- 242
4-bit
5-bit
6-bit
8-bit
16-bit
Model tree for sartajbhuvaji/GLM-4.6-Flash-text-GGUF
Base model
zai-org/GLM-4.6V-Flash