Instructions to use felipedpm/z-image-turbo-GGUF-confyui with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use felipedpm/z-image-turbo-GGUF-confyui with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf felipedpm/z-image-turbo-GGUF-confyui:UD-Q5_K_XL # Run inference directly in the terminal: llama cli -hf felipedpm/z-image-turbo-GGUF-confyui:UD-Q5_K_XL
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf felipedpm/z-image-turbo-GGUF-confyui:UD-Q5_K_XL # Run inference directly in the terminal: llama cli -hf felipedpm/z-image-turbo-GGUF-confyui:UD-Q5_K_XL
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf felipedpm/z-image-turbo-GGUF-confyui:UD-Q5_K_XL # Run inference directly in the terminal: ./llama-cli -hf felipedpm/z-image-turbo-GGUF-confyui:UD-Q5_K_XL
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf felipedpm/z-image-turbo-GGUF-confyui:UD-Q5_K_XL # Run inference directly in the terminal: ./build/bin/llama-cli -hf felipedpm/z-image-turbo-GGUF-confyui:UD-Q5_K_XL
Use Docker
docker model run hf.co/felipedpm/z-image-turbo-GGUF-confyui:UD-Q5_K_XL
- LM Studio
- Jan
- Ollama
How to use felipedpm/z-image-turbo-GGUF-confyui with Ollama:
ollama run hf.co/felipedpm/z-image-turbo-GGUF-confyui:UD-Q5_K_XL
- Unsloth Studio
How to use felipedpm/z-image-turbo-GGUF-confyui with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for felipedpm/z-image-turbo-GGUF-confyui to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for felipedpm/z-image-turbo-GGUF-confyui to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for felipedpm/z-image-turbo-GGUF-confyui to start chatting
- Pi
How to use felipedpm/z-image-turbo-GGUF-confyui with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf felipedpm/z-image-turbo-GGUF-confyui:UD-Q5_K_XL
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "felipedpm/z-image-turbo-GGUF-confyui:UD-Q5_K_XL" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use felipedpm/z-image-turbo-GGUF-confyui with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf felipedpm/z-image-turbo-GGUF-confyui:UD-Q5_K_XL
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "felipedpm/z-image-turbo-GGUF-confyui:UD-Q5_K_XL" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use felipedpm/z-image-turbo-GGUF-confyui with Docker Model Runner:
docker model run hf.co/felipedpm/z-image-turbo-GGUF-confyui:UD-Q5_K_XL
- Lemonade
How to use felipedpm/z-image-turbo-GGUF-confyui with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull felipedpm/z-image-turbo-GGUF-confyui:UD-Q5_K_XL
Run and chat with the model
lemonade run user.z-image-turbo-GGUF-confyui-UD-Q5_K_XL
List all available models
lemonade list
- Hermes Agent
How to use felipedpm/z-image-turbo-GGUF-confyui with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf felipedpm/z-image-turbo-GGUF-confyui:UD-Q5_K_XL
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default felipedpm/z-image-turbo-GGUF-confyui:UD-Q5_K_XL
Run Hermes
hermes
- Atomic Chat
z-image-turbo (GGUF Version) for ComfyUI
This repository contains the GGUF quantized weights for the z-image-turbo model, optimized to run in environments with limited VRAM resources (though still demanding) using ComfyUI.
The goal of this upload is to enable the execution of this pipeline by leveraging the efficiency of the GGUF format for both the UNET and the Text Encoder (Qwen). Feel free to download and use just the workflow (json) in the models tab and versions!
π Files and Structure
The models are organized within the repository folders as follows:
UNET:
models/unet/z_image_turbo-Q8_0.gguf- Q8 quantized version of the main diffusion model.
Text Encoder:
models/text_encoders/Qwen3-4B-UD-Q5_K_XL.gguf- Qwen3 4B LLM quantized in Q5, used for prompt processing.
VAE:
models/vae/ae.safetensors- Standard Variational Autoencoder for decoding the image.
Feel free to download and use just the workflow!
βοΈ Installation in ComfyUI
To use these models, you will need custom Nodes that support GGUF loading (such as City96's ComfyUI-GGUF or similar).
Download the files:
- Move the
.gguffile from theunetfolder to:ComfyUI/models/unet/ - Move the
.gguffile from thetext_encodersfolder to:ComfyUI/models/clip/(ortext_encodersdepending on your node loader). - Move the
.safetensorsfile from thevaefolder to:ComfyUI/models/vae/
- Move the
Recommended Nodes:
- Use UnetLoaderGGUF to load
z_image_turbo-Q8_0.gguf. - Use a GGUF-compatible CLIP/Text Encoder Loader to load
Qwen3-4B.
- Use UnetLoaderGGUF to load
π» Hardware Requirements and Performance
β οΈ Warning: This is a heavy workflow.
Even with GGUF quantization, the model requires considerable hardware due to the size of the text encoder and the UNET.
- Minimum GPU: 12GB VRAM (NVIDIA RTX 3060/4070 or higher).
- System RAM: 32GB Recommended (System may need to offload data to RAM).
Generation Time
On a GPU with 12GB VRAM, the estimated generation time per image ranges between:
- 15 to 30 seconds. making considerably fast!
Model Information Check out the original model card Z-Image Turbo for detailed information about the model.
π Useful Links
- Downloads last month
- 2,076
5-bit
8-bit



