Instructions to use benxh/Qwen2.5-VL-7B-Instruct-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use benxh/Qwen2.5-VL-7B-Instruct-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf benxh/Qwen2.5-VL-7B-Instruct-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf benxh/Qwen2.5-VL-7B-Instruct-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf benxh/Qwen2.5-VL-7B-Instruct-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf benxh/Qwen2.5-VL-7B-Instruct-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf benxh/Qwen2.5-VL-7B-Instruct-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf benxh/Qwen2.5-VL-7B-Instruct-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf benxh/Qwen2.5-VL-7B-Instruct-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf benxh/Qwen2.5-VL-7B-Instruct-GGUF:Q4_K_M
Use Docker
docker model run hf.co/benxh/Qwen2.5-VL-7B-Instruct-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use benxh/Qwen2.5-VL-7B-Instruct-GGUF with Ollama:
ollama run hf.co/benxh/Qwen2.5-VL-7B-Instruct-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use benxh/Qwen2.5-VL-7B-Instruct-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf benxh/Qwen2.5-VL-7B-Instruct-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "benxh/Qwen2.5-VL-7B-Instruct-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use benxh/Qwen2.5-VL-7B-Instruct-GGUF with Docker Model Runner:
docker model run hf.co/benxh/Qwen2.5-VL-7B-Instruct-GGUF:Q4_K_M
- Lemonade
How to use benxh/Qwen2.5-VL-7B-Instruct-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull benxh/Qwen2.5-VL-7B-Instruct-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Qwen2.5-VL-7B-Instruct-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use benxh/Qwen2.5-VL-7B-Instruct-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf benxh/Qwen2.5-VL-7B-Instruct-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default benxh/Qwen2.5-VL-7B-Instruct-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use benxh/Qwen2.5-VL-7B-Instruct-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf benxh/Qwen2.5-VL-7B-Instruct-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "benxh/Qwen2.5-VL-7B-Instruct-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Wrong format?
LM-Studio did not recognise this model as VL, unlike previous models.
Any update on this?
FYI, both llama-llava-cli and llama-cli make it produce gibberish:
llama-cli -c 0 --temp 0.2 -m Qwen2.5-VL-7B-Instruct-Q4_K_M.gguf -p "Provide a full description."
...
<|im_start|>system
You are a helpful assistant<|im_end|>
<|im_start|>user
Hello<|im_end|>
<|im_start|>assistant
Hi there<|im_end|>
<|im_start|>user
How are you?<|im_end|>
<|im_start|>assistant
system_info: n_threads = 6 (n_threads_batch = 6) / 12 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | LLAMAFILE = 1 | OPENMP = 1 | AARCH64_REPACK = 1 |
main: interactive mode on.
sampler seed: 1971700995
sampler params:
repeat_last_n = 64, repeat_penalty = 1,000, frequency_penalty = 0,000, presence_penalty = 0,000
dry_multiplier = 0,000, dry_base = 1,750, dry_allowed_length = 2, dry_penalty_last_n = 128000
top_k = 40, top_p = 0,950, min_p = 0,050, xtc_probability = 0,000, xtc_threshold = 0,100, typical_p = 1,000, temp = 0,200
mirostat = 0, mirostat_lr = 0,100, mirostat_ent = 5,000
sampler chain: logits -> logit-bias -> penalties -> dry -> top-k -> typical -> top-p -> min-p -> xtc -> temp-ext -> dist
generate: n_ctx = 128000, n_batch = 2048, n_predict = -1, n_keep = 0
== Running in interactive mode. ==
- Press Ctrl+C to interject at any time.
- Press Return to return control to the AI.
- To return control without starting a new line, end your input with '/'.
- If you want to submit another line, end your input with '\'.
system
Provide a full description.
> Describe yourself and your goals
1141:14:44:14:114 *1:41:44 *1:4 A4 A4:1:4 A. *4: * A: * A4:4 * * A: * * * * * A:4: * * * * * * * * * * A:4 * A * * 1: * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * A * * * * * A * * * * A * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * A * * * * * * * * * * * A * * * * * * A * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * A * * * * * * * * * * * * * * * * * * A * * * * * * * * * * * * * * * * * * * * * * * * * * * * *^C^C^C^C^C^C
Ref, hand-compiled:
llama-cli --version
version: 4591 (7919256c)
built with Ubuntu clang version 18.1.8
works as a regular language model but no vision support(LM Studio)
Do we need a mmproj file to enable the vision? Was this split out of the model when converted to gguf?
For me both my bot this quant generate gibberish.
I assume the vision adapter extractor (..._surgery...py) is not implemented (or not correctly) in llama.cpp (not even tried as the language part not working - at least for me -)
so the vision adapters not extracted.
This was a quick and dirty conversion, mmproj is missing, and it only works properly as a LM not a VLM. Await a better version, with official support