How to use from
Hermes Agent
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf Abiray/Qwen3.8-27B-Q4_K_M-GGUF:Q4_K_M
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default Abiray/Qwen3.8-27B-Q4_K_M-GGUF:Q4_K_M
Run Hermes
hermes
Quick Links

Qwen3.8-27B GGUF (Vision-Language)

This repository contains the GGUF format quantization of the Qwen3.8-27B model, a native vision-language model.

The model has been quantized using llama.cpp (release b10430) to the Q4_K_M format.

File Details

  • Model Name: Qwen3.8-27B
  • Quantization: Q4_K_M
  • Main Model: Qwen3.8-27B-Q4_K_M.gguf (16.8 GB)
  • Vision Adapter: mmproj-F16.gguf (Required for vision/multimodal tasks)
  • Quantization Tool: llama.cpp (b10430)

About Qwen3.8-27B

Qwen3.8-27B is the most capable generation in the Qwen open-model family, built on the architectural foundation of Qwen3.5. It is a native vision-language model that understands images and videos, delivering substantial gains across coding, professional work, research, and long-horizon agentic tasks.

For more information, please visit the original model card.

Usage

llama.cpp (Text Only)

To run this model as a text-only model using llama.cpp from the command line:

./llama-cli -m Qwen3.8-27B-Q4_K_M.gguf -p "Write a Python function to merge two sorted linked lists." -n 5124

llama.cpp (Vision-Language)

To utilize the vision capabilities, you must provide the mmproj file using the --mmproj flag:

./llama-cli -m Qwen3.8-27B-Q4_K_M.gguf \
  --mmproj mmproj-F16.gguf \
  --image path/to/your/image.jpg \
  -p "Describe the contents of this image."

LM Studio / Ollama / GPT4All

This GGUF file is compatible with popular local LLM inference tools. When configuring the model in your preferred client, ensure you attach the mmproj-F16.gguf file in the "Vision Adapter" or "Multimedia Projection" setting to enable image processing.

Acknowledgements

The original model was developed and released by the Qwen Team.

Downloads last month
45,238
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Abiray/Qwen3.8-27B-Q4_K_M-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(948)
this model