📦 Qwen3.5-2B GGUF (Quantized)

Custom quantized versions of the Qwen3.5-2B model in .gguf format, optimized for efficient CPU inference via llama.cpp. Designed for lightweight deployment on systems with limited RAM (4–8 GB).

📁 Available Files

File Size Quantization Effective BPW Purpose
Qwen3.5-2B-Q4_K_M.gguf ~1.2 GB Q4_K_M ~5.03 Primary: Best balance of quality & speed for chat/code
Qwen3.5-2B-fp16.gguf ~4.1 GB FP16 16.0 🔧 Reference format for re-quantization or GPU inference

⚡ Quick Start

1. Download

# Download only the recommended Q4_K_M version (~1.2 GB)
hf download YOUR_USERNAME/Qwen3.5-2B-GGUF \
  --include "Qwen3.5-2B-Q4_K_M.gguf" \
  --local-dir ./models
Downloads last month
17
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for vremiks/Qwen3.5-2B-GGUF

Finetuned
Qwen/Qwen3.5-2B
Quantized
(150)
this model