File size: 1,845 Bytes
ca3ea78
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
---
license: other
base_model: Qwen/Qwen3-8B
pipeline_tag: text-generation
tags:
- gguf
- local-llm
- llama.cpp
- lm-studio
- quantized
- imatrix
- sub-4-bit
- qwen3
- base_model:Qwen/Qwen3-8B-Base
---

# Qwen3-8B — iMatrix GGUF

GGUF quantizations of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B), published by [Liodon AI](https://huggingface.co/liodon-ai).

## Quick Start

**llama.cpp**
```bash
llama-cli -hf liodon-ai/Qwen3-8B-imatrix-GGUF:Q4_K_M
```

**Ollama**
```bash
ollama run hf.co/liodon-ai/Qwen3-8B-imatrix-GGUF:Q4_K_M
```

**LM Studio / Jan** — search `liodon-ai/Qwen3-8B-imatrix-GGUF` and pick your quant.

## Quants

| Quant | Size | VRAM est. | Notes |
|-------|------|-----------|-------|
| `IQ2_M` | 3.05 GB | ~4 GB | 2-bit, iMatrix — smallest usable |
| `IQ3_M` | 3.90 GB | ~4 GB | 3-bit, iMatrix — great quality/size tradeoff |
| `IQ4_XS` | 4.56 GB | ~5 GB | 4-bit extra-small, iMatrix |
| `Q4_K_M` | 5.03 GB | ~6 GB | 4-bit, iMatrix-calibrated (recommended) |
| `Q5_K_M` | 5.85 GB | ~7 GB | 5-bit, iMatrix-calibrated |
| `Q6_K` | 6.73 GB | ~8 GB | 6-bit, iMatrix-calibrated, near-lossless |
| `Q8_0` | 8.71 GB | ~10 GB | 8-bit, essentially lossless |


## What is iMatrix?

Standard quantization treats all weights equally. iMatrix runs 128 calibration chunks through
the full-precision model to find which weights matter most, then allocates more precision where
it counts. At Q2/Q3/Q4 this means noticeably better coherence and instruction-following —
**same file size, better output**.

Calibration: 2M tokens of [WikiText-103](https://huggingface.co/datasets/wikitext).

> Also see plain (non-iMatrix) quants: `liodon-ai/Qwen3-8B-GGUF`

## Source

- **Model**: [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)
- **License**: other

---
*Quantized by [Liodon AI](https://huggingface.co/liodon-ai)*