File size: 4,455 Bytes
b158e17
 
 
 
0a6d10a
 
b158e17
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
02e366a
 
 
 
 
 
 
 
b158e17
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
---
license: apache-2.0
base_model: bottlecapai/ThinkingCap-Qwen3.6-27B
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
tags:
  - awq
  - int4
  - w4a16
  - compressed-tensors
  - llm-compressor
language:
  - en
---

# ThinkingCap-Qwen3.6-27B-AWQ

AWQ (W4A16) quantization of
[bottlecapai/ThinkingCap-Qwen3.6-27B](https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B)
— a token-efficient **reasoning** finetune of [Qwen/Qwen3.6-27B](https://huggingface.co/Qwen/Qwen3.6-27B) by BottleCap AI (`qwen3_5`: dense hybrid GatedDeltaNet linear-attention + full-attention over 64 text layers in a 3:1 pattern, plus a vision tower and an MTP head; multimodal image-text-to-text). It matches Qwen3.6-27B answer quality while emitting ~50% fewer thinking tokens on average.

**Variant**: AWQ **W4A16** — 4-bit symmetric integer weights, group size 128, with activation-aware scaling. Activations stay BF16.
**Quantized by**: [sahilchachra](https://huggingface.co/sahilchachra)
**Tooling**: `llm-compressor` (`AWQModifier` + `QuantizationModifier`) -> `compressed-tensors` `pack-quantized`

> This is a quantized derivative. Weights, behavior, and license follow the base
> model — see the
> [original card](https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B) for full details, benchmarks, and citation.

## What is quantized

Quantized to 4-bit:

- full-attention `self_attn.{q,k,v,o}_proj`
- `mlp.{gate,up,down}_proj` (all text layers)

Kept in **BF16**: GatedDeltaNet `linear_attn` (mamba) layers, vision tower (`model.visual.*`, 27 blocks), MTP head, token embeddings, lm_head, all norms (incl. q_norm / k_norm).

### Note on what's 4-bit vs BF16 (quality/speed tradeoff)

This is a **hybrid** architecture: 48 of the 64 layers are GatedDeltaNet **linear-attention ("Mamba") layers**. Only the full-attention `self_attn.{q,k,v,o}` projections and the per-layer `mlp.{gate,up,down}` projections are quantized to 4-bit; the **Mamba `linear_attn` layers are deliberately kept in BF16** (along with the vision tower, MTP head, embeddings, lm_head and norms), because they are quantization-sensitive and keeping them full-precision preserves the reasoning quality and the concise `<think>` behavior.

As a result, roughly **two-thirds of the weight bytes read per token stay BF16** (~18 GB BF16 vs ~9 GB of 4-bit weights), so this variant's memory footprint is close to an 8-bit build and it is tuned for **quality rather than peak throughput** — on Blackwell it can run a little slower than a full W8A8/FP8 build of the base model.

The Mamba layers **can also be quantized** (to shrink the model further and speed up memory-bound decoding), but **accuracy may take a hit** — this build intentionally trades that extra speed for output quality.

## Calibration

AWQ: 128 sequences x 512 tokens of **GSM8K** (`openai/gsm8k`, config `main`, train split) — the model's headline in-domain reasoning dataset — rendered through the model's own chat template with the `<think>…</think>` reasoning format, so calibration matches the model's real (thinking) inference distribution. NVFP4 is data-free (no calibration).

## Prompt template & sampling

This is a **reasoning ("thinking") model**. Use the Qwen3.6 chat template — ChatML (`<|im_start|>role … <|im_end|>`) with a `<think>…</think>` reasoning trace, thinking enabled by default. Apply it via `tokenizer.apply_chat_template(messages, add_generation_prompt=True)` (or the processor for image inputs); do not hand-format prompts. See the base [Qwen/Qwen3.6-27B](https://huggingface.co/Qwen/Qwen3.6-27B) for full usage details.

**Recommended sampling**: `temperature=1.0, top_p=0.95, top_k=20, min_p=0.0` with thinking on (the base model's recommended sampling, per the original card).


## Usage (vLLM)

```python
from vllm import LLM, SamplingParams

# This is a multimodal checkpoint: the vision tower is kept in BF16
# (only the text / MoE weights are 4-bit). vLLM builds the full model.
llm = LLM(
    model="sahilchachra/ThinkingCap-Qwen3.6-27B-AWQ",
    trust_remote_code=True,
)
out = llm.chat(
    [{"role": "user", "content": "Hello!"}],
    SamplingParams(temperature=0.6, top_p=0.95, max_tokens=512),
)
print(out[0].outputs[0].text)
```

Serving via the CLI, pass the flag directly:

```bash
vllm serve sahilchachra/ThinkingCap-Qwen3.6-27B-AWQ \
    --trust-remote-code \
    --max-model-len 262144 --reasoning-parser qwen3
```