sahilchachra commited on
Commit
9805754
·
verified ·
1 Parent(s): c53ba11

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +74 -0
README.md ADDED
@@ -0,0 +1,74 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: bottlecapai/ThinkingCap-Qwen3.6-27B
4
+ base_model_relation: quantized
5
+ pipeline_tag: text-generation
6
+ tags:
7
+ - nvfp4
8
+ - fp4
9
+ - nvfp4a16
10
+ - compressed-tensors
11
+ - llm-compressor
12
+ language:
13
+ - en
14
+ ---
15
+
16
+ # ThinkingCap-Qwen3.6-27B-NVFP4A16
17
+
18
+ NVFP4A16 quantization of
19
+ [bottlecapai/ThinkingCap-Qwen3.6-27B](https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B)
20
+ — a token-efficient **reasoning** finetune of [Qwen/Qwen3.6-27B](https://huggingface.co/Qwen/Qwen3.6-27B) by BottleCap AI (`qwen3_5`: dense hybrid GatedDeltaNet linear-attention + full-attention over 64 text layers in a 3:1 pattern, plus a vision tower and an MTP head; multimodal image-text-to-text). It matches Qwen3.6-27B answer quality while emitting ~50% fewer thinking tokens on average.
21
+
22
+ **Variant**: NVFP4 **A16** — 4-bit NVFP4 (FP4 E2M1, group size 16) weights, activations BF16. Native on NVIDIA Blackwell.
23
+ **Quantized by**: [sahilchachra](https://huggingface.co/sahilchachra)
24
+ **Tooling**: `llm-compressor` `model_free_ptq` (data-free, RTN) -> `compressed-tensors`
25
+
26
+ > This is a quantized derivative. Weights, behavior, and license follow the base
27
+ > model — see the
28
+ > [original card](https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B) for full details, benchmarks, and citation.
29
+
30
+ ## What is quantized
31
+
32
+ Quantized to 4-bit:
33
+
34
+ - full-attention `self_attn.{q,k,v,o}_proj`
35
+ - `mlp.{gate,up,down}_proj` (all text layers)
36
+
37
+ Kept in **BF16**: GatedDeltaNet `linear_attn` (mamba) layers, vision tower (`model.visual.*`, 27 blocks), MTP head, token embeddings, lm_head, all norms (incl. q_norm / k_norm).
38
+
39
+ ## Calibration
40
+
41
+ Data-free — weight-only (`model_free_ptq`, round-to-nearest); no calibration data. Weights are quantized by streaming the safetensors from disk.
42
+
43
+ ## Prompt template & sampling
44
+
45
+ This is a **reasoning ("thinking") model**. Use the Qwen3.6 chat template — ChatML (`<|im_start|>role … <|im_end|>`) with a `<think>…</think>` reasoning trace, thinking enabled by default. Apply it via `tokenizer.apply_chat_template(messages, add_generation_prompt=True)` (or the processor for image inputs); do not hand-format prompts. See the base [Qwen/Qwen3.6-27B](https://huggingface.co/Qwen/Qwen3.6-27B) for full usage details.
46
+
47
+ **Recommended sampling**: `temperature=1.0, top_p=0.95, top_k=20, min_p=0.0` with thinking on (the base model's recommended sampling, per the original card).
48
+
49
+
50
+ ## Usage (vLLM)
51
+
52
+ ```python
53
+ from vllm import LLM, SamplingParams
54
+
55
+ # This is a multimodal checkpoint: the vision tower is kept in BF16
56
+ # (only the text / MoE weights are 4-bit). vLLM builds the full model.
57
+ llm = LLM(
58
+ model="sahilchachra/ThinkingCap-Qwen3.6-27B-NVFP4A16",
59
+ trust_remote_code=True,
60
+ )
61
+ out = llm.chat(
62
+ [{"role": "user", "content": "Hello!"}],
63
+ SamplingParams(temperature=0.6, top_p=0.95, max_tokens=512),
64
+ )
65
+ print(out[0].outputs[0].text)
66
+ ```
67
+
68
+ Serving via the CLI, pass the flag directly:
69
+
70
+ ```bash
71
+ vllm serve sahilchachra/ThinkingCap-Qwen3.6-27B-NVFP4A16 \
72
+ --trust-remote-code \
73
+ --max-model-len 262144 --reasoning-parser qwen3
74
+ ```