sahilchachra's picture
Upload README.md with huggingface_hub
9805754 verified
|
Raw
History Blame
3.07 kB
metadata
license: apache-2.0
base_model: bottlecapai/ThinkingCap-Qwen3.6-27B
base_model_relation: quantized
pipeline_tag: text-generation
tags:
  - nvfp4
  - fp4
  - nvfp4a16
  - compressed-tensors
  - llm-compressor
language:
  - en

ThinkingCap-Qwen3.6-27B-NVFP4A16

NVFP4A16 quantization of bottlecapai/ThinkingCap-Qwen3.6-27B — a token-efficient reasoning finetune of Qwen/Qwen3.6-27B by BottleCap AI (qwen3_5: dense hybrid GatedDeltaNet linear-attention + full-attention over 64 text layers in a 3:1 pattern, plus a vision tower and an MTP head; multimodal image-text-to-text). It matches Qwen3.6-27B answer quality while emitting ~50% fewer thinking tokens on average.

Variant: NVFP4 A16 — 4-bit NVFP4 (FP4 E2M1, group size 16) weights, activations BF16. Native on NVIDIA Blackwell. Quantized by: sahilchachra Tooling: llm-compressor model_free_ptq (data-free, RTN) -> compressed-tensors

This is a quantized derivative. Weights, behavior, and license follow the base model — see the original card for full details, benchmarks, and citation.

What is quantized

Quantized to 4-bit:

  • full-attention self_attn.{q,k,v,o}_proj
  • mlp.{gate,up,down}_proj (all text layers)

Kept in BF16: GatedDeltaNet linear_attn (mamba) layers, vision tower (model.visual.*, 27 blocks), MTP head, token embeddings, lm_head, all norms (incl. q_norm / k_norm).

Calibration

Data-free — weight-only (model_free_ptq, round-to-nearest); no calibration data. Weights are quantized by streaming the safetensors from disk.

Prompt template & sampling

This is a reasoning ("thinking") model. Use the Qwen3.6 chat template — ChatML (<|im_start|>role … <|im_end|>) with a <think>…</think> reasoning trace, thinking enabled by default. Apply it via tokenizer.apply_chat_template(messages, add_generation_prompt=True) (or the processor for image inputs); do not hand-format prompts. See the base Qwen/Qwen3.6-27B for full usage details.

Recommended sampling: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0 with thinking on (the base model's recommended sampling, per the original card).

Usage (vLLM)

from vllm import LLM, SamplingParams

# This is a multimodal checkpoint: the vision tower is kept in BF16
# (only the text / MoE weights are 4-bit). vLLM builds the full model.
llm = LLM(
    model="sahilchachra/ThinkingCap-Qwen3.6-27B-NVFP4A16",
    trust_remote_code=True,
)
out = llm.chat(
    [{"role": "user", "content": "Hello!"}],
    SamplingParams(temperature=0.6, top_p=0.95, max_tokens=512),
)
print(out[0].outputs[0].text)

Serving via the CLI, pass the flag directly:

vllm serve sahilchachra/ThinkingCap-Qwen3.6-27B-NVFP4A16 \
    --trust-remote-code \
    --max-model-len 262144 --reasoning-parser qwen3