Qwen3-8B — 2-bit BPDQ (packed)

Qwen3-8B with the transformer weights stored as 2-bit bit-planes and kept that way for the whole run. They are never expanded into a dense fp16 tensor. 3.0 GB on disk, 3.65 GiB resident, against 15.3 GiB for the same weights as fp16.

Runs on one CUDA GPU or an Apple M-series Mac.

Run it

Three files sit on top of the weights: bpdq.py, demo.ipynb and this README. bpdq.py works out what it is running on and picks the kernels.

import sys; sys.path.insert(0, MODEL_DIR)
import bpdq

model, tok = bpdq.load(MODEL_DIR)                      # cuda | mps | cpu
ids = tok("The capital of France is", return_tensors="pt").to(model.device)
out, s_per_step = bpdq.decode(model, ids.input_ids, 64,
                              sampling=bpdq.Sampling.for_mode(think=False))
print(tok.decode(out[0]))

decode takes a batch; ragged prompts are left-padded (tok.padding_side = "left", pass the attention_mask). Pass a transformers streamer to stream.

On CUDA, serve through vLLM — importing bpdq registers the method:

import sys; sys.path.insert(0, MODEL_DIR)
import bpdq
from vllm import LLM

llm = LLM(model=MODEL_DIR, quantization="bpdq", dtype="float16",
          max_num_seqs=64, reasoning_parser="qwen3")

demo.ipynb detects the machine and gives one generate() either way.

python bpdq.py selftest [MODEL_DIR]   # kernels vs a dense reconstruction
python bpdq.py bench    [MODEL_DIR]   # per-matmul and end-to-end
python bpdq.py chat      MODEL_DIR    # interactive, streaming

Decoding

bpdq.Sampling.for_mode(think) carries Qwen3's published settings:

temperature top_p top_k min_p
think 0.6 0.95 20 0
nothink 0.7 0.80 20 0

Do not decode this checkpoint greedily. At 2 bits it falls into restating the same paragraph. Sampling also sets no_repeat_ngram=12 on the torch path and presence_penalty=0.3 on the vLLM path — each backend's native control for that.

It rarely stops thinking on its own. On vLLM, Sampling(think_budget=N).vllm() sets thinking_token_budget, which needs LLM(..., reasoning_parser="qwen3") so it knows where </think> is. The torch path has no equivalent; give think mode a generous max_tokens there.

Measured

RTX 4090, vLLM 0.28, 64-token decode. The fp16 column is these same weights reconstructed dense and served by the same vLLM, so only the matmul differs.

concurrent seqs packed 2-bit dense fp16
1 322 tok/s 59 5.46x
4 1246 226 5.51x
16 3452 890 3.88x
32 5327 1703 3.13x
64 6761 3104 2.18x
128 7695 5353 1.44x
256 8130 8110 1.00x
weights resident 3.65 GiB 15.26 GiB

One sequence reads 2.58 GB of weights per token; at the 918 GB/s a device-to-device copy reaches, that bounds decode at 356 tok/s. At 256–512 tokens gate_up and down run at 153–168 TFLOPS (cuBLAS fp16 peaks near 166 TFLOPS on this card).

Apple M4 (10-core GPU, 17 GB unified), 230-token prompt. fp16 Qwen3-8B does not fit on this machine.

batch decode prefill
1 31.3 tok/s 169 tok/s
4 110.6 173
8 144.1 168
16 137.7 168

Per matmul against dense fp16 of the same shape: 4.7–5.6x at 1 token, 3.2–3.9x at 8, 1.1–1.4x at 32, 0.7–0.9x from 64 up. From 1 to 4 tokens a matmul takes at most 1.25x the time to read its weights at the 98 GB/s a device-to-device copy reaches.

Quality

value
WikiText-2 PPL (seq 2048) 18.58    (fp16 Qwen3-8B: 9.72)
KL to fp16 (32 chunks) 0.864
GSM8K think, 8-shot, 250q, 1024 tok 0.588

The packed path and a dense reconstruction of the same weights agree to fp16 round-off — worst relative error 5e-4 over all 252 weight matrices, checked by bpdq.py selftest.

Format

w[r, c] = sum_i coeffs[g(c), r, i] * bit_i(r, c) + coeffs[g(c), r, msbits]

Two bit-planes per weight, group_size = 256, one fp16 coefficient set per (group, output row) plus a per-group offset — about 2.2 bits per weight. embed_tokens and lm_head ride as 8-bit GPTQ-v2 integer sets. embed_tokens is the only tensor expanded to fp16 at load; lm_head stays 8-bit on CUDA (compute capability 8.0+) and Apple.

At load the packed tensors are repacked, not decoded. On CUDA (compute capability 8.0+) the codes are laid out in tensor-core fragment order:

codes   int32  [out / 16, in / 128, 32, 4]     2-bit: one lane's A fragments for 128 columns
codes   int32  [out / 16, in / 128, 32, 16]    8-bit lm_head
table   fp16   [groups, out, 4] | [groups, out, 2]   value table | (scale, 1024 + zero)

Elsewhere:

codes   int32  [in_features / 16, out_features]    16 columns per word, 2 bits each
coeffs  fp32   [groups, 3, out_features]

out_features is innermost so a warp (or SIMD group) covering 32 rows reads one contiguous line, and a column's two bits are interleaved into one code so a weight is peeled with code = word & 3; word >>= 2.

Kernels

CUDA is one hand-written kernel, compiled at import with NVRTC (no CUDA toolkit needed). Each lane loads exactly its mma.m16n8k16 A fragments, turns 2-bit codes into fp16 with one byte permute into the per-(group, row) value table (8-bit: permute, subtract, scale), and multiplies against activations staged through shared memory with ldmatrix. The tile shape is picked per batch size; K is split across CTAs with an in-kernel deterministic reduction when a launch would not fill the card. Output is fp16. While a forward runs at most 8 tokens under vLLM, a second stream loads each weight matrix into L2 two matmuls ahead of its use, paced by a progress flag the matmuls write. GPUs below compute capability 8.0 get a Triton kernel instead.

Apple multiplies without dequantising. Per 256-column group a row's output is c0 * S0 + c1 * S1 + bias * sum(x), where S0 and S1 are the sums of the activations each bit-plane selects. For every 8 columns a threadgroup builds a 256-entry table of all subset sums of those 8 activations, 4 tokens per entry, and one byte of a bit-plane indexes it; the planes are the checkpoint's own, transposed at load. From 8 tokens on, the kernel is bound by random reads of those tables. The 8-bit lm_head has its own Metal kernel. One-token decode attention folds each KV head's query heads into one pass instead of repeating the KV cache, and the decode step runs under torch.compile; the first step compiles for about 30 s.

Limits

  • TP = 1. No tensor parallelism.
  • On Apple, batches of 64 tokens and up run at 0.7–0.9x of a dense fp16 matmul. The win is at decode batch sizes.
  • At 2 bits the model invents confident detail about anything obscure, and past roughly 250 tokens long open-ended answers drift. Use it for short exchanges and bounded tasks.
  • The torch decode loop does not capture CUDA graphs; transformers bakes part of the attention mask into a capture and the replayed step drifts. vLLM's own capture is fine.
  • Greedy output is not bit-stable across cache implementations — this loop uses a StaticCache, model.generate a dynamic one. Both are faithful to the weights.
  • Launch shapes were measured on an RTX 4090 and an M4; other chips run correctly but may not be at their own optimum.

License

Apache-2.0, inheriting Qwen/Qwen3-8B.

Downloads last month
1,039
Safetensors
Model size
0.8B params
Tensor type
I32
·
F16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for gitarist/Qwen3-8B-BPDQ-2bit

Finetuned
Qwen/Qwen3-8B
Quantized
(458)
this model