Llama-2-13b — KronQ W2A16 (packed int2)
Paper: arXiv:2607.07964 · Code: GitHub
Llama-2-13b quantized to 2-bit weights / 16-bit activations with KronQ (Kronecker-factored Hessian quantization). Weights are stored packed int2 (~3.9 GB vs 26 GB fp16); a fused dequant + bidirectional-incoherence (BiIP) CUDA kernel unpacks them on the fly.
Results (WikiText-2, seqlen 2048)
Perplexity: 6.99
Zero-shot accuracy:
| PIQA | ARC-E | ARC-C | HellaSwag | WinoGrande | BoolQ | OBQA | Average |
|---|---|---|---|---|---|---|---|
| 72.42 | 62.46 | 36.35 | 59.42 | 65.11 | 69.42 | 37.00 | 57.45 |
(lm-evaluation-harness, 0-shot. acc_norm for PIQA/HellaSwag/ARC/OBQA, acc for WinoGrande/BoolQ.)
Usage
KronQ-packed checkpoint (model.safetensors carries biip_w_codes/scale/zero + BiIP buffers, see kronq_packed_config.json). Load with the KronQ runtime:
python eval_pretrained.py meta-llama/Llama-2-13b-hf donghyunli/Llama-2-13b-KronQ-W2A16 --ppl --zs
Recipe
Per-channel asymmetric W2, weight-only (a_bits=16), --alpha 0.25, bidirectional incoherence processing (BiIP), act_order, raw H_G. Calibrated on 128 WikiText-2 sequences.
License
Derivative of Llama-2-13b — subject to the Llama 2 Community License.
Model tree for donghyunli/Llama-2-13b-KronQ-W2A16
Base model
meta-llama/Llama-2-13b-hf