donghyunli's picture
Add library_name metadata and clean up tags (#1)
4817c51
|
Raw
History Blame Contribute Delete
1.55 kB
metadata
base_model: meta-llama/Llama-2-13b-hf
language:
  - en
license: llama2
pipeline_tag: text-generation
library_name: transformers
tags:
  - kronq
  - quantization
  - group-quantization
  - fake-quant
  - fp16

Llama-2-13b — KronQ W3A16 g128 (fake-quant fp16)

Paper: arXiv:2607.07964 · Code: GitHub

⚠️ Fake-quant fp16 checkpoint. 3-bit group-128 weights stored in fp16 (KronQ does not pack int3) — same size as bf16, for PPL/accuracy reproduction. For deployable low-bit see the W4A16-g128 / W2A16-g128 (packed) repos.

Llama-2-13b quantized to 3-bit weights (group 128) with KronQ, exported as a standard fp16 model.

Results (WikiText-2, seqlen 2048)

Perplexity: 5.135

Zero-shot accuracy:

PIQA ARC-E ARC-C HellaSwag WinoGrande BoolQ OBQA Average
79.22 76.30 48.72 77.53 72.14 81.62 44.60 68.59

(lm-evaluation-harness 0-shot.)

Usage

Loads as a standard fp16 model (no KronQ code):

from transformers import AutoModelForCausalLM
m = AutoModelForCausalLM.from_pretrained("donghyunli/Llama-2-13b-KronQ-W3A16-g128-fake", torch_dtype="float16", device_map="auto")

Recipe

Group-128 asymmetric W3, weight-only, --alpha 0.25, --act_order, BiIP, raw H_G.

License

Derivative of Llama-2-13b — llama2 license.