| --- |
| base_model: |
| - meta-llama/Llama-2-70b-hf |
| base_model_relation: quantized |
| license: llama2 |
| --- |
| # Model Card |
|
|
| - Base model: `meta-llama/Llama-2-70b-hf` |
| - Quantization method: SqueezeLLM |
| - Target bit-width: 2 |
| - Backend kernel: Any-Precision-LLM kernel (`ap-gemv`) |
| - Calibration data: RedPajama (1024 sentences / 4096 tokens) |
| - Calibration objective: Next-token prediction |
|
|
| # How to run |
| - Follow the instruction in https://github.com/snu-mllab/GuidedQuant. |
|
|
| # References |
| - [Model Paper](https://arxiv.org/abs/2505.07004) |