Hariharasubramanian commited on
Commit
c6af05b
·
verified ·
1 Parent(s): 4697915

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +69 -0
README.md ADDED
@@ -0,0 +1,69 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: llama3.2
3
+ library_name: qualcomm-ai-runtime
4
+ tags:
5
+ - qualcomm
6
+ - qnn
7
+ - htp
8
+ - npu
9
+ - edge-ai
10
+ - z4-quantized
11
+ - llama
12
+ base_model: meta-llama/Llama-3.2-1B-Instruct
13
+ ---
14
+
15
+ # Llama 3.2 1B Instruct - QNN HTP Z4 Quantized
16
+
17
+ Pre-compiled model binary for **Qualcomm Hexagon HTP NPU** inference using **QNN GenAI Transformer** backend.
18
+
19
+ ## Model Details
20
+
21
+ | Property | Value |
22
+ |----------|-------|
23
+ | Base Model | [meta-llama/Llama-3.2-1B-Instruct](https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct) |
24
+ | Quantization | Z4 (Qualcomm 4-bit) |
25
+ | Format | QNN GenAI Transformer single binary |
26
+ | SDK Version | QAIRT v2.38.0.250901 |
27
+ | Target Hardware | Qualcomm Hexagon HTP v73+ NPU |
28
+ | Binary Size | ~1.6 GB |
29
+ | Tensors | 147 (Z4 for weights, F32 for norms) |
30
+
31
+ ## Tested Hardware
32
+
33
+ - **Qualcomm IQ-9075 EVK** (QCS9075 SoC, Hexagon HTP v73)
34
+
35
+ ## Usage
36
+
37
+ This binary is designed for use with the **Qualcomm Genie** runtime (`libGenie.so`) or `genie-t2t-run` CLI.
38
+
39
+ ### With genie-t2t-run
40
+
41
+ ```bash
42
+ cd /path/to/model/
43
+ LD_LIBRARY_PATH=/path/to/qnn-libs:/usr/lib \
44
+ ADSP_LIBRARY_PATH="/usr/lib/dsp/cdsp;/usr/lib/dsp/cdsp1" \
45
+ genie-t2t-run -c genie_config.json -p "Hello, how are you?"
46
+ ```
47
+
48
+ ### genie_config.json
49
+
50
+ ```json
51
+ {
52
+ "dialog": {
53
+ "backend": "QnnGenAiTransformer",
54
+ "model-path": "llama3.2-1b-instruct-z4.bin",
55
+ "tokenizer": "tokenizer.json"
56
+ }
57
+ }
58
+ ```
59
+
60
+ ## Compilation
61
+
62
+ Compiled using QAIRT SDK v2.38.0 GenAI Transformer Composer:
63
+ - Source: `meta-llama/Llama-3.2-1B-Instruct` (HuggingFace)
64
+ - Quantization: Z4 (4-bit weights, F32 normalization layers)
65
+ - Compile time: ~0.3 minutes on x86_64
66
+
67
+ ## License
68
+
69
+ This model inherits the [Llama 3.2 Community License](https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct/blob/main/LICENSE).