AvoCahDoe commited on
Commit
415456c
·
verified ·
1 Parent(s): 3644411

RL-MPQ repo index (README + config.json) — 2026-06-11T15:57:56.128906

Browse files
Files changed (1) hide show
  1. README.md +41 -20
README.md CHANGED
@@ -2,20 +2,36 @@
2
  license: llama2
3
  base_model: meta-llama/Llama-2-7b-hf
4
  pipeline_tag: text-generation
 
 
5
  tags:
 
 
6
  - rl-mpq
7
  - mixed-precision
8
  - quantization
9
- - text-generation
 
10
  library_name: transformers
 
 
11
  ---
12
 
13
- # RL-MPQ — LLAMA-2-7B (all scenarios)
 
 
 
 
14
 
15
- This repository groups **five RL-MPQ quantization scenarios** for
16
- [meta-llama/Llama-2-7b-hf](https://huggingface.co/meta-llama/Llama-2-7b-hf) in one place.
17
 
18
- Each scenario is a separate subfolder with fake-quantized FP16 weights, tokenizer, and `rlmpq_policy.json`.
 
 
 
 
 
19
 
20
  ## Scenarios
21
 
@@ -25,41 +41,46 @@ Each scenario is a separate subfolder with fake-quantized FP16 weights, tokenize
25
  - [`Aggressive/`](./Aggressive/) — load with `subfolder="Aggressive"`
26
  - [`Extreme_Survival/`](./Extreme_Survival/) — load with `subfolder="Extreme_Survival"`
27
 
28
- ## Summary table
29
 
30
- | Scenario | Avg bits | Compression | WikiText-2 PPL |
31
- |----------|----------|-------------|----------------|
32
  | [High_Fidelity](./High_Fidelity/) | 6.5 | 2.4615x | 4.9808 |
33
  | [Conservative](./Conservative/) | 5.125 | 3.122x | 5.0276 |
34
  | [Balanced](./Balanced/) | 4.375 | 3.6571x | 5.0437 |
35
  | [Aggressive](./Aggressive/) | 3.5938 | 4.4522x | 5.2614 |
36
  | [Extreme_Survival](./Extreme_Survival/) | 2.9688 | 5.3895x | 10.9577 |
37
 
38
- ## Quick load (example: Balanced)
39
 
40
  ```python
41
  from transformers import AutoModelForCausalLM, AutoTokenizer
42
 
43
  repo = "AvoCahDoe/llama-2-7b-rlmpq"
44
- scenario = "Balanced" # or High_Fidelity, Conservative, Aggressive, Extreme_Survival
45
 
46
  model = AutoModelForCausalLM.from_pretrained(repo, subfolder=scenario, torch_dtype="float16")
47
  tokenizer = AutoTokenizer.from_pretrained(repo, subfolder=scenario)
48
  ```
49
 
50
- ## Method
 
51
 
52
- PPO-trained per-layer bit-width policies (Phase 3), validated with env replay + WikiText-2 PPL (Phase 4-new).
53
- Weights are fake-quantized in FP16 storage (not GPTQ/AWQ packed format).
54
 
55
- ## Download stats
56
-
57
- Hub download counts increment when `config.json` is fetched (root index or scenario subfolder).
58
- Use `from_pretrained(repo, subfolder="<scenario>")` so the scenario `config.json` is requested.
59
- The root `config.json` is a collection index only — always pass `subfolder` to load weights.
60
 
61
  ## Citation
62
 
63
- RL-NMP-Model-Quantasation — RL-MPQ thesis framework.
 
 
 
 
 
 
 
64
 
65
- Generated: 2026-06-11T15:48:16.624893
 
2
  license: llama2
3
  base_model: meta-llama/Llama-2-7b-hf
4
  pipeline_tag: text-generation
5
+ language:
6
+ - en
7
  tags:
8
+ - llama
9
+ - text-generation
10
  - rl-mpq
11
  - mixed-precision
12
  - quantization
13
+ - fake-quantization
14
+ - llama-2
15
  library_name: transformers
16
+ datasets:
17
+ - wikitext
18
  ---
19
 
20
+ # Llama 2 7B — RL-MPQ Quantized (Thesis Release)
21
+
22
+ **Quantized variant of [meta-llama/Llama-2-7b-hf](https://huggingface.co/meta-llama/Llama-2-7b-hf)** using **RL-MPQ**
23
+ (Reinforcement Learning Mixed-Precision Quantization): per-layer bit-width policies
24
+ trained with PPO, validated on WikiText-2 perplexity.
25
 
26
+ This repo ships **five compression scenarios** as subfolders — from near-FP16 fidelity
27
+ to aggressive survival mode — so you can pick the bits-vs-quality trade-off for your thesis experiments.
28
 
29
+ | | |
30
+ |---|---|
31
+ | **Base model** | [meta-llama/Llama-2-7b-hf](https://huggingface.co/meta-llama/Llama-2-7b-hf) |
32
+ | **Method** | RL-MPQ (PPO per-layer bit policy) |
33
+ | **Format** | Fake-quant FP16 weights + `rlmpq_policy.json` |
34
+ | **Recommended start** | `subfolder="Balanced"` |
35
 
36
  ## Scenarios
37
 
 
41
  - [`Aggressive/`](./Aggressive/) — load with `subfolder="Aggressive"`
42
  - [`Extreme_Survival/`](./Extreme_Survival/) — load with `subfolder="Extreme_Survival"`
43
 
44
+ ## Results (WikiText-2)
45
 
46
+ | Scenario | Avg bits | Compression vs FP16 | Perplexity |
47
+ |----------|----------|---------------------|------------|
48
  | [High_Fidelity](./High_Fidelity/) | 6.5 | 2.4615x | 4.9808 |
49
  | [Conservative](./Conservative/) | 5.125 | 3.122x | 5.0276 |
50
  | [Balanced](./Balanced/) | 4.375 | 3.6571x | 5.0437 |
51
  | [Aggressive](./Aggressive/) | 3.5938 | 4.4522x | 5.2614 |
52
  | [Extreme_Survival](./Extreme_Survival/) | 2.9688 | 5.3895x | 10.9577 |
53
 
54
+ ## Usage
55
 
56
  ```python
57
  from transformers import AutoModelForCausalLM, AutoTokenizer
58
 
59
  repo = "AvoCahDoe/llama-2-7b-rlmpq"
60
+ scenario = "Balanced" # High_Fidelity | Conservative | Aggressive | Extreme_Survival
61
 
62
  model = AutoModelForCausalLM.from_pretrained(repo, subfolder=scenario, torch_dtype="float16")
63
  tokenizer = AutoTokenizer.from_pretrained(repo, subfolder=scenario)
64
  ```
65
 
66
+ > **Important:** Always pass `subfolder=<scenario>`. Root `config.json` describes the
67
+ > collection; weights and tokenizer live inside each scenario folder.
68
 
69
+ ## Method (thesis summary)
 
70
 
71
+ 1. **Phase 3** — PPO agent selects per-layer bit widths under scenario-specific reward targets.
72
+ 2. **Phase 4** — Policies replayed on real weights; WikiText-2 PPL measures quality retention.
73
+ 3. **Export** — Fake-quantized FP16 checkpoints (compatible with Hugging Face Transformers).
 
 
74
 
75
  ## Citation
76
 
77
+ ```bibtex
78
+ @misc{rlmpq2026,
79
+ title = {RL-MPQ: Reinforcement Learning Mixed-Precision Quantization},
80
+ author = {AvoCahDoe},
81
+ year = {2026},
82
+ url = {https://huggingface.co/AvoCahDoe/llama-2-7b-rlmpq}
83
+ }
84
+ ```
85
 
86
+ Part of the **RL-NMP-Model-Quantasation** thesis framework.