--- license: gemma base_model: google/gemma-2-9b pipeline_tag: text-generation language: - en tags: - gemma - text-generation - rl-mpq - mixed-precision - quantization - fake-quantization - conservative library_name: transformers datasets: - wikitext widget: - text: "The capital of France is" --- # Gemma 2 9B — RL-MPQ Conservative Standalone **RL-MPQ** (Reinforcement Learning Mixed-Precision Quantization) checkpoint for the **Conservative** scenario — a quantized variant of [google/gemma-2-9b](https://huggingface.co/google/gemma-2-9b). | Field | Value | |-------|-------| | **Base model** | [google/gemma-2-9b](https://huggingface.co/google/gemma-2-9b) | | **Scenario** | Conservative | | **Avg bits / weight** | 5.1429 | | **Compression vs FP16** | 3.1111× | | **WikiText-2 PPL** | 116.5244 | | **Layers** | 42 | | **Bit distribution** | `{'4': 30, '8': 12}` | | **Format** | Fake-quant FP16 + `rlmpq_policy.json` | **Collection:** [RL-MPQ — Gemma 2 9B](https://huggingface.co/collections/AvoCahDoe/rl-mpq-gemma-2-9b-6a2b1d9443205eab61ff447b) — all five scenarios for Gemma 2 9B. ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer repo = "AvoCahDoe/gemma-2-9b-rlmpq-conservative" model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="float16") tokenizer = AutoTokenizer.from_pretrained(repo) ``` ## Other Gemma 2 9B scenarios | Scenario | Avg bits | Compression | WikiText-2 PPL | |----------|----------|-------------|----------------| | [Aggressive](https://huggingface.co/AvoCahDoe/gemma-2-9b-rlmpq-aggressive) | 3.6667 | 4.3636x | 162.8437 | | [Balanced](https://huggingface.co/AvoCahDoe/gemma-2-9b-rlmpq-balanced) | 4.2857 | 3.7333x | 127.0798 | | [Extreme Survival](https://huggingface.co/AvoCahDoe/gemma-2-9b-rlmpq-extreme-survival) | 2.7857 | 5.7436x | 424.7991 | | [High Fidelity](https://huggingface.co/AvoCahDoe/gemma-2-9b-rlmpq-high-fidelity) | 7.0476 | 2.2703x | 104.8098 | Grouped archive (all scenarios in one repo): [AvoCahDoe/gemma-2-9b-rlmpq](https://huggingface.co/AvoCahDoe/gemma-2-9b-rlmpq) ## Method 1. **Phase 3** — PPO agent assigns per-layer bit widths under the Conservative reward target. 2. **Phase 4** — Policy replayed on real weights; WikiText-2 perplexity validates quality. 3. **Export** — Fake-quantized FP16 weights compatible with Hugging Face Transformers. ## Files | File | Description | |------|-------------| | `config.json` | Llama architecture + RL-MPQ metadata | | `model.safetensors` | Fake-quantized weights | | `rlmpq_policy.json` | Per-layer bit-width policy | | `rlmpq_metrics.json` | Validation & PPL summary | ## Citation ```bibtex @misc{rlmpq_gemma_2_9b_conservative_2026, title = {RL-MPQ Conservative: Gemma 2 9B Mixed-Precision Quantization}, author = {AvoCahDoe}, year = {2026}, url = {https://huggingface.co/AvoCahDoe/gemma-2-9b-rlmpq-conservative} } ```