--- license: apache-2.0 base_model: Qwen/Qwen2.5-7B pipeline_tag: text-generation language: - en tags: - qwen2 - text-generation - rl-mpq - mixed-precision - quantization - fake-quantization - aggressive library_name: transformers datasets: - wikitext widget: - text: "The capital of France is" --- # Qwen 2.5 7B — RL-MPQ Aggressive Standalone **RL-MPQ** (Reinforcement Learning Mixed-Precision Quantization) checkpoint for the **Aggressive** scenario — a quantized variant of [Qwen/Qwen2.5-7B](https://huggingface.co/Qwen/Qwen2.5-7B). | Field | Value | |-------|-------| | **Base model** | [Qwen/Qwen2.5-7B](https://huggingface.co/Qwen/Qwen2.5-7B) | | **Scenario** | Aggressive | | **Avg bits / weight** | 3.1429 | | **Compression vs FP16** | 5.0909× | | **WikiText-2 PPL** | 9.3678 | | **Layers** | 28 | | **Bit distribution** | `{'3': 24, '4': 4}` | | **Format** | Fake-quant FP16 + `rlmpq_policy.json` | ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer repo = "AvoCahDoe/qwen2-5-7b-rlmpq-aggressive" model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="float16") tokenizer = AutoTokenizer.from_pretrained(repo) ``` ## Other Qwen 2.5 7B scenarios | Scenario | Avg bits | Compression | WikiText-2 PPL | |----------|----------|-------------|----------------| | [Balanced](https://huggingface.co/AvoCahDoe/qwen2-5-7b-rlmpq-balanced) | 3.3929 | 4.7158x | 8.9305 | | [Conservative](https://huggingface.co/AvoCahDoe/qwen2-5-7b-rlmpq-conservative) | 3.6786 | 4.3495x | 8.4114 | | [Extreme Survival](https://huggingface.co/AvoCahDoe/qwen2-5-7b-rlmpq-extreme-survival) | 2.4643 | 6.4928x | 497.4791 | | [High Fidelity](https://huggingface.co/AvoCahDoe/qwen2-5-7b-rlmpq-high-fidelity) | 3.75 | 4.2667x | 8.208 | Grouped archive (all scenarios in one repo): [AvoCahDoe/qwen2-5-7b-rlmpq](https://huggingface.co/AvoCahDoe/qwen2-5-7b-rlmpq) ## Method 1. **Phase 3** — PPO agent assigns per-layer bit widths under the Aggressive reward target. 2. **Phase 4** — Policy replayed on real weights; WikiText-2 perplexity validates quality. 3. **Export** — Fake-quantized FP16 weights compatible with Hugging Face Transformers. ## Files | File | Description | |------|-------------| | `config.json` | Llama architecture + RL-MPQ metadata | | `model.safetensors` | Fake-quantized weights | | `rlmpq_policy.json` | Per-layer bit-width policy | | `rlmpq_metrics.json` | Validation & PPL summary | ## Citation ```bibtex @misc{rlmpq_qwen2_5_7b_aggressive_2026, title = {RL-MPQ Aggressive: Qwen 2.5 7B Mixed-Precision Quantization}, author = {AvoCahDoe}, year = {2026}, url = {https://huggingface.co/AvoCahDoe/qwen2-5-7b-rlmpq-aggressive} } ```