Text Generation
Transformers
Safetensors
English
qwen2
rl-mpq
mixed-precision
quantization
fake-quantization
aggressive
conversational
text-generation-inference
Instructions to use AvoCahDoe/qwen2-5-7b-rlmpq-aggressive with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AvoCahDoe/qwen2-5-7b-rlmpq-aggressive with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AvoCahDoe/qwen2-5-7b-rlmpq-aggressive") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("AvoCahDoe/qwen2-5-7b-rlmpq-aggressive") model = AutoModelForCausalLM.from_pretrained("AvoCahDoe/qwen2-5-7b-rlmpq-aggressive", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AvoCahDoe/qwen2-5-7b-rlmpq-aggressive with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AvoCahDoe/qwen2-5-7b-rlmpq-aggressive" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AvoCahDoe/qwen2-5-7b-rlmpq-aggressive", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/AvoCahDoe/qwen2-5-7b-rlmpq-aggressive
- SGLang
How to use AvoCahDoe/qwen2-5-7b-rlmpq-aggressive with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AvoCahDoe/qwen2-5-7b-rlmpq-aggressive" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AvoCahDoe/qwen2-5-7b-rlmpq-aggressive", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AvoCahDoe/qwen2-5-7b-rlmpq-aggressive" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AvoCahDoe/qwen2-5-7b-rlmpq-aggressive", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use AvoCahDoe/qwen2-5-7b-rlmpq-aggressive with Docker Model Runner:
docker model run hf.co/AvoCahDoe/qwen2-5-7b-rlmpq-aggressive
File size: 2,725 Bytes
83951ed | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 | ---
license: apache-2.0
base_model: Qwen/Qwen2.5-7B
pipeline_tag: text-generation
language:
- en
tags:
- qwen2
- text-generation
- rl-mpq
- mixed-precision
- quantization
- fake-quantization
- aggressive
library_name: transformers
datasets:
- wikitext
widget:
- text: "The capital of France is"
---
# Qwen 2.5 7B — RL-MPQ Aggressive
Standalone **RL-MPQ** (Reinforcement Learning Mixed-Precision Quantization) checkpoint for the
**Aggressive** scenario — a quantized variant of
[Qwen/Qwen2.5-7B](https://huggingface.co/Qwen/Qwen2.5-7B).
| Field | Value |
|-------|-------|
| **Base model** | [Qwen/Qwen2.5-7B](https://huggingface.co/Qwen/Qwen2.5-7B) |
| **Scenario** | Aggressive |
| **Avg bits / weight** | 3.1429 |
| **Compression vs FP16** | 5.0909× |
| **WikiText-2 PPL** | 9.3678 |
| **Layers** | 28 |
| **Bit distribution** | `{'3': 24, '4': 4}` |
| **Format** | Fake-quant FP16 + `rlmpq_policy.json` |
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "AvoCahDoe/qwen2-5-7b-rlmpq-aggressive"
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="float16")
tokenizer = AutoTokenizer.from_pretrained(repo)
```
## Other Qwen 2.5 7B scenarios
| Scenario | Avg bits | Compression | WikiText-2 PPL |
|----------|----------|-------------|----------------|
| [Balanced](https://huggingface.co/AvoCahDoe/qwen2-5-7b-rlmpq-balanced) | 3.3929 | 4.7158x | 8.9305 |
| [Conservative](https://huggingface.co/AvoCahDoe/qwen2-5-7b-rlmpq-conservative) | 3.6786 | 4.3495x | 8.4114 |
| [Extreme Survival](https://huggingface.co/AvoCahDoe/qwen2-5-7b-rlmpq-extreme-survival) | 2.4643 | 6.4928x | 497.4791 |
| [High Fidelity](https://huggingface.co/AvoCahDoe/qwen2-5-7b-rlmpq-high-fidelity) | 3.75 | 4.2667x | 8.208 |
Grouped archive (all scenarios in one repo):
[AvoCahDoe/qwen2-5-7b-rlmpq](https://huggingface.co/AvoCahDoe/qwen2-5-7b-rlmpq)
## Method
1. **Phase 3** — PPO agent assigns per-layer bit widths under the Aggressive reward target.
2. **Phase 4** — Policy replayed on real weights; WikiText-2 perplexity validates quality.
3. **Export** — Fake-quantized FP16 weights compatible with Hugging Face Transformers.
## Files
| File | Description |
|------|-------------|
| `config.json` | Llama architecture + RL-MPQ metadata |
| `model.safetensors` | Fake-quantized weights |
| `rlmpq_policy.json` | Per-layer bit-width policy |
| `rlmpq_metrics.json` | Validation & PPL summary |
## Citation
```bibtex
@misc{rlmpq_qwen2_5_7b_aggressive_2026,
title = {RL-MPQ Aggressive: Qwen 2.5 7B Mixed-Precision Quantization},
author = {AvoCahDoe},
year = {2026},
url = {https://huggingface.co/AvoCahDoe/qwen2-5-7b-rlmpq-aggressive}
}
```
|