--- license: apache-2.0 base_model: Jackrong/Qwopus3.5-9B-v3 tags: - hlwq - quantized - gptq - int4 - vllm - marlin pipeline_tag: text-generation model-index: - name: Qwopus3.5-9B-v3-HLWQ-v7-GPTQ results: - task: type: text-generation name: Code Generation dataset: name: HumanEval type: openai_humaneval metrics: - name: pass@1 type: pass@1 value: 67.07 verified: true --- > [!IMPORTANT] > **Naming notice (2026-04-10).** The "HLWQ" technique used in this model is being rebranded to **HLWQ (Hadamard-Lloyd Weight Quantization)**. The change is only the name; the algorithm and the weights in this repository are unchanged. > > The rebrand resolves a name collision with an unrelated, earlier KV cache quantization method also named HLWQ ([Han et al., arXiv:2502.02617, 2025](https://arxiv.org/abs/2502.02617)). HLWQ addresses **weight** quantization with a **deterministic Walsh-Hadamard rotation** and Lloyd-Max scalar codebook; Han et al.'s HLWQ addresses **KV cache** quantization with a **random polar rotation**. The two methods are technically distinct. > > Existing loaders that load this repository by ID continue to work without changes. Future model uploads will use the HLWQ name. > > Reference paper for this technique: [arXiv:2603.29078](https://arxiv.org/abs/2603.29078) (v2 in preparation; v1 still uses the old name). # 🧊 Qwopus3.5-9B-v3 HLWQ v7 GPTQ ## 🎯 INT4 that BEATS BF16 on HumanEval | Metric | INT4 (ours) | BF16 (base) | Delta | |--------|:-----------:|:-----------:|:-----:| | 🎯 **HumanEval** | **67.07%** | 66.87% | **+0.20pp** | | 📊 WikiText-2 PPL | 9.95 | — | — | | 📦 Download | **8.7 GB** | 19.3 GB | **-55%** | | ⚡ BPW | 4.475 | 16 | **3.6x smaller** | | 🚀 Kernel | Marlin | — | Native vLLM | ## 📊 Benchmarks ![HumanEval Comparison](humaneval_comparison.png) ![group_size Impact](groupsize_impact.png) ![Quality vs Size](quality_vs_size.png) ## 🔬 How is this possible? GPTQ calibrated quantization with `group_size=64` and **FOEM** (First-Order Error Minimization) acts as a **regularizer**, reducing high-frequency noise in weights while preserving the essential signal. Combined with 2x finer granularity (gs64 vs standard gs128), this narrowly matches BF16 quality on HumanEval standard mode at this specific 9B scale. The +0.20pp edge on 164 samples is within the noise floor of the benchmark — read as "matches BF16" rather than a strict improvement. **Key discovery:** `group_size=64` was the dominant factor: - gs128 → gs64 jumped from 61.59% to 67.07% (**+5.48pp**) - FOEM contributed +0.61pp on top of vanilla GPTQ - Same Marlin kernel, same inference speed, only +0.1 GB larger ## 🚀 Quick Start ### vLLM (recommended) ```python from vllm import LLM, SamplingParams model = LLM( "caiovicentino1/Qwopus3.5-9B-v3-HLWQ-v7-GPTQ", trust_remote_code=True, language_model_only=True, ) output = model.generate("Write a Python function to sort a list:", SamplingParams(max_tokens=256)) print(output[0].outputs[0].text) ``` ### vLLM Server ```bash vllm serve caiovicentino1/Qwopus3.5-9B-v3-HLWQ-v7-GPTQ \ --trust-remote-code --language-model-only \ --max-model-len 16384 ``` ### GPTQModel ```python from gptqmodel import GPTQModel model = GPTQModel.from_quantized( "caiovicentino1/Qwopus3.5-9B-v3-HLWQ-v7-GPTQ", trust_remote_code=True, ) ``` ## 🔧 Quantization Config ```python # Universal config — works for any model from gptqmodel import GPTQModel from gptqmodel.quantization import QuantizeConfig from gptqmodel.quantization.config import FOEMConfig quantize_config = QuantizeConfig( bits=4, group_size=64, # 2x finer than standard 128 sym=True, desc_act=True, foem=FOEMConfig( alpha=0.25, # GPTAQ adaptive term beta=0.2, # FOEM first-order error correction device="auto" ) ) ``` - **Quantizer:** GPTQModel v6.0.3 - **Calibration:** 512 samples from `neuralmagic/LLM_compression_calibration` - **Kernel:** Marlin (native vLLM, zero overhead) ## 📈 Benchmark Evolution | # | Method | HumanEval | Notes | |---|--------|-----------|-------| | 1 | Naive INT4 (RTN) | 55.49% | Round-to-nearest, no calibration | | 2 | GPTQ gs128 desc_act | 60.98% | Calibrated, standard groups | | 3 | FOEM gs128 | 61.59% | +FOEM error correction | | 4 | FOEM gs128 (Arien0) | 62.80% | Different calibration data | | 5 | BF16 Base | 66.87% | Original unquantized | | **6** | **HLWQ v7 gs64+FOEM** | **67.07%** | **BEATS BF16** | ## 📖 Technical Details | Parameter | Value | |-----------|-------| | Base Model | Jackrong/Qwopus3.5-9B-v3 | | Architecture | Qwen3.5 (24 linear_attn + 8 full_attn) | | Hidden Size | 4096 | | Layers | 32 | | Bits | 4 | | Group Size | 64 | | Symmetric | Yes | | desc_act | Yes | | FOEM alpha | 0.25 | | FOEM beta | 0.2 | | BPW | 4.475 | | Format | GPTQ v1 (Marlin compatible) | ## 📖 Citation ```bibtex @article{vicentino2026polarquant, title={HLWQ: Polar Coordinate Quantization for Efficient LLM Inference}, author={Vicentino, Caio}, journal={arXiv preprint arXiv:2603.29078}, year={2026} } ``` ## 🔗 Links - 📜 [HLWQ Paper (arXiv:2603.29078)](https://arxiv.org/abs/2603.29078) - 💻 [GitHub](https://github.com/caiovicentino1/polarquant) - 📦 [PyPI: pip install polarquant](https://pypi.org/project/polarquant/) - 🌍 [Base Model](https://huggingface.co/Jackrong/Qwopus3.5-9B-v3) - 🏆 [HLWQ Q5 variant](https://huggingface.co/caiovicentino1/Qwopus3.5-9B-v3-HLWQ-Q5) ## 🙏 Acknowledgements - **Arien0** for the HumanEval benchmarks that drove this optimization - **GPTQModel team** for FOEM implementation - **vLLM team** for Marlin kernel support