XReyRobert's picture
Add personal project notice
ca73e3d verified
|
Raw
History Blame
6.59 kB
metadata
base_model: Jackrong/Qwopus3.6-27B-v2
library_name: transformers
tags:
  - qwen3.6
  - qwopus
  - gptq
  - gptq-pro
  - marlin
  - vllm
  - int4
  - quantized
license: other

Qwopus3.6-27B-v2 GPTQ-Pro v1 RTX 3090 benchmark

Qwopus3.6-27B-v2 GPTQ-Pro v1

This is the first public-ready GPTQ-Pro 4-bit release of Jackrong/Qwopus3.6-27B-v2, built to make this excellent Qwopus/Qwen3.6 model practical to run in vLLM with GPTQ-Marlin kernels and long-context inference.

The goal is simple: preserve as much of the original model's character and capability as possible while making it efficient enough for real local/homelab serving, including RTX 3090-class deployments.

This is not a new fine-tune. It is a quantized derivative of the original Qwopus3.6-27B-v2 model.

Source and credits

Source model:

Quantization methodology and reference recipe:

Thanks to Jackrong for the original Qwopus3.6 model, and to groxaxo for GPTQ-Pro and the Qwen3.6 GPTQ-Pro recipe this quantization was aligned with.

Quantization recipe

Setting Value
Method GPTQ-Pro / GPTQModel
Bits 4
Group size 128
Symmetric quantization true
Desc act false
True sequential true
Calibration dataset WikiText-2 raw train
Calibration samples 256
Sequence length 2048
MSE 2.0
Damp percent 0.05
Damp auto increment 0.01
FOEM alpha 0.25
FOEM beta 0.2
Batch size 1

Preserved modules include vision, lm_head, embeddings, norms, and MTP modules. Post-publication validation showed the internal ns256-v2 artifact preserves MTP metadata but does not include actual mtp.* tensors in model.safetensors.index.json, so this release should be treated as non-MTP for vLLM speculative decoding.

Post-save compatibility patch:

  • pad_token_id=248055
  • tokenizer class patched to Qwen2TokenizerFast when needed for vLLM compatibility

Intended serving setup

This checkpoint is intended for local/homelab vLLM serving on RTX 3090-class hardware.

Recommended vLLM options:

vllm serve XReyRobert/Qwopus3.6-27B-v2-GPTQ-Pro-v1 \
  --served-model-name qwopus3.6-27b-v2-gptq-pro-v1 \
  --language-model-only \
  --dtype float16 \
  --quantization gptq_marlin \
  --disable-custom-all-reduce \
  --tensor-parallel-size 1 \
  --max-model-len 131072 \
  --max-num-seqs 1 \
  --kv-cache-dtype fp8_e5m2 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder \
  --enable-prefix-caching \
  --max-cudagraph-capture-size 32 \
  --gpu-memory-utilization 0.95 \
  --trust-remote-code

Reasoning / thinking mode

This model preserves Qwen3-style reasoning behavior. The Hermes validation below was run with thinking enabled.

MTP / speculative decoding status

This v1 artifact should be considered text-only and non-MTP for vLLM speculative decoding as published. config.json advertises mtp_num_hidden_layers=1, but the weight index does not contain source mtp.* tensors. Enabling vLLM MTP against this unpatched artifact produced essentially zero accepted draft tokens and poor throughput.

A separate experimental follow-up artifact restores real MTP tensors and quantizes the large MTP linears:

XReyRobert/Qwopus3.6-27B-v2-MTP-GPTQ-Pro-v1

That MTP-GPTQ artifact works and reaches good draft acceptance, but it was still slower than this non-MTP baseline on a single RTX 3090. For practical 100k-131k serving on 1x RTX 3090, this non-MTP artifact remains the preferred choice.

RTX 3090 validation status

This checkpoint was validated locally on an RTX 3090 24GB with vLLM, max_model_len=131072, kv_cache_dtype=fp8_e5m2, prefix caching enabled, and Hermes running with thinking enabled.

Observed Hermes/vLLM workload metrics:

Metric Observed value Notes
Requests observed 15 Hermes session calls
vLLM request success count 15/15 No vLLM errors observed during the sample
Average prompt size 33,172 tokens Real multi-turn Hermes workload
Average output size 322 tokens Real Hermes responses
Average time to first token 5.70s Prometheus TTFT summary
Average end-to-end request latency 13.07s Includes prefill, decode, and serving overhead
Average time per output token 0.0230s/token vLLM TPOT summary
Decode throughput from TPOT about 43.5 tok/s Decode-only estimate
Prefix cache hit ratio 83.2% cumulative vLLM prefix-cache counters
Live 60s prompt throughput about 1,917 prompt tok/s Aggregate observed window
Live 60s generation throughput about 19.1 generated tok/s Aggregate over full window, including prefill and idle mix
Live 60s prefix-cache hit ratio 78.9% Delta over the observed window

These are practical Hermes session metrics, not a synthetic benchmark. They are useful for RTX 3090-class homelab serving expectations, especially multi-turn long-context usage with prefix caching.

Compatibility notes

This artifact was built and validated for text-only vLLM serving without speculative decoding. Do not enable MTP on this artifact as published; the mtp.* tensors are absent from the weight index. Vision-related modules were not validated for vision use in this release.

Limitations

  • Experimental quantization.
  • MTP/speculative decoding is not supported by this published artifact because mtp.* tensors are missing.
  • Quality has not yet been benchmarked against the BF16 source model.
  • RTX 3090 metrics above are observed Hermes workload numbers, not a controlled benchmark suite.
  • Use at your own risk, especially for long-context or tool-calling workflows.

References

Personal project notice

This repository is a personal research project. The model, benchmarks, opinions, and documentation are my own and are not affiliated with, sponsored by, or endorsed by my employer or any organization I am associated with.