Qwopus3.5
Collection
9 items • Updated
How to use mlx-works/Qwopus3.5-9B-Coder-oQ4-mtp with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Qwopus3.5-9B-Coder-oQ4-mtp mlx-works/Qwopus3.5-9B-Coder-oQ4-mtp
This model was quantized using oQ (oMLX v0.5.1) mixed-precision quantization.
Note: Results are for reference only and may vary depending on hardware, software configuration, and workload.
| Test | TTFT(ms) | TPOT(ms) | pp TPS | tg TPS | E2E(s) | Throughput | Peak Mem |
|---|---|---|---|---|---|---|---|
| pp1024/tg128 | 1735.2 | 26.43 | 590.1 tok/s | 38.1 tok/s | 5.092 | 226.2 tok/s | 5.96 GB |
| pp4096/tg128 | 6142.9 | 28.49 | 666.8 tok/s | 35.4 tok/s | 9.761 | 432.7 tok/s | 6.57 GB |
| Batch | tg TPS | Speedup | pp TPS | pp TPS/req | TTFT(ms) | E2E(s) |
|---|---|---|---|---|---|---|
| 1x | 38.1 tok/s | 1.00x | 590.1 tok/s | 590.1 tok/s | 1735.2 | 5.092 |
| 2x | 48.4 tok/s | 1.27x | 606.6 tok/s | 303.3 tok/s | 3376.0 | 8.662 |
| 4x | 53.2 tok/s | 1.40x | 578.0 tok/s | 144.5 tok/s | 6889.9 | 16.712 |
Note: Each benchmark round tests only 30 questions. Results are for reference only.
| Benchmark | Accuracy | Correct | Total | Time(s) | Think |
|---|---|---|---|---|---|
| MMLU | 80.0% | 24 | 30 | 57.1 | No |
| TRUTHFULQA | 86.7% | 26 | 30 | 24.4 | No |
| GSM8K | 83.3% | 25 | 30 | 355.2 | No |
| MATHQA | 40.0% | 12 | 30 | 25.6 | No |
| HUMANEVAL | 80.0% | 24 | 30 | 157.7 | No |
4-bit