ZixiQi's picture
Add acceptance-rate Performance section (MT-Bench + SPEED-Bench, per-position)
f0a1fea verified
|
Raw
History Blame Contribute Delete
904 Bytes
---
license: mit
library_name: transformers
base_model: Inferact/MiniMax-M3-EAGLE3-GQA
pipeline_tag: text-generation
tags:
- eagle3
- speculative-decoding
- draft-model
- gqa
- vllm
- nvfp4
- quantized
---
# MiniMax-M3-EAGLE3-GQA-NVFP4
W4A4 NVFP4 MLP-quantized version of [`Inferact/MiniMax-M3-EAGLE3-GQA`](https://huggingface.co/Inferact/MiniMax-M3-EAGLE3-GQA).
## Performance
Mean accepted length and draft accept rate measured end-to-end against `MiniMaxAI/MiniMax-M3-MXFP8` served with vLLM at `tensor-parallel-size=4`, `num_speculative_tokens=3`, greedy sampling (`temperature=0`, `top_p=1.0`), `max-concurrency=16`.
| Dataset | n | Mean accepted length | Draft accept rate | Per-position accept rate (pos 1 / 2 / 3) |
|---|---:|---:|---:|---:|
| MT-Bench | 64 | 2.663 | 55.42% | 0.742 / 0.534 / 0.386 |
| SPEED-Bench (qualitative) | 64 | 2.633 | 54.43% | 0.736 / 0.526 / 0.371 |