ZixiQi commited on
Commit
f0a1fea
·
verified ·
1 Parent(s): e6b2bbf

Add acceptance-rate Performance section (MT-Bench + SPEED-Bench, per-position)

Browse files
Files changed (1) hide show
  1. README.md +9 -0
README.md CHANGED
@@ -16,3 +16,12 @@ tags:
16
  # MiniMax-M3-EAGLE3-GQA-NVFP4
17
 
18
  W4A4 NVFP4 MLP-quantized version of [`Inferact/MiniMax-M3-EAGLE3-GQA`](https://huggingface.co/Inferact/MiniMax-M3-EAGLE3-GQA).
 
 
 
 
 
 
 
 
 
 
16
  # MiniMax-M3-EAGLE3-GQA-NVFP4
17
 
18
  W4A4 NVFP4 MLP-quantized version of [`Inferact/MiniMax-M3-EAGLE3-GQA`](https://huggingface.co/Inferact/MiniMax-M3-EAGLE3-GQA).
19
+
20
+ ## Performance
21
+
22
+ Mean accepted length and draft accept rate measured end-to-end against `MiniMaxAI/MiniMax-M3-MXFP8` served with vLLM at `tensor-parallel-size=4`, `num_speculative_tokens=3`, greedy sampling (`temperature=0`, `top_p=1.0`), `max-concurrency=16`.
23
+
24
+ | Dataset | n | Mean accepted length | Draft accept rate | Per-position accept rate (pos 1 / 2 / 3) |
25
+ |---|---:|---:|---:|---:|
26
+ | MT-Bench | 64 | 2.663 | 55.42% | 0.742 / 0.534 / 0.386 |
27
+ | SPEED-Bench (qualitative) | 64 | 2.633 | 54.43% | 0.736 / 0.526 / 0.371 |