--- license: mit library_name: transformers base_model: Inferact/MiniMax-M3-EAGLE3-GQA pipeline_tag: text-generation tags: - eagle3 - speculative-decoding - draft-model - gqa - vllm - nvfp4 - quantized --- # MiniMax-M3-EAGLE3-GQA-NVFP4 W4A4 NVFP4 MLP-quantized version of [`Inferact/MiniMax-M3-EAGLE3-GQA`](https://huggingface.co/Inferact/MiniMax-M3-EAGLE3-GQA). ## Performance Mean accepted length and draft accept rate measured end-to-end against `MiniMaxAI/MiniMax-M3-MXFP8` served with vLLM at `tensor-parallel-size=4`, `num_speculative_tokens=3`, greedy sampling (`temperature=0`, `top_p=1.0`), `max-concurrency=16`. | Dataset | n | Mean accepted length | Draft accept rate | Per-position accept rate (pos 1 / 2 / 3) | |---|---:|---:|---:|---:| | MT-Bench | 64 | 2.663 | 55.42% | 0.742 / 0.534 / 0.386 | | SPEED-Bench (qualitative) | 64 | 2.633 | 54.43% | 0.736 / 0.526 / 0.371 |