--- base_model: - RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic library_name: speculators license: apache-2.0 tags: - speculative-decoding - dflash - speculators --- # NVIDIA-Nemotron-3-Ultra-550B-A55B-speculator.dflash This is a DFlash speculator model for [RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic](https://huggingface.co/RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic). ## Training Details This model was trained using the [Speculators](https://github.com/vllm-project/speculators) library on a subset of [Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered](https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered) and the `train_sft` split of [HuggingFaceH4/ultrachat_200k](https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k). Responses were regenerated by NVIDIA-Nemotron-3-Ultra-550B-A55B. Training compute for this model was generously provided by [Lambda](https://lambda.ai/), a leading cloud platform for AI training and inference.
Commands Using the [Speculators](https://github.com/vllm-project/speculators) library and the helper scripts provided in the repo. ### Prepare data ```bash # In virtual environment with speculators installed python scripts/prepare_data.py \ --model RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic \ --data ./regenerated_data.jsonl \ --output ./output \ --seq-length 8192 ``` ### Launch vLLM ```bash # In (separate) virtual environment with vllm installed CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 vllm_venv/bin/python scripts/launch_vllm.py \ RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic \ --target-layer-ids 7 47 88 \ -- --port 8000 \ --gpu-memory-utilization 0.9 \ --disable-uvicorn-access-log \ --tensor-parallel-size 8 ``` ### Launch training Must be run once vLLM has finished launching and is running in the background. ```bash # In virtual environment with speculators installed CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 torchrun \ --standalone \ --nproc_per_node 8 \ scripts/train.py \ --verifier-name-or-path RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic \ --speculator-type dflash \ --num-layers 5 \ --data-path ./output \ --vllm-endpoint http://localhost:8000/v1 \ --save-path ./output/checkpoints \ --epochs 5 \ --lr 0.0006 \ --total-seq-len 8192 \ --on-missing generate \ --on-generate delete \ --seed 42 \ --log-freq 10 \ --draft-vocab-size 32000 \ --draft-arch llama \ --target-layer-ids 7 47 88 \ --draft-hidden-act silu \ --scheduler-type cosine \ --max-anchors 3072 \ --prefetch-factor 2 \ --num-workers 16 \ --sliding-window-indices 0 1 2 3 4 \ --optimizer muon ```
## Model Specifications | | | |---|---| | **Base Model** | RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic | | **Chat Template** | RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic (use `/chat/completions` endpoint) | | **Format** | Safetensors | | **License** | Apache 2.0 | | **Validation Hardware** | Nvidia B200 | ## Deployment ```bash # Deploy with speculative decoding vllm serve RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic \ --tensor-parallel-size 8 \ --max-model-len 16384 \ --speculative-config '{ "model": "RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-speculator.dflash", "num_speculative_tokens": 7, "method": "dflash" }' ``` ## Acceptance Rate Per-position token acceptance rates across datasets: | Dataset | Pos 0 | Pos 1 | Pos 2 | Pos 3 | Pos 4 | Pos 5 | Pos 6 | Pos 7 | Avg. Length | |---------|-------|-------|-------|-------|-------|-------|-------|-------|------------| | HumanEval | 76.3% | 53.9% | 37.3% | 25.8% | 18.8% | 14.0% | 10.5% | 7.5% | 3.44 | | math_reasoning | 88.4% | 74.8% | 61.3% | 50.7% | 41.0% | 32.9% | 25.4% | 19.3% | 4.94 | | qa | 64.9% | 37.9% | 21.5% | 12.8% | 7.7% | 4.5% | 2.7% | 1.5% | 2.53 | | question | 68.4% | 41.3% | 24.1% | 15.3% | 10.4% | 7.1% | 4.8% | 3.0% | 2.74 | | rag | 75.3% | 52.0% | 34.6% | 23.7% | 16.2% | 11.1% | 7.4% | 4.3% | 3.25 | | summarization | 71.6% | 44.5% | 25.6% | 14.9% | 8.3% | 4.7% | 2.3% | 1.1% | 2.73 | | tool_call | 73.3% | 48.4% | 29.2% | 18.5% | 12.3% | 8.0% | 5.2% | 3.5% | 2.98 | | translation | 64.8% | 41.0% | 24.9% | 14.6% | 8.3% | 4.7% | 2.7% | 1.4% | 2.62 | | writing | 68.2% | 41.0% | 24.1% | 15.1% | 10.2% | 7.0% | 4.7% | 3.0% | 2.73 | ## Performance Eval We used **NVIDIA B200** for performance evaluation. ![Performance Evaluation - HumanEval Latency Speedup](speedup_HumanEval_latency.png) ## Long Context Benchmarking Per-position token acceptance rates across long context datasets: | Dataset | Context Length | Pos 0 | Pos 1 | Pos 2 | Pos 3 | Pos 4 | Pos 5 | Pos 6 | Avg. Length | |---------|---------------|-------|-------|-------|-------|-------|-------|-------|------------| | Academic | 2000 | 70.2% | 44.4% | 27.6% | 13.9% | 8.7% | 5.0% | 3.6% | 2.73 | | Academic | 4000 | 71.2% | 39.3% | 18.6% | 10.2% | 4.4% | 1.7% | 1.2% | 2.47 | | Academic | 6000 | 63.3% | 35.8% | 15.8% | 6.0% | 2.0% | 1.1% | 1.1% | 2.25 | | Academic | 8000 | 68.2% | 38.3% | 19.5% | 7.8% | 3.3% | 0.7% | 0.2% | 2.38 | | Academic | 10000 | 65.6% | 35.4% | 15.1% | 7.2% | 4.0% | 1.6% | 1.1% | 2.30 | | Academic | 12000 | 67.1% | 34.8% | 14.7% | 6.0% | 3.0% | 1.4% | 0.9% | 2.28 | | Academic | 14000 | 66.3% | 34.3% | 15.6% | 6.7% | 2.8% | 1.2% | 0.9% | 2.28 | | Agent history QA | 2000 | 69.5% | 42.4% | 24.1% | 14.0% | 9.9% | 3.9% | 2.4% | 2.66 | | Agent history QA | 4000 | 64.5% | 37.9% | 21.6% | 11.3% | 7.6% | 2.9% | 1.6% | 2.47 | | Agent history QA | 6000 | 70.3% | 40.9% | 23.9% | 14.7% | 8.4% | 2.2% | 1.2% | 2.62 | | Agent history QA | 8000 | 69.9% | 41.2% | 25.2% | 14.8% | 8.7% | 2.7% | 1.2% | 2.64 | | Agent history QA | 10000 | 67.3% | 42.0% | 21.0% | 10.1% | 6.0% | 1.6% | 0.0% | 2.48 | | Agent history QA | 12000 | 67.5% | 43.6% | 26.8% | 14.5% | 7.0% | 2.5% | 1.0% | 2.63 | | Agent history QA | 14000 | 68.0% | 42.3% | 21.7% | 11.3% | 6.6% | 1.4% | 0.6% | 2.52 | | Code repo QA | 2000 | 67.2% | 38.1% | 21.1% | 10.6% | 7.2% | 4.0% | 1.7% | 2.50 | | Code repo QA | 4000 | 65.7% | 34.1% | 18.1% | 7.3% | 1.8% | 0.9% | 0.2% | 2.28 | | Code repo QA | 6000 | 63.9% | 33.6% | 12.9% | 6.3% | 3.6% | 0.7% | 0.4% | 2.21 | | Code repo QA | 8000 | 60.2% | 30.1% | 11.7% | 3.6% | 0.9% | 0.6% | 0.4% | 2.07 | | Code repo QA | 10000 | 60.2% | 31.1% | 12.4% | 4.2% | 1.5% | 0.2% | 0.0% | 2.10 | | Code repo QA | 12000 | 59.0% | 31.5% | 10.9% | 3.2% | 0.9% | 0.4% | 0.2% | 2.06 | | Code repo QA | 14000 | 60.8% | 29.2% | 11.8% | 4.4% | 2.0% | 0.3% | 0.0% | 2.09 | | Detective | 2000 | 65.4% | 34.2% | 15.2% | 6.8% | 2.5% | 1.5% | 0.8% | 2.26 | | Detective | 4000 | 61.8% | 32.9% | 11.3% | 3.8% | 1.7% | 0.5% | 0.0% | 2.12 | | Detective | 6000 | 62.3% | 29.2% | 11.8% | 4.1% | 1.8% | 0.0% | 0.0% | 2.09 | | Detective | 8000 | 62.8% | 30.7% | 11.2% | 4.8% | 2.2% | 0.5% | 0.2% | 2.12 | | Detective | 10000 | 63.1% | 29.9% | 9.8% | 3.9% | 1.0% | 0.3% | 0.3% | 2.08 | | Detective | 12000 | 53.7% | 26.4% | 10.4% | 3.4% | 1.1% | 0.3% | 0.0% | 1.95 | | Detective | 14000 | 52.5% | 26.6% | 9.6% | 2.0% | 1.1% | 0.3% | 0.0% | 1.92 | | Dialogue history QA | 2000 | 75.0% | 50.0% | 29.2% | 17.0% | 8.3% | 5.6% | 3.8% | 2.89 | | Dialogue history QA | 4000 | 72.6% | 41.2% | 22.6% | 11.2% | 6.5% | 3.8% | 1.2% | 2.59 | | Dialogue history QA | 6000 | 73.5% | 41.3% | 19.1% | 9.5% | 5.0% | 3.1% | 1.2% | 2.53 | | Dialogue history QA | 8000 | 69.1% | 34.6% | 13.6% | 3.4% | 1.0% | 0.3% | 0.0% | 2.22 | | Dialogue history QA | 10000 | 78.3% | 44.0% | 24.0% | 13.0% | 5.3% | 3.0% | 2.3% | 2.70 | | Dialogue history QA | 12000 | 65.7% | 31.7% | 16.5% | 6.3% | 3.6% | 1.5% | 0.5% | 2.26 | | Dialogue history QA | 14000 | 66.9% | 37.2% | 17.7% | 9.8% | 6.0% | 3.2% | 1.5% | 2.42 | | Event ordering | 2000 | 65.8% | 38.3% | 19.2% | 11.0% | 5.5% | 2.5% | 0.4% | 2.43 | | Event ordering | 4000 | 60.6% | 31.0% | 14.9% | 6.4% | 2.4% | 0.8% | 0.7% | 2.17 | | Event ordering | 6000 | 60.8% | 32.8% | 15.0% | 6.3% | 2.0% | 0.7% | 0.2% | 2.18 | | Event ordering | 8000 | 60.9% | 31.1% | 14.7% | 5.4% | 1.7% | 0.8% | 0.3% | 2.15 | | Event ordering | 10000 | 58.8% | 27.8% | 13.1% | 5.4% | 1.5% | 0.7% | 0.0% | 2.07 | | Event ordering | 12000 | 61.9% | 31.4% | 14.7% | 4.9% | 1.7% | 0.8% | 0.2% | 2.16 | | Event ordering | 14000 | 62.6% | 33.6% | 15.6% | 6.4% | 2.2% | 1.2% | 0.5% | 2.22 | | Financial | 2000 | 70.0% | 44.7% | 28.4% | 17.8% | 12.9% | 8.2% | 4.2% | 2.86 | | Financial | 4000 | 66.5% | 36.9% | 19.0% | 11.7% | 6.7% | 4.2% | 1.5% | 2.47 | | Financial | 6000 | 65.8% | 39.0% | 20.2% | 9.8% | 6.7% | 3.5% | 1.0% | 2.46 | | Financial | 8000 | 64.3% | 37.5% | 19.8% | 11.1% | 6.4% | 3.5% | 1.7% | 2.44 | | Financial | 10000 | 64.6% | 37.1% | 19.1% | 9.4% | 5.2% | 2.8% | 1.3% | 2.40 | | Financial | 12000 | 64.9% | 39.1% | 19.7% | 10.9% | 7.1% | 3.3% | 1.3% | 2.46 | | Financial | 14000 | 67.1% | 41.0% | 21.9% | 11.8% | 6.8% | 3.6% | 2.2% | 2.54 | | Governmental | 2000 | 73.0% | 49.8% | 30.9% | 16.2% | 10.1% | 6.5% | 3.6% | 2.90 | | Governmental | 4000 | 72.2% | 46.6% | 27.6% | 15.8% | 9.4% | 5.7% | 2.4% | 2.80 | | Governmental | 6000 | 70.4% | 44.4% | 24.7% | 12.3% | 7.0% | 3.5% | 1.0% | 2.63 | | Governmental | 8000 | 74.1% | 44.8% | 24.0% | 13.7% | 7.6% | 4.0% | 1.7% | 2.70 | | Governmental | 10000 | 71.1% | 43.8% | 22.6% | 11.4% | 5.7% | 3.7% | 2.2% | 2.60 | | Governmental | 12000 | 70.6% | 43.8% | 24.3% | 13.8% | 6.2% | 2.9% | 1.9% | 2.63 | | Governmental | 14000 | 70.8% | 39.6% | 23.0% | 11.6% | 6.0% | 3.4% | 1.6% | 2.56 | | Knowledge graph reasoning | 2000 | 76.5% | 43.1% | 15.7% | 9.8% | 3.9% | 0.0% | 0.0% | 2.49 | | Knowledge graph reasoning | 4000 | 76.9% | 43.6% | 15.4% | 7.7% | 5.1% | 0.0% | 0.0% | 2.49 | | Knowledge graph reasoning | 6000 | 78.2% | 36.4% | 9.1% | 5.5% | 1.8% | 0.0% | 0.0% | 2.31 | | Knowledge graph reasoning | 8000 | 69.6% | 30.4% | 12.5% | 8.9% | 5.4% | 0.0% | 0.0% | 2.27 | | Knowledge graph reasoning | 10000 | 71.9% | 31.2% | 3.1% | 0.0% | 0.0% | 0.0% | 0.0% | 2.06 | | Knowledge graph reasoning | 12000 | 70.2% | 31.6% | 14.0% | 7.0% | 1.8% | 0.0% | 0.0% | 2.25 | | Knowledge graph reasoning | 14000 | 59.6% | 28.1% | 12.3% | 8.8% | 3.5% | 0.0% | 0.0% | 2.12 | | Legal | 2000 | 74.7% | 48.4% | 29.4% | 17.5% | 10.3% | 4.7% | 2.9% | 2.88 | | Legal | 4000 | 70.3% | 42.2% | 23.1% | 12.7% | 5.6% | 2.6% | 1.6% | 2.58 | | Legal | 6000 | 71.6% | 43.1% | 19.7% | 11.1% | 5.6% | 2.4% | 1.0% | 2.54 | | Legal | 8000 | 71.4% | 39.2% | 16.0% | 5.9% | 3.5% | 1.7% | 0.9% | 2.39 | | Legal | 10000 | 68.7% | 39.5% | 17.4% | 10.1% | 5.7% | 2.5% | 1.0% | 2.45 | | Legal | 12000 | 68.5% | 40.8% | 19.1% | 7.5% | 3.4% | 1.3% | 0.8% | 2.41 | | Legal | 14000 | 64.5% | 33.2% | 16.8% | 7.1% | 3.2% | 1.2% | 0.9% | 2.27 | | Literary | 2000 | 66.9% | 36.0% | 16.8% | 9.4% | 4.7% | 2.9% | 1.8% | 2.38 | | Literary | 4000 | 66.1% | 34.6% | 15.8% | 8.0% | 3.5% | 1.5% | 0.9% | 2.30 | | Literary | 6000 | 60.4% | 31.2% | 13.2% | 6.3% | 1.5% | 0.7% | 0.5% | 2.14 | | Literary | 8000 | 61.1% | 30.1% | 13.4% | 5.8% | 2.0% | 0.7% | 0.5% | 2.14 | | Literary | 10000 | 62.9% | 31.5% | 12.4% | 5.4% | 1.7% | 0.7% | 0.3% | 2.15 | | Literary | 12000 | 63.1% | 32.5% | 12.5% | 3.9% | 1.3% | 0.7% | 0.2% | 2.14 | | Literary | 14000 | 60.5% | 31.3% | 13.4% | 4.5% | 2.5% | 1.0% | 0.5% | 2.14 | | Many-shot learning | 2000 | 68.2% | 38.2% | 18.9% | 8.2% | 4.8% | 1.0% | 1.0% | 2.40 | | Many-shot learning | 4000 | 63.2% | 30.5% | 13.6% | 5.4% | 3.0% | 1.4% | 1.1% | 2.18 | | Many-shot learning | 6000 | 65.0% | 27.6% | 6.7% | 1.8% | 1.1% | 0.2% | 0.2% | 2.02 | | Many-shot learning | 8000 | 64.6% | 39.2% | 23.5% | 9.7% | 4.0% | 1.9% | 0.9% | 2.44 | | Many-shot learning | 10000 | 64.5% | 36.5% | 10.8% | 4.5% | 1.7% | 0.9% | 0.0% | 2.19 | | Many-shot learning | 12000 | 70.5% | 43.3% | 19.6% | 5.9% | 3.0% | 0.4% | 0.0% | 2.43 | | Many-shot learning | 14000 | 77.8% | 46.4% | 15.2% | 7.2% | 3.5% | 0.8% | 0.0% | 2.51 | | Multi-news | 2000 | 73.5% | 47.3% | 31.4% | 18.3% | 10.6% | 7.0% | 3.4% | 2.92 | | Multi-news | 4000 | 74.4% | 44.2% | 23.5% | 13.4% | 7.1% | 4.4% | 2.1% | 2.69 | | Multi-news | 6000 | 69.2% | 43.8% | 23.7% | 11.2% | 6.1% | 4.1% | 2.2% | 2.60 | | Multi-news | 8000 | 66.1% | 36.0% | 17.9% | 9.4% | 4.4% | 2.6% | 1.1% | 2.37 | | Multi-news | 10000 | 66.6% | 39.6% | 21.3% | 11.1% | 5.5% | 3.3% | 2.1% | 2.50 | | Multi-news | 12000 | 67.1% | 39.9% | 18.5% | 9.8% | 5.8% | 3.3% | 1.5% | 2.46 | | Multi-news | 14000 | 63.7% | 39.3% | 20.2% | 9.2% | 4.3% | 2.4% | 1.3% | 2.40 | | New language translation | 2000 | 67.0% | 36.0% | 15.9% | 6.9% | 3.2% | 1.5% | 1.0% | 2.32 | | New language translation | 4000 | 65.4% | 35.4% | 14.5% | 4.1% | 2.0% | 1.5% | 1.5% | 2.24 | | New language translation | 6000 | 64.0% | 34.6% | 11.6% | 3.9% | 1.1% | 0.5% | 0.0% | 2.16 | | New language translation | 8000 | 65.1% | 31.6% | 11.4% | 2.1% | 0.6% | 0.4% | 0.0% | 2.11 | | New language translation | 10000 | 73.9% | 37.1% | 12.7% | 3.0% | 1.5% | 0.4% | 0.0% | 2.29 | | New language translation | 12000 | 71.7% | 31.7% | 13.0% | 3.7% | 1.2% | 0.6% | 0.0% | 2.22 | | New language translation | 14000 | 69.2% | 34.6% | 12.5% | 3.4% | 1.0% | 0.4% | 0.0% | 2.21 | | Table QA | 2000 | 71.2% | 42.2% | 23.8% | 13.7% | 5.9% | 2.4% | 1.7% | 2.61 | | Table QA | 4000 | 69.5% | 39.6% | 21.8% | 11.4% | 6.4% | 3.0% | 1.3% | 2.53 | | Table QA | 6000 | 68.6% | 41.5% | 20.8% | 10.3% | 5.5% | 2.3% | 1.3% | 2.50 | | Table QA | 8000 | 66.8% | 41.7% | 21.2% | 12.0% | 7.0% | 4.4% | 2.0% | 2.55 | | Table QA | 10000 | 65.1% | 38.0% | 19.5% | 9.3% | 6.2% | 3.5% | 1.4% | 2.43 | | Table QA | 12000 | 65.2% | 35.4% | 17.4% | 9.1% | 6.1% | 3.6% | 1.0% | 2.38 | | Table QA | 14000 | 67.1% | 37.4% | 17.4% | 9.0% | 5.1% | 2.5% | 1.0% | 2.39 | | User guide QA | 2000 | 66.8% | 43.1% | 23.5% | 9.7% | 4.0% | 1.0% | 0.4% | 2.48 | | User guide QA | 4000 | 65.3% | 37.2% | 15.7% | 6.4% | 3.0% | 1.3% | 0.7% | 2.30 | | User guide QA | 6000 | 61.5% | 28.4% | 9.8% | 4.2% | 2.1% | 0.2% | 0.2% | 2.06 | | User guide QA | 8000 | 65.6% | 34.0% | 15.3% | 4.6% | 1.2% | 0.4% | 0.2% | 2.21 | | User guide QA | 10000 | 60.5% | 30.3% | 10.3% | 3.2% | 0.6% | 0.2% | 0.2% | 2.05 | | User guide QA | 12000 | 63.1% | 31.8% | 13.4% | 5.9% | 2.0% | 0.7% | 0.2% | 2.17 | | User guide QA | 14000 | 62.4% | 30.3% | 11.1% | 4.4% | 1.7% | 0.9% | 0.5% | 2.11 | ## References **Paper**: [DFlash: Block Diffusion for Flash Speculative Decoding](https://arxiv.org/abs/2602.06036)