NVIDIA-Nemotron-3-Ultra-550B-A55B-speculator.dflash
This is a DFlash speculator model for RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic.
Training Details
This model was trained using the Speculators library on a subset of Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered and the train_sft split of HuggingFaceH4/ultrachat_200k. Responses were regenerated by NVIDIA-Nemotron-3-Ultra-550B-A55B. Training compute for this model was generously provided by Lambda, a leading cloud platform for AI training and inference.
Commands
Using the Speculators library and the helper scripts provided in the repo.
Prepare data
# In virtual environment with speculators installed
python scripts/prepare_data.py \
--model RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic \
--data ./regenerated_data.jsonl \
--output ./output \
--seq-length 8192
Launch vLLM
# In (separate) virtual environment with vllm installed
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 vllm_venv/bin/python scripts/launch_vllm.py \
RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic \
--target-layer-ids 7 47 88 \
-- --port 8000 \
--gpu-memory-utilization 0.9 \
--disable-uvicorn-access-log \
--tensor-parallel-size 8
Launch training
Must be run once vLLM has finished launching and is running in the background.
# In virtual environment with speculators installed
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 torchrun \
--standalone \
--nproc_per_node 8 \
scripts/train.py \
--verifier-name-or-path RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic \
--speculator-type dflash \
--num-layers 5 \
--data-path ./output \
--vllm-endpoint http://localhost:8000/v1 \
--save-path ./output/checkpoints \
--epochs 5 \
--lr 0.0006 \
--total-seq-len 8192 \
--on-missing generate \
--on-generate delete \
--seed 42 \
--log-freq 10 \
--draft-vocab-size 32000 \
--draft-arch llama \
--target-layer-ids 7 47 88 \
--draft-hidden-act silu \
--scheduler-type cosine \
--max-anchors 3072 \
--prefetch-factor 2 \
--num-workers 16 \
--sliding-window-indices 0 1 2 3 4 \
--optimizer muon
Model Specifications
| Base Model | RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic |
| Chat Template | RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic (use /chat/completions endpoint) |
| Format | Safetensors |
| License | Apache 2.0 |
| Validation Hardware | Nvidia B200 |
Deployment
# Deploy with speculative decoding
vllm serve RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic \
--tensor-parallel-size 8 \
--max-model-len 16384 \
--speculative-config '{
"model": "RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-speculator.dflash",
"num_speculative_tokens": 7,
"method": "dflash"
}'
Acceptance Rate
Per-position token acceptance rates across datasets:
| Dataset | Pos 0 | Pos 1 | Pos 2 | Pos 3 | Pos 4 | Pos 5 | Pos 6 | Pos 7 | Avg. Length |
|---|---|---|---|---|---|---|---|---|---|
| HumanEval | 76.3% | 53.9% | 37.3% | 25.8% | 18.8% | 14.0% | 10.5% | 7.5% | 3.44 |
| math_reasoning | 88.4% | 74.8% | 61.3% | 50.7% | 41.0% | 32.9% | 25.4% | 19.3% | 4.94 |
| qa | 64.9% | 37.9% | 21.5% | 12.8% | 7.7% | 4.5% | 2.7% | 1.5% | 2.53 |
| question | 68.4% | 41.3% | 24.1% | 15.3% | 10.4% | 7.1% | 4.8% | 3.0% | 2.74 |
| rag | 75.3% | 52.0% | 34.6% | 23.7% | 16.2% | 11.1% | 7.4% | 4.3% | 3.25 |
| summarization | 71.6% | 44.5% | 25.6% | 14.9% | 8.3% | 4.7% | 2.3% | 1.1% | 2.73 |
| tool_call | 73.3% | 48.4% | 29.2% | 18.5% | 12.3% | 8.0% | 5.2% | 3.5% | 2.98 |
| translation | 64.8% | 41.0% | 24.9% | 14.6% | 8.3% | 4.7% | 2.7% | 1.4% | 2.62 |
| writing | 68.2% | 41.0% | 24.1% | 15.1% | 10.2% | 7.0% | 4.7% | 3.0% | 2.73 |
Performance Eval
We used NVIDIA B200 for performance evaluation.
Long Context Benchmarking
Per-position token acceptance rates across long context datasets:
| Dataset | Context Length | Pos 0 | Pos 1 | Pos 2 | Pos 3 | Pos 4 | Pos 5 | Pos 6 | Avg. Length |
|---|---|---|---|---|---|---|---|---|---|
| Academic | 2000 | 70.2% | 44.4% | 27.6% | 13.9% | 8.7% | 5.0% | 3.6% | 2.73 |
| Academic | 4000 | 71.2% | 39.3% | 18.6% | 10.2% | 4.4% | 1.7% | 1.2% | 2.47 |
| Academic | 6000 | 63.3% | 35.8% | 15.8% | 6.0% | 2.0% | 1.1% | 1.1% | 2.25 |
| Academic | 8000 | 68.2% | 38.3% | 19.5% | 7.8% | 3.3% | 0.7% | 0.2% | 2.38 |
| Academic | 10000 | 65.6% | 35.4% | 15.1% | 7.2% | 4.0% | 1.6% | 1.1% | 2.30 |
| Academic | 12000 | 67.1% | 34.8% | 14.7% | 6.0% | 3.0% | 1.4% | 0.9% | 2.28 |
| Academic | 14000 | 66.3% | 34.3% | 15.6% | 6.7% | 2.8% | 1.2% | 0.9% | 2.28 |
| Agent history QA | 2000 | 69.5% | 42.4% | 24.1% | 14.0% | 9.9% | 3.9% | 2.4% | 2.66 |
| Agent history QA | 4000 | 64.5% | 37.9% | 21.6% | 11.3% | 7.6% | 2.9% | 1.6% | 2.47 |
| Agent history QA | 6000 | 70.3% | 40.9% | 23.9% | 14.7% | 8.4% | 2.2% | 1.2% | 2.62 |
| Agent history QA | 8000 | 69.9% | 41.2% | 25.2% | 14.8% | 8.7% | 2.7% | 1.2% | 2.64 |
| Agent history QA | 10000 | 67.3% | 42.0% | 21.0% | 10.1% | 6.0% | 1.6% | 0.0% | 2.48 |
| Agent history QA | 12000 | 67.5% | 43.6% | 26.8% | 14.5% | 7.0% | 2.5% | 1.0% | 2.63 |
| Agent history QA | 14000 | 68.0% | 42.3% | 21.7% | 11.3% | 6.6% | 1.4% | 0.6% | 2.52 |
| Code repo QA | 2000 | 67.2% | 38.1% | 21.1% | 10.6% | 7.2% | 4.0% | 1.7% | 2.50 |
| Code repo QA | 4000 | 65.7% | 34.1% | 18.1% | 7.3% | 1.8% | 0.9% | 0.2% | 2.28 |
| Code repo QA | 6000 | 63.9% | 33.6% | 12.9% | 6.3% | 3.6% | 0.7% | 0.4% | 2.21 |
| Code repo QA | 8000 | 60.2% | 30.1% | 11.7% | 3.6% | 0.9% | 0.6% | 0.4% | 2.07 |
| Code repo QA | 10000 | 60.2% | 31.1% | 12.4% | 4.2% | 1.5% | 0.2% | 0.0% | 2.10 |
| Code repo QA | 12000 | 59.0% | 31.5% | 10.9% | 3.2% | 0.9% | 0.4% | 0.2% | 2.06 |
| Code repo QA | 14000 | 60.8% | 29.2% | 11.8% | 4.4% | 2.0% | 0.3% | 0.0% | 2.09 |
| Detective | 2000 | 65.4% | 34.2% | 15.2% | 6.8% | 2.5% | 1.5% | 0.8% | 2.26 |
| Detective | 4000 | 61.8% | 32.9% | 11.3% | 3.8% | 1.7% | 0.5% | 0.0% | 2.12 |
| Detective | 6000 | 62.3% | 29.2% | 11.8% | 4.1% | 1.8% | 0.0% | 0.0% | 2.09 |
| Detective | 8000 | 62.8% | 30.7% | 11.2% | 4.8% | 2.2% | 0.5% | 0.2% | 2.12 |
| Detective | 10000 | 63.1% | 29.9% | 9.8% | 3.9% | 1.0% | 0.3% | 0.3% | 2.08 |
| Detective | 12000 | 53.7% | 26.4% | 10.4% | 3.4% | 1.1% | 0.3% | 0.0% | 1.95 |
| Detective | 14000 | 52.5% | 26.6% | 9.6% | 2.0% | 1.1% | 0.3% | 0.0% | 1.92 |
| Dialogue history QA | 2000 | 75.0% | 50.0% | 29.2% | 17.0% | 8.3% | 5.6% | 3.8% | 2.89 |
| Dialogue history QA | 4000 | 72.6% | 41.2% | 22.6% | 11.2% | 6.5% | 3.8% | 1.2% | 2.59 |
| Dialogue history QA | 6000 | 73.5% | 41.3% | 19.1% | 9.5% | 5.0% | 3.1% | 1.2% | 2.53 |
| Dialogue history QA | 8000 | 69.1% | 34.6% | 13.6% | 3.4% | 1.0% | 0.3% | 0.0% | 2.22 |
| Dialogue history QA | 10000 | 78.3% | 44.0% | 24.0% | 13.0% | 5.3% | 3.0% | 2.3% | 2.70 |
| Dialogue history QA | 12000 | 65.7% | 31.7% | 16.5% | 6.3% | 3.6% | 1.5% | 0.5% | 2.26 |
| Dialogue history QA | 14000 | 66.9% | 37.2% | 17.7% | 9.8% | 6.0% | 3.2% | 1.5% | 2.42 |
| Event ordering | 2000 | 65.8% | 38.3% | 19.2% | 11.0% | 5.5% | 2.5% | 0.4% | 2.43 |
| Event ordering | 4000 | 60.6% | 31.0% | 14.9% | 6.4% | 2.4% | 0.8% | 0.7% | 2.17 |
| Event ordering | 6000 | 60.8% | 32.8% | 15.0% | 6.3% | 2.0% | 0.7% | 0.2% | 2.18 |
| Event ordering | 8000 | 60.9% | 31.1% | 14.7% | 5.4% | 1.7% | 0.8% | 0.3% | 2.15 |
| Event ordering | 10000 | 58.8% | 27.8% | 13.1% | 5.4% | 1.5% | 0.7% | 0.0% | 2.07 |
| Event ordering | 12000 | 61.9% | 31.4% | 14.7% | 4.9% | 1.7% | 0.8% | 0.2% | 2.16 |
| Event ordering | 14000 | 62.6% | 33.6% | 15.6% | 6.4% | 2.2% | 1.2% | 0.5% | 2.22 |
| Financial | 2000 | 70.0% | 44.7% | 28.4% | 17.8% | 12.9% | 8.2% | 4.2% | 2.86 |
| Financial | 4000 | 66.5% | 36.9% | 19.0% | 11.7% | 6.7% | 4.2% | 1.5% | 2.47 |
| Financial | 6000 | 65.8% | 39.0% | 20.2% | 9.8% | 6.7% | 3.5% | 1.0% | 2.46 |
| Financial | 8000 | 64.3% | 37.5% | 19.8% | 11.1% | 6.4% | 3.5% | 1.7% | 2.44 |
| Financial | 10000 | 64.6% | 37.1% | 19.1% | 9.4% | 5.2% | 2.8% | 1.3% | 2.40 |
| Financial | 12000 | 64.9% | 39.1% | 19.7% | 10.9% | 7.1% | 3.3% | 1.3% | 2.46 |
| Financial | 14000 | 67.1% | 41.0% | 21.9% | 11.8% | 6.8% | 3.6% | 2.2% | 2.54 |
| Governmental | 2000 | 73.0% | 49.8% | 30.9% | 16.2% | 10.1% | 6.5% | 3.6% | 2.90 |
| Governmental | 4000 | 72.2% | 46.6% | 27.6% | 15.8% | 9.4% | 5.7% | 2.4% | 2.80 |
| Governmental | 6000 | 70.4% | 44.4% | 24.7% | 12.3% | 7.0% | 3.5% | 1.0% | 2.63 |
| Governmental | 8000 | 74.1% | 44.8% | 24.0% | 13.7% | 7.6% | 4.0% | 1.7% | 2.70 |
| Governmental | 10000 | 71.1% | 43.8% | 22.6% | 11.4% | 5.7% | 3.7% | 2.2% | 2.60 |
| Governmental | 12000 | 70.6% | 43.8% | 24.3% | 13.8% | 6.2% | 2.9% | 1.9% | 2.63 |
| Governmental | 14000 | 70.8% | 39.6% | 23.0% | 11.6% | 6.0% | 3.4% | 1.6% | 2.56 |
| Knowledge graph reasoning | 2000 | 76.5% | 43.1% | 15.7% | 9.8% | 3.9% | 0.0% | 0.0% | 2.49 |
| Knowledge graph reasoning | 4000 | 76.9% | 43.6% | 15.4% | 7.7% | 5.1% | 0.0% | 0.0% | 2.49 |
| Knowledge graph reasoning | 6000 | 78.2% | 36.4% | 9.1% | 5.5% | 1.8% | 0.0% | 0.0% | 2.31 |
| Knowledge graph reasoning | 8000 | 69.6% | 30.4% | 12.5% | 8.9% | 5.4% | 0.0% | 0.0% | 2.27 |
| Knowledge graph reasoning | 10000 | 71.9% | 31.2% | 3.1% | 0.0% | 0.0% | 0.0% | 0.0% | 2.06 |
| Knowledge graph reasoning | 12000 | 70.2% | 31.6% | 14.0% | 7.0% | 1.8% | 0.0% | 0.0% | 2.25 |
| Knowledge graph reasoning | 14000 | 59.6% | 28.1% | 12.3% | 8.8% | 3.5% | 0.0% | 0.0% | 2.12 |
| Legal | 2000 | 74.7% | 48.4% | 29.4% | 17.5% | 10.3% | 4.7% | 2.9% | 2.88 |
| Legal | 4000 | 70.3% | 42.2% | 23.1% | 12.7% | 5.6% | 2.6% | 1.6% | 2.58 |
| Legal | 6000 | 71.6% | 43.1% | 19.7% | 11.1% | 5.6% | 2.4% | 1.0% | 2.54 |
| Legal | 8000 | 71.4% | 39.2% | 16.0% | 5.9% | 3.5% | 1.7% | 0.9% | 2.39 |
| Legal | 10000 | 68.7% | 39.5% | 17.4% | 10.1% | 5.7% | 2.5% | 1.0% | 2.45 |
| Legal | 12000 | 68.5% | 40.8% | 19.1% | 7.5% | 3.4% | 1.3% | 0.8% | 2.41 |
| Legal | 14000 | 64.5% | 33.2% | 16.8% | 7.1% | 3.2% | 1.2% | 0.9% | 2.27 |
| Literary | 2000 | 66.9% | 36.0% | 16.8% | 9.4% | 4.7% | 2.9% | 1.8% | 2.38 |
| Literary | 4000 | 66.1% | 34.6% | 15.8% | 8.0% | 3.5% | 1.5% | 0.9% | 2.30 |
| Literary | 6000 | 60.4% | 31.2% | 13.2% | 6.3% | 1.5% | 0.7% | 0.5% | 2.14 |
| Literary | 8000 | 61.1% | 30.1% | 13.4% | 5.8% | 2.0% | 0.7% | 0.5% | 2.14 |
| Literary | 10000 | 62.9% | 31.5% | 12.4% | 5.4% | 1.7% | 0.7% | 0.3% | 2.15 |
| Literary | 12000 | 63.1% | 32.5% | 12.5% | 3.9% | 1.3% | 0.7% | 0.2% | 2.14 |
| Literary | 14000 | 60.5% | 31.3% | 13.4% | 4.5% | 2.5% | 1.0% | 0.5% | 2.14 |
| Many-shot learning | 2000 | 68.2% | 38.2% | 18.9% | 8.2% | 4.8% | 1.0% | 1.0% | 2.40 |
| Many-shot learning | 4000 | 63.2% | 30.5% | 13.6% | 5.4% | 3.0% | 1.4% | 1.1% | 2.18 |
| Many-shot learning | 6000 | 65.0% | 27.6% | 6.7% | 1.8% | 1.1% | 0.2% | 0.2% | 2.02 |
| Many-shot learning | 8000 | 64.6% | 39.2% | 23.5% | 9.7% | 4.0% | 1.9% | 0.9% | 2.44 |
| Many-shot learning | 10000 | 64.5% | 36.5% | 10.8% | 4.5% | 1.7% | 0.9% | 0.0% | 2.19 |
| Many-shot learning | 12000 | 70.5% | 43.3% | 19.6% | 5.9% | 3.0% | 0.4% | 0.0% | 2.43 |
| Many-shot learning | 14000 | 77.8% | 46.4% | 15.2% | 7.2% | 3.5% | 0.8% | 0.0% | 2.51 |
| Multi-news | 2000 | 73.5% | 47.3% | 31.4% | 18.3% | 10.6% | 7.0% | 3.4% | 2.92 |
| Multi-news | 4000 | 74.4% | 44.2% | 23.5% | 13.4% | 7.1% | 4.4% | 2.1% | 2.69 |
| Multi-news | 6000 | 69.2% | 43.8% | 23.7% | 11.2% | 6.1% | 4.1% | 2.2% | 2.60 |
| Multi-news | 8000 | 66.1% | 36.0% | 17.9% | 9.4% | 4.4% | 2.6% | 1.1% | 2.37 |
| Multi-news | 10000 | 66.6% | 39.6% | 21.3% | 11.1% | 5.5% | 3.3% | 2.1% | 2.50 |
| Multi-news | 12000 | 67.1% | 39.9% | 18.5% | 9.8% | 5.8% | 3.3% | 1.5% | 2.46 |
| Multi-news | 14000 | 63.7% | 39.3% | 20.2% | 9.2% | 4.3% | 2.4% | 1.3% | 2.40 |
| New language translation | 2000 | 67.0% | 36.0% | 15.9% | 6.9% | 3.2% | 1.5% | 1.0% | 2.32 |
| New language translation | 4000 | 65.4% | 35.4% | 14.5% | 4.1% | 2.0% | 1.5% | 1.5% | 2.24 |
| New language translation | 6000 | 64.0% | 34.6% | 11.6% | 3.9% | 1.1% | 0.5% | 0.0% | 2.16 |
| New language translation | 8000 | 65.1% | 31.6% | 11.4% | 2.1% | 0.6% | 0.4% | 0.0% | 2.11 |
| New language translation | 10000 | 73.9% | 37.1% | 12.7% | 3.0% | 1.5% | 0.4% | 0.0% | 2.29 |
| New language translation | 12000 | 71.7% | 31.7% | 13.0% | 3.7% | 1.2% | 0.6% | 0.0% | 2.22 |
| New language translation | 14000 | 69.2% | 34.6% | 12.5% | 3.4% | 1.0% | 0.4% | 0.0% | 2.21 |
| Table QA | 2000 | 71.2% | 42.2% | 23.8% | 13.7% | 5.9% | 2.4% | 1.7% | 2.61 |
| Table QA | 4000 | 69.5% | 39.6% | 21.8% | 11.4% | 6.4% | 3.0% | 1.3% | 2.53 |
| Table QA | 6000 | 68.6% | 41.5% | 20.8% | 10.3% | 5.5% | 2.3% | 1.3% | 2.50 |
| Table QA | 8000 | 66.8% | 41.7% | 21.2% | 12.0% | 7.0% | 4.4% | 2.0% | 2.55 |
| Table QA | 10000 | 65.1% | 38.0% | 19.5% | 9.3% | 6.2% | 3.5% | 1.4% | 2.43 |
| Table QA | 12000 | 65.2% | 35.4% | 17.4% | 9.1% | 6.1% | 3.6% | 1.0% | 2.38 |
| Table QA | 14000 | 67.1% | 37.4% | 17.4% | 9.0% | 5.1% | 2.5% | 1.0% | 2.39 |
| User guide QA | 2000 | 66.8% | 43.1% | 23.5% | 9.7% | 4.0% | 1.0% | 0.4% | 2.48 |
| User guide QA | 4000 | 65.3% | 37.2% | 15.7% | 6.4% | 3.0% | 1.3% | 0.7% | 2.30 |
| User guide QA | 6000 | 61.5% | 28.4% | 9.8% | 4.2% | 2.1% | 0.2% | 0.2% | 2.06 |
| User guide QA | 8000 | 65.6% | 34.0% | 15.3% | 4.6% | 1.2% | 0.4% | 0.2% | 2.21 |
| User guide QA | 10000 | 60.5% | 30.3% | 10.3% | 3.2% | 0.6% | 0.2% | 0.2% | 2.05 |
| User guide QA | 12000 | 63.1% | 31.8% | 13.4% | 5.9% | 2.0% | 0.7% | 0.2% | 2.17 |
| User guide QA | 14000 | 62.4% | 30.3% | 11.1% | 4.4% | 1.7% | 0.9% | 0.5% | 2.11 |
References
Paper: DFlash: Block Diffusion for Flash Speculative Decoding
- Downloads last month
- 135
