NVIDIA-Nemotron-3-Ultra-550B-A55B-speculator.dflash

This is a DFlash speculator model for RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic.

Training Details

This model was trained using the Speculators library on a subset of Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered and the train_sft split of HuggingFaceH4/ultrachat_200k. Responses were regenerated by NVIDIA-Nemotron-3-Ultra-550B-A55B. Training compute for this model was generously provided by Lambda, a leading cloud platform for AI training and inference.

Commands

Using the Speculators library and the helper scripts provided in the repo.

Prepare data

# In virtual environment with speculators installed
python scripts/prepare_data.py \
  --model RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic \
  --data ./regenerated_data.jsonl \
  --output ./output \
  --seq-length 8192

Launch vLLM

# In (separate) virtual environment with vllm installed
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 vllm_venv/bin/python scripts/launch_vllm.py \
  RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic \
  --target-layer-ids 7 47 88 \
  -- --port 8000 \
  --gpu-memory-utilization 0.9 \
  --disable-uvicorn-access-log \
  --tensor-parallel-size 8

Launch training

Must be run once vLLM has finished launching and is running in the background.

# In virtual environment with speculators installed
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 torchrun \
  --standalone \
  --nproc_per_node 8 \
  scripts/train.py \
  --verifier-name-or-path RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic \
  --speculator-type dflash \
  --num-layers 5 \
  --data-path ./output \
  --vllm-endpoint http://localhost:8000/v1 \
  --save-path ./output/checkpoints \
  --epochs 5 \
  --lr 0.0006 \
  --total-seq-len 8192 \
  --on-missing generate \
  --on-generate delete \
  --seed 42 \
  --log-freq 10 \
  --draft-vocab-size 32000 \
  --draft-arch llama \
  --target-layer-ids 7 47 88 \
  --draft-hidden-act silu \
  --scheduler-type cosine \
  --max-anchors 3072 \
  --prefetch-factor 2 \
  --num-workers 16 \
  --sliding-window-indices 0 1 2 3 4 \
  --optimizer muon

Model Specifications

Base Model RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic
Chat Template RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic (use /chat/completions endpoint)
Format Safetensors
License Apache 2.0
Validation Hardware Nvidia B200

Deployment

# Deploy with speculative decoding
vllm serve RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-dynamic \
    --tensor-parallel-size 8 \
    --max-model-len 16384 \
    --speculative-config '{
        "model": "RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-speculator.dflash",
        "num_speculative_tokens": 7,
        "method": "dflash"
    }'

Acceptance Rate

Per-position token acceptance rates across datasets:

Dataset Pos 0 Pos 1 Pos 2 Pos 3 Pos 4 Pos 5 Pos 6 Pos 7 Avg. Length
HumanEval 76.3% 53.9% 37.3% 25.8% 18.8% 14.0% 10.5% 7.5% 3.44
math_reasoning 88.4% 74.8% 61.3% 50.7% 41.0% 32.9% 25.4% 19.3% 4.94
qa 64.9% 37.9% 21.5% 12.8% 7.7% 4.5% 2.7% 1.5% 2.53
question 68.4% 41.3% 24.1% 15.3% 10.4% 7.1% 4.8% 3.0% 2.74
rag 75.3% 52.0% 34.6% 23.7% 16.2% 11.1% 7.4% 4.3% 3.25
summarization 71.6% 44.5% 25.6% 14.9% 8.3% 4.7% 2.3% 1.1% 2.73
tool_call 73.3% 48.4% 29.2% 18.5% 12.3% 8.0% 5.2% 3.5% 2.98
translation 64.8% 41.0% 24.9% 14.6% 8.3% 4.7% 2.7% 1.4% 2.62
writing 68.2% 41.0% 24.1% 15.1% 10.2% 7.0% 4.7% 3.0% 2.73

Performance Eval

We used NVIDIA B200 for performance evaluation.

Performance Evaluation - HumanEval Latency Speedup

Long Context Benchmarking

Per-position token acceptance rates across long context datasets:

Dataset Context Length Pos 0 Pos 1 Pos 2 Pos 3 Pos 4 Pos 5 Pos 6 Avg. Length
Academic 2000 70.2% 44.4% 27.6% 13.9% 8.7% 5.0% 3.6% 2.73
Academic 4000 71.2% 39.3% 18.6% 10.2% 4.4% 1.7% 1.2% 2.47
Academic 6000 63.3% 35.8% 15.8% 6.0% 2.0% 1.1% 1.1% 2.25
Academic 8000 68.2% 38.3% 19.5% 7.8% 3.3% 0.7% 0.2% 2.38
Academic 10000 65.6% 35.4% 15.1% 7.2% 4.0% 1.6% 1.1% 2.30
Academic 12000 67.1% 34.8% 14.7% 6.0% 3.0% 1.4% 0.9% 2.28
Academic 14000 66.3% 34.3% 15.6% 6.7% 2.8% 1.2% 0.9% 2.28
Agent history QA 2000 69.5% 42.4% 24.1% 14.0% 9.9% 3.9% 2.4% 2.66
Agent history QA 4000 64.5% 37.9% 21.6% 11.3% 7.6% 2.9% 1.6% 2.47
Agent history QA 6000 70.3% 40.9% 23.9% 14.7% 8.4% 2.2% 1.2% 2.62
Agent history QA 8000 69.9% 41.2% 25.2% 14.8% 8.7% 2.7% 1.2% 2.64
Agent history QA 10000 67.3% 42.0% 21.0% 10.1% 6.0% 1.6% 0.0% 2.48
Agent history QA 12000 67.5% 43.6% 26.8% 14.5% 7.0% 2.5% 1.0% 2.63
Agent history QA 14000 68.0% 42.3% 21.7% 11.3% 6.6% 1.4% 0.6% 2.52
Code repo QA 2000 67.2% 38.1% 21.1% 10.6% 7.2% 4.0% 1.7% 2.50
Code repo QA 4000 65.7% 34.1% 18.1% 7.3% 1.8% 0.9% 0.2% 2.28
Code repo QA 6000 63.9% 33.6% 12.9% 6.3% 3.6% 0.7% 0.4% 2.21
Code repo QA 8000 60.2% 30.1% 11.7% 3.6% 0.9% 0.6% 0.4% 2.07
Code repo QA 10000 60.2% 31.1% 12.4% 4.2% 1.5% 0.2% 0.0% 2.10
Code repo QA 12000 59.0% 31.5% 10.9% 3.2% 0.9% 0.4% 0.2% 2.06
Code repo QA 14000 60.8% 29.2% 11.8% 4.4% 2.0% 0.3% 0.0% 2.09
Detective 2000 65.4% 34.2% 15.2% 6.8% 2.5% 1.5% 0.8% 2.26
Detective 4000 61.8% 32.9% 11.3% 3.8% 1.7% 0.5% 0.0% 2.12
Detective 6000 62.3% 29.2% 11.8% 4.1% 1.8% 0.0% 0.0% 2.09
Detective 8000 62.8% 30.7% 11.2% 4.8% 2.2% 0.5% 0.2% 2.12
Detective 10000 63.1% 29.9% 9.8% 3.9% 1.0% 0.3% 0.3% 2.08
Detective 12000 53.7% 26.4% 10.4% 3.4% 1.1% 0.3% 0.0% 1.95
Detective 14000 52.5% 26.6% 9.6% 2.0% 1.1% 0.3% 0.0% 1.92
Dialogue history QA 2000 75.0% 50.0% 29.2% 17.0% 8.3% 5.6% 3.8% 2.89
Dialogue history QA 4000 72.6% 41.2% 22.6% 11.2% 6.5% 3.8% 1.2% 2.59
Dialogue history QA 6000 73.5% 41.3% 19.1% 9.5% 5.0% 3.1% 1.2% 2.53
Dialogue history QA 8000 69.1% 34.6% 13.6% 3.4% 1.0% 0.3% 0.0% 2.22
Dialogue history QA 10000 78.3% 44.0% 24.0% 13.0% 5.3% 3.0% 2.3% 2.70
Dialogue history QA 12000 65.7% 31.7% 16.5% 6.3% 3.6% 1.5% 0.5% 2.26
Dialogue history QA 14000 66.9% 37.2% 17.7% 9.8% 6.0% 3.2% 1.5% 2.42
Event ordering 2000 65.8% 38.3% 19.2% 11.0% 5.5% 2.5% 0.4% 2.43
Event ordering 4000 60.6% 31.0% 14.9% 6.4% 2.4% 0.8% 0.7% 2.17
Event ordering 6000 60.8% 32.8% 15.0% 6.3% 2.0% 0.7% 0.2% 2.18
Event ordering 8000 60.9% 31.1% 14.7% 5.4% 1.7% 0.8% 0.3% 2.15
Event ordering 10000 58.8% 27.8% 13.1% 5.4% 1.5% 0.7% 0.0% 2.07
Event ordering 12000 61.9% 31.4% 14.7% 4.9% 1.7% 0.8% 0.2% 2.16
Event ordering 14000 62.6% 33.6% 15.6% 6.4% 2.2% 1.2% 0.5% 2.22
Financial 2000 70.0% 44.7% 28.4% 17.8% 12.9% 8.2% 4.2% 2.86
Financial 4000 66.5% 36.9% 19.0% 11.7% 6.7% 4.2% 1.5% 2.47
Financial 6000 65.8% 39.0% 20.2% 9.8% 6.7% 3.5% 1.0% 2.46
Financial 8000 64.3% 37.5% 19.8% 11.1% 6.4% 3.5% 1.7% 2.44
Financial 10000 64.6% 37.1% 19.1% 9.4% 5.2% 2.8% 1.3% 2.40
Financial 12000 64.9% 39.1% 19.7% 10.9% 7.1% 3.3% 1.3% 2.46
Financial 14000 67.1% 41.0% 21.9% 11.8% 6.8% 3.6% 2.2% 2.54
Governmental 2000 73.0% 49.8% 30.9% 16.2% 10.1% 6.5% 3.6% 2.90
Governmental 4000 72.2% 46.6% 27.6% 15.8% 9.4% 5.7% 2.4% 2.80
Governmental 6000 70.4% 44.4% 24.7% 12.3% 7.0% 3.5% 1.0% 2.63
Governmental 8000 74.1% 44.8% 24.0% 13.7% 7.6% 4.0% 1.7% 2.70
Governmental 10000 71.1% 43.8% 22.6% 11.4% 5.7% 3.7% 2.2% 2.60
Governmental 12000 70.6% 43.8% 24.3% 13.8% 6.2% 2.9% 1.9% 2.63
Governmental 14000 70.8% 39.6% 23.0% 11.6% 6.0% 3.4% 1.6% 2.56
Knowledge graph reasoning 2000 76.5% 43.1% 15.7% 9.8% 3.9% 0.0% 0.0% 2.49
Knowledge graph reasoning 4000 76.9% 43.6% 15.4% 7.7% 5.1% 0.0% 0.0% 2.49
Knowledge graph reasoning 6000 78.2% 36.4% 9.1% 5.5% 1.8% 0.0% 0.0% 2.31
Knowledge graph reasoning 8000 69.6% 30.4% 12.5% 8.9% 5.4% 0.0% 0.0% 2.27
Knowledge graph reasoning 10000 71.9% 31.2% 3.1% 0.0% 0.0% 0.0% 0.0% 2.06
Knowledge graph reasoning 12000 70.2% 31.6% 14.0% 7.0% 1.8% 0.0% 0.0% 2.25
Knowledge graph reasoning 14000 59.6% 28.1% 12.3% 8.8% 3.5% 0.0% 0.0% 2.12
Legal 2000 74.7% 48.4% 29.4% 17.5% 10.3% 4.7% 2.9% 2.88
Legal 4000 70.3% 42.2% 23.1% 12.7% 5.6% 2.6% 1.6% 2.58
Legal 6000 71.6% 43.1% 19.7% 11.1% 5.6% 2.4% 1.0% 2.54
Legal 8000 71.4% 39.2% 16.0% 5.9% 3.5% 1.7% 0.9% 2.39
Legal 10000 68.7% 39.5% 17.4% 10.1% 5.7% 2.5% 1.0% 2.45
Legal 12000 68.5% 40.8% 19.1% 7.5% 3.4% 1.3% 0.8% 2.41
Legal 14000 64.5% 33.2% 16.8% 7.1% 3.2% 1.2% 0.9% 2.27
Literary 2000 66.9% 36.0% 16.8% 9.4% 4.7% 2.9% 1.8% 2.38
Literary 4000 66.1% 34.6% 15.8% 8.0% 3.5% 1.5% 0.9% 2.30
Literary 6000 60.4% 31.2% 13.2% 6.3% 1.5% 0.7% 0.5% 2.14
Literary 8000 61.1% 30.1% 13.4% 5.8% 2.0% 0.7% 0.5% 2.14
Literary 10000 62.9% 31.5% 12.4% 5.4% 1.7% 0.7% 0.3% 2.15
Literary 12000 63.1% 32.5% 12.5% 3.9% 1.3% 0.7% 0.2% 2.14
Literary 14000 60.5% 31.3% 13.4% 4.5% 2.5% 1.0% 0.5% 2.14
Many-shot learning 2000 68.2% 38.2% 18.9% 8.2% 4.8% 1.0% 1.0% 2.40
Many-shot learning 4000 63.2% 30.5% 13.6% 5.4% 3.0% 1.4% 1.1% 2.18
Many-shot learning 6000 65.0% 27.6% 6.7% 1.8% 1.1% 0.2% 0.2% 2.02
Many-shot learning 8000 64.6% 39.2% 23.5% 9.7% 4.0% 1.9% 0.9% 2.44
Many-shot learning 10000 64.5% 36.5% 10.8% 4.5% 1.7% 0.9% 0.0% 2.19
Many-shot learning 12000 70.5% 43.3% 19.6% 5.9% 3.0% 0.4% 0.0% 2.43
Many-shot learning 14000 77.8% 46.4% 15.2% 7.2% 3.5% 0.8% 0.0% 2.51
Multi-news 2000 73.5% 47.3% 31.4% 18.3% 10.6% 7.0% 3.4% 2.92
Multi-news 4000 74.4% 44.2% 23.5% 13.4% 7.1% 4.4% 2.1% 2.69
Multi-news 6000 69.2% 43.8% 23.7% 11.2% 6.1% 4.1% 2.2% 2.60
Multi-news 8000 66.1% 36.0% 17.9% 9.4% 4.4% 2.6% 1.1% 2.37
Multi-news 10000 66.6% 39.6% 21.3% 11.1% 5.5% 3.3% 2.1% 2.50
Multi-news 12000 67.1% 39.9% 18.5% 9.8% 5.8% 3.3% 1.5% 2.46
Multi-news 14000 63.7% 39.3% 20.2% 9.2% 4.3% 2.4% 1.3% 2.40
New language translation 2000 67.0% 36.0% 15.9% 6.9% 3.2% 1.5% 1.0% 2.32
New language translation 4000 65.4% 35.4% 14.5% 4.1% 2.0% 1.5% 1.5% 2.24
New language translation 6000 64.0% 34.6% 11.6% 3.9% 1.1% 0.5% 0.0% 2.16
New language translation 8000 65.1% 31.6% 11.4% 2.1% 0.6% 0.4% 0.0% 2.11
New language translation 10000 73.9% 37.1% 12.7% 3.0% 1.5% 0.4% 0.0% 2.29
New language translation 12000 71.7% 31.7% 13.0% 3.7% 1.2% 0.6% 0.0% 2.22
New language translation 14000 69.2% 34.6% 12.5% 3.4% 1.0% 0.4% 0.0% 2.21
Table QA 2000 71.2% 42.2% 23.8% 13.7% 5.9% 2.4% 1.7% 2.61
Table QA 4000 69.5% 39.6% 21.8% 11.4% 6.4% 3.0% 1.3% 2.53
Table QA 6000 68.6% 41.5% 20.8% 10.3% 5.5% 2.3% 1.3% 2.50
Table QA 8000 66.8% 41.7% 21.2% 12.0% 7.0% 4.4% 2.0% 2.55
Table QA 10000 65.1% 38.0% 19.5% 9.3% 6.2% 3.5% 1.4% 2.43
Table QA 12000 65.2% 35.4% 17.4% 9.1% 6.1% 3.6% 1.0% 2.38
Table QA 14000 67.1% 37.4% 17.4% 9.0% 5.1% 2.5% 1.0% 2.39
User guide QA 2000 66.8% 43.1% 23.5% 9.7% 4.0% 1.0% 0.4% 2.48
User guide QA 4000 65.3% 37.2% 15.7% 6.4% 3.0% 1.3% 0.7% 2.30
User guide QA 6000 61.5% 28.4% 9.8% 4.2% 2.1% 0.2% 0.2% 2.06
User guide QA 8000 65.6% 34.0% 15.3% 4.6% 1.2% 0.4% 0.2% 2.21
User guide QA 10000 60.5% 30.3% 10.3% 3.2% 0.6% 0.2% 0.2% 2.05
User guide QA 12000 63.1% 31.8% 13.4% 5.9% 2.0% 0.7% 0.2% 2.17
User guide QA 14000 62.4% 30.3% 11.1% 4.4% 1.7% 0.9% 0.5% 2.11

References

Paper: DFlash: Block Diffusion for Flash Speculative Decoding

Downloads last month
135
Safetensors
Model size
3B params
Tensor type
I64
·
BF16
·
BOOL
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-speculator.dflash

Collection including RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-speculator.dflash

Paper for RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-speculator.dflash