GR00T-N1.6-3B-p150 / GPU_COMPARISON.md
changh95's picture
Stage-2 release: DiT denoise as one persistent megakernel (snapshot 009de892, image 0be2c39e4a01)
8ea4fb0 verified
|
Raw
History Blame
10.5 kB

GR00T-N1.6-3B-p150 β€” Blackhole p150a vs RTX 5090 (same host, same weights, same observation)

Facts only. The GPU rows were measured on the build host during the port's reference-environment study (2026-09-14, docs/research/ref-env-n16.md Β§5 of the port project); the p150a rows are the port's own benchmark of the frozen snapshot this image ships (Stage-2: megakernel denoise), the same-tree TTNN-denoise run it replaced, and the served numbers of this image. The GPU numbers were not re-measured for this file; every number names its source.

What was run

Model (GPU) NVIDIA's unmodified Gr00tPolicy from Isaac-GR00T n1.6.1-release β€” Eagle-Block2A-2B-v2 backbone (SigLIP2 252Β² β†’ 81 tokens, Qwen3-1.7B[:16]) + AlternateVLDiT action head, 4 Euler flow-matching steps β€” i.e. the network the p150a port is PCC-gated against
Weights nvidia/GR00T-N1.6-3B @ d0814e7ecb19202e7c8468b46098b0b7ef3a6d61 (= tt-model.yaml weights.revision = serve.env.TT_WEIGHTS_REVISION), from the HF cache
Observation demo_data/gr1.PickNPlace episode 0, frame 0 (GR1: one 256Γ—256 ego_view_bg_crop_pad_res256_freq20 frame, 5 state groups 7/7/6/6/3, instruction "pick the pear from the counter and place it in the plate") β€” the same step shipped as media/demo_observation.json + media/demo_ego_view.png and used by the smoke test; batch 1
GPU NVIDIA GeForce RTX 5090 (sm_120), the same host as the p150a; GPU otherwise idle (nvidia-smi compute-app list recorded in every run manifest)
GPU software Python 3.10.20, torch 2.7.1+cu128 (cuDNN 90701), transformers 4.51.3, flash_attn 2.7.4.post1 (sm_120 SASS, imports and runs natively), numpy 1.26.4 β€” the uv.lock of Isaac-GR00T-n1.6 (ref-env-n16.md Β§1)
GPU loop medians of 20 warm iterations after warm-up, torch.cuda.synchronize() around each call, one process per configuration
GPU timing definitions e2e = policy.get_action incl. ~2.5 ms of CPU preprocessing (albumentations, PIL resize, tokenisation) and the numpy decode; backbone and action head are the GPU sub-steps (action head = 4 DiT steps + encoders/decoder)
p150a (device, Stage 2 = this image's path) one Blackhole p150a, the frozen snapshot 009de892c895 of the port on tt-metal 668c2907575, benchmarks/bench_e2e.py defaults: dit_backend=megakernel (the DiT denoise as one persistent generic_op, bfp8_b DRAM weight arena), mixed_dit elsewhere (bf16), per_stage traces, 2 command queues with CQ-1 input uploads, device embedding gather, worker_l1_size 1,395,712: 10 warm-up calls, then 50 calls of encode β†’ upload β†’ 4 Γ— execute_trace β†’ D2H β†’ decode, host wall clock medians; a second 50-call pass with a sync after every trace gives the per-trace split β€” code/models/experimental/gr00t/benchmarks/results/bench_e2e_n16_20260918-050632.json
p150a (device, TTNN denoise on the same tree) the same tree and settings with --dit-backend ttnn (the Stage-1 op sequence in the denoise trace, firmware-default L1) β€” benchmarks/results/bench_e2e_n16_20260918-032258.json
p150a (served) this image, POST /predict on the demo observation, seed path, 50 requests after 5 warm-ups (code/scripts/bench_http.py), server-side timing_ms

Correctness reference (what "same output" means here)

The port is gated against the fp32 CPU run of NVIDIA's code (the golden). NVIDIA's own bf16 + flash-attention GPU path sits at a measurable distance from that golden, which bounds what any bf16 implementation can be asked to reproduce (code/models/experimental/gr00t/common/golden_tap_map.json, thresholds.n16):

tensor official bf16 GPU vs fp32 golden (PCC) p150a Stage-2 (megakernel) vs fp32 golden (PCC) p150a Stage-1 image (2026-09-15) gate
backbone_features (LLM output, 116 tokens Γ— 2048) 0.99679 0.99878 0.99878 0.9917
action_pred_valid (16 steps Γ— 29 action dims, normalised) 0.99956 0.99922 0.99921 0.998
unnormalized_actions β€” left_arm / right_arm / right_hand β€” 0.99995 / 0.99990 / 0.99308 0.99995 / 0.99991 / 0.99312 0.9999 / 0.9998 / 0.99

p150a Stage-2 values: tests/tt/results/test_e2e_{untraced_taps_vs_golden,predict_actions_vs_golden}_n16_20260918-0450*.json (the snapshot validation). Two intermediate taps of the megakernel path miss their per-tap gates on the unused state-token row (action_decoder_out[k=1] 0.999594 vs 0.9997, action_decoder_out[k=2] 0.999598 vs 0.99978; 26/28 gated rows pass); the DiT block intermediates (dit_out[k], two of which missed on the Stage-1 path) are not observable inside the megakernel; every action tap passes β€” see the card's Caveats. Megakernel vs TTNN denoise on device (same tree, mk_default_summary.md Β§3.1): action_pred_normalized PCC β‰₯ 0.999994 on every N1.6 reference sample.

GPU latency (RTX 5090, batch 1, medians of 20 iterations, ms)

configuration e2e get_action data proc (CPU) backbone (GPU) action head = 4 DiT steps (GPU)
official: bf16 + flash-attn 2, deterministic algorithms OFF, cudnn.benchmark ON, eager PyTorch 63.1 2.5 16.8 36.2
official + torch.compile(DiT.forward, mode="max-autotune") (as standalone_inference_script.py; compile 41 s; inductor skipped some Triton mm configs on sm_120 "out of shared memory") 33.1 2.4 17.1 13.5
golden run bf16_fa2 (deterministic algorithms ON, CUBLAS_WORKSPACE_CONFIG=:4096:8, cudnn.benchmark OFF) 97.3 2.6 25.3 66.8
golden run bf16_sdpa (same flags) 93.2 2.9 23.6 66.9
golden run fp32_math (same flags, fp32, math SDPA) 114.5 2.6 32.8 80.4

The official rows reproduce NVIDIA's README table for the RTX 5090 (eager 58 ms, compiled 37 ms e2e; backbone 18 ms; action head 38 β†’ 16 ms). The "golden" rows are the settings the fp32 goldens were dumped with; the deterministic mode costs ~35 ms here; on N1.7 the same CUBLAS_WORKSPACE_CONFIG=:4096:8 env var alone was shown to double the small-M DiT GEMMs (docs/research/golden-multisample.md Β§3) β€” not separately attributed for N1.6 β€” and it is not a serving configuration. The eager action head is dispatch-bound (128 DiT-block evaluations in 36 ms), which is why torch.compile gains 2.7Γ— on it β€” the regime the p150a's persistent megakernel addresses with one launch per denoise.

p150a latency (batch 1, ms)

row Stage 2 (this image): median p90 TTNN denoise, same tree: median p90 source
bench_e2e end-to-end (host encode + upload + 4 traces + D2H + host decode) 40.67 40.77 56.20 56.60 stages.e2e of bench_e2e_n16_20260918-050632.json / …-032258.json
device = upload + traces + D2H 39.17 39.22 54.39 54.77 stages.device
4 traces (one sync) 37.76 37.78 52.65 52.70 stages.traces
trace vision (SigLIP2 tower + unshuffle projector) 10.80 11.00 10.82 10.87 stages.trace/vision (split pass)
trace llm (Qwen3-1.7B[:16], S 116 β†’ 128) 9.55 9.75 9.55 9.61 stages.trace/llm
trace adapter (VL-LN, K/V hoist, state encoder; intermediates in DRAM under the megakernel) 2.24 2.44 2.01 2.07 stages.trace/adapter
trace denoise (4 Euler steps Γ— 32 DiT blocks + decoder) 15.35 (3.84 per step) 15.52 30.43 30.53 stages.trace/denoise
upload / D2H / host encode / host decode 1.38 / 0.04 / 1.41 / 0.10 1.41 / 0.05 / 1.48 / 0.10 1.66 / 0.07 / 1.61 / 0.17 1.94 / 0.09 / 1.79 / 0.23 stages.*
served timing_ms.device (this image, 50 requests after 5 warm-ups) 39.6 39.9 56.8 (Stage-1 image feb77048a2ed) 57.3 bench_http.py run of the container validation (bench-mk-n16-c1-20260918-054614.json; Stage-1: bench-n16-c1-20260915-014844.json)
served timing_ms.total (handler wall incl. base64 + PNG decode) 42.1 42.5 59.4 59.8 same
served client wall (localhost, incl. HTTP + JSON) 44.3 44.8 61.0 61.6 same

The Stage-1 release of 2026-09-15 measured 59.42 ms e2e (benchmarks/results/e2e_stage1_n16.json: 1 command queue, host embedding gather); the same-tree TTNN column above already carries the 2-CQ uploads and the device embedding gather that the tree gained since, so the megakernel's own effect is the difference between the two columns: βˆ’15.5 ms e2e (βˆ’27.6 %), all of it in the denoise trace (βˆ’15.1 ms, βˆ’49.6 %) minus the adapter's DRAM placement (+0.2 ms).

Comparison with matching definitions

Ratio = GPU ms / p150a ms (> 1 means the p150a is faster).

row p150a (Stage 2) GPU ratio
end-to-end policy call: bench_e2e e2e 40.67 (encode + device + decode) vs get_action eager bf16 e2e 63.1 (data proc + backbone + head) 40.67 63.1 1.55 (p150a faster)
same vs get_action with torch.compile(DiT) 40.67 33.1 0.81 (GPU 1.23Γ— faster)
model only: p150a device 39.17 (upload + traces + D2H) vs GPU backbone + head 53.0 (eager) 39.17 53.0 1.35 (p150a faster)
model only vs GPU compiled backbone + head 30.6 39.17 30.6 0.78 (GPU 1.28Γ— faster)
denoise: p150a denoise trace 15.35 vs GPU action head 36.2 (eager) / 13.5 (compiled) 15.35 36.2 / 13.5 2.36 / 0.88 (GPU 1.14Γ— faster)
vision + LLM: p150a vision + llm traces 20.35 vs GPU backbone 16.8 20.35 16.8 0.83 (GPU 1.21Γ— faster)

Reading: against NVIDIA's eager bf16 path the Stage-2 p150a port is 1.55Γ— faster end-to-end (the Stage-1 release was 1.06Γ—); against the compiled DiT the GPU is still 1.23Γ— faster (1.80Γ— against Stage 1). The denoise gap closed from 30.4 vs 13.5 ms to 15.35 vs 13.5 ms; the remaining e2e gap is now the backbone (vision + llm 20.35 ms vs 16.8 ms on the GPU) plus the p150a's upload / readback, not the DiT. The GPU's data proc (2.5 ms) and the p150a's host encode (1.4 ms) are the same kind of host work; HTTP/JSON framing is excluded on both sides except in the "served client wall" row.

Not measured / not claimed: p150a power (no power or efficiency comparison); TensorRT or FP8 on the GPU; the compiled GPU path with the p150a's exact host chain; N1.6's gr1_long_text layout. GPU numbers were not re-run for this card; they are copied with their source lines from ref-env-n16.md Β§5 (which also records the machine state of each run).