license: other
license_name: nvidia-open-model-license-agreement
license_link: >-
https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/
library_name: gr00t
base_model: nvidia/GR00T-N1.7-3B
base_model_relation: finetune
datasets:
- cloudwalk-research/gr00t-g1-grab-bottle-right-hand-v11
tags:
- gr00t
- gr00t-n1.7
- vla
- vision-language-action
- humanoid
- robotics
- imitation-learning
- diffusion-policy
- unitree-g1
- sonic-wbc
- arxiv:2503.14734
- arxiv:2511.07820
language:
- en
pipeline_tag: robotics
GR00T N1.7-3B Fine-Tune v11 β Unitree G1 "grab the bottle" (right hand, SONIC WBC, Break-Down + Speed-Up merge, 40k steps)
v11 fine-tune of NVIDIA GR00T N1.7-3B (paper) on a Unitree G1 humanoid driven by the SONIC whole-body controller.
Single-handed pick task ("grab the bottle", RIGHT hand). Trained at the CloudWalk Robotics Lab (CW-RL), 2026-07, on the "Break Down + Speed Up" merged set (355 episodes / 87,148 frames) for the UNITREE_G1_SONIC embodiment. Released as a reference fine-tune for teams building manipulation policies on the GR00T + SONIC + MuJoCo/G1 stack.
What makes v11 different β the merged set curated by "Break Down", then resampled by "Speed Up". v11's dataset is the merge of the two raw sources β gr00t-g1-grab-bottle-right-hand-105ep-v1 (DS1, 105 ep) + β¦-worst-positions-empty-115ep-v3 (DS2, 115 ep) β first run through Break-Down wandering removal (rising segments of the distance-to-goal curve removed; each kept sub-segment becomes its own episode), then through a DP resampling ("Speed Up") applied independently per sub-segment with a dynamic per-frame target (2 mm/frame that scales down near the goal, d_ref 100 mm, so the fine approach motion stays dense). Result: 355 episodes (196 DS1 sub-segments + 139 DS2 sub-segments + 20 DS2 as-is) / 87,148 frames (~73,303 training samples). This makes v11 the direct sibling of v10 (Break Down only, no speedup): the v10βv11 pair isolates the Speed-Up effect given a fixed Break-Down pass. The natural comparisons are v11 vs v10 (does DP speedup help already-Break-Down-curated data?) and v11 vs the v2 champion; secondarily v11 vs v9 (both use the same gentle 2 mm DP speedup, but v9 on the v2-curated data and v11 on the Break-Down-merged data).
Predecessors: v2 (β¦-210ep-v2-finetune) β the current validated production champion (checkpoint-20000, 11/12) β v4 (radius-5), v5 (radius-20), v6 (merged radius-20), v7 (speedup-3mm raw), v8 (speedup-3mm + cycle-removed raw), v9 (speedup-2mm on curated), v10 (Break-Down merge) β all eval pending β and v1 (105ep).
This is a behavior-cloning fine-tune of the full 3B model. Status: trained 2026-07-08 (40,000 steps, save every 10,000 β {checkpoint-10000, checkpoint-20000, checkpoint-30000, checkpoint-40000}, all four published here; wall-clock 2 h 12 min, W&B offline redacted). Closed-loop comparison against v10 (speed-up isolation) and the v2 champion is pending (see Evaluation).
Quick facts
| Base | GR00T-N1.7-3B β Qwen3-VL vision-language backbone + flow-matching diffusion-transformer (DiT) action head |
| Parameters | 3.14 B total / 1.62 B trainable (51.5%) |
| Dataset | cloudwalk-research/gr00t-g1-grab-bottle-right-hand-v11 β 355 episodes, 87,148 frames @ 50 Hz (~73,303 training samples), 480Γ640 ego_view camera (no wrist cams); merged DS1 (105ep) + DS2 (115ep worst-positions), Break-Down v1-wandering curation then DP "Speed Up" (2 mm/frame dynamic target) |
| Robot target | Unitree G1 (29-DoF body) + Inspire FTP hands (7-DoF/hand) + SONIC whole-body controller |
| Embodiment tag | UNITREE_G1_SONIC (unitree_g1_sonic) |
| State space | 43-D (left_leg 6 + right_leg 6 + waist 3 + left_arm 7 + left_hand 7 + right_arm 7 + right_hand 7) |
| Action space | [40 Γ 78] = 40-step horizon Γ (64 motion_token + 7 left_hand_joints + 7 right_hand_joints) |
| Hardware | 6Γ NVIDIA B200 (sm_100 / Blackwell) |
| Mixed precision | bf16 |
| Optimizer | AdamW, lr 1e-4 cosine, warmup_ratio 0.05, weight_decay 1e-5 |
| Steps / batch | 40,000 / global batch 48 (8 per GPU Γ 6 GPUs); checkpoints saved every 10,000 β {checkpoint-10000, checkpoint-20000, checkpoint-30000, checkpoint-40000} |
| Epochs | 5.51 @ 10k / 11.02 @ 20k / 16.52 @ 30k / 22.03 @ 40k (1.92M frame-views Γ· 87,148 frames) |
| Augmentation | color jitter (brightness 0.3, contrast 0.4, saturation 0.5, hue 0.08) |
| Wall-clock | 2 h 12 min 21 s for 40k steps on 6Γ B200; ~5.3 it/s |
| Final train loss | per rung: 10k = 0.0698 Β· 20k = 0.0412 Β· 30k = 0.0240 Β· 40k = 0.0302 (per-step log; min 0.0144, mean 0.0649 over 4000 logged steps). No held-out split β a fit probe, not generalization. |
| W&B run | offline redacted (project g1_grab_bottle) |
Repository contents
checkpoint-10000/ # 10k steps (~5.5 ep)
checkpoint-20000/ # 20k steps (~11.0 ep) β under-side bracket of v2's sweet spot
checkpoint-30000/ # 30k steps (~16.5 ep) β over-side bracket, natural production candidate
checkpoint-40000/ # 40k steps (~22.0 ep) β most-trained rung (overfit-knee probe)
README.md # this file
All four
{checkpoint-10000, checkpoint-20000, checkpoint-30000, checkpoint-40000}are published so the closed-loop eval can pick the best one.{20k, 30k}bracket v2's validated ~15.3-epoch sweet spot (20k just under, 30k just over); 40k pushes to ~22 epochs to locate the overfit knee.
Each checkpoint-NNNNN/ is a self-contained, deploy-ready snapshot:
model-00001-of-00002.safetensors+model-00002-of-00002.safetensors(~6.5 GB, bf16)model.safetensors.index.jsonconfig.jsonembodiment_id.json(containsunitree_g1_sonic)processor_config.jsonstatistics.json(dataset normalization stats)experiment_cfg/(training config snapshot)
DeepSpeed ZeRO partition states, optimizer, scheduler and RNG files are intentionally omitted β they are needed only to resume training and would add several GB per checkpoint.
Evaluation
Epoch math. This run trains for 40,000 steps = 22.03 epochs on the 87,148-frame set (30,000 = 16.52; 20,000 = 11.02; 10,000 = 5.51). v2's validated sweet spot is ~15.3 epochs, so {20k, 30k} bracket it (11.0 ep just under, 16.5 ep just over) and 40k (22.0 ep) probes past it for the overfit knee. The production checkpoint is chosen by closed-loop comparison, not by training loss. There is no held-out split; train loss is a fit probe, not a generalization measure.
Closed-loop evaluation β TBD. checkpoint-{10000,20000,30000,40000} will be run closed-loop on the G1 + SONIC stack and compared against each other, the v2 checkpoint-20000 champion (11/12 across hand-placed bottle poses), and the v4βv10 lineage. The central questions are v11 vs v10 (does the DP "Speed Up" help data already curated by Break Down? β the v10βv11 pair isolates the speedup at matched steps {10k,20k,30k}) and v11 vs v2. If neither rung matches or beats the v2 champion, v2 remains in production. (This section is updated with the verdict once the eval runs.)
How to download
from huggingface_hub import snapshot_download
local = snapshot_download(
repo_id="cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-355ep-v11-finetune",
repo_type="model",
allow_patterns=["checkpoint-30000/*"], # or checkpoint-10000/* / -20000/* / -40000/*
)
print(local)
How to deploy
Start the GR00T policy server (from an Isaac-GR00T environment) pointing at the downloaded checkpoint:
python -m gr00t.eval.run_gr00t_server \
--model-path <local>/checkpoint-30000 \
--embodiment-tag UNITREE_G1_SONIC \
--device cuda:0 --host 0.0.0.0 --port 5550
Closed-loop control of the G1 (sim or real) is driven by the SONIC whole-body controller in GR00T-WholeBodyControl: the policy emits motion_token + hand-joint targets that the SONIC WBC decodes into whole-body joint commands. The server must be launched with the same UNITREE_G1_SONIC embodiment tag used in training. See the NVlabs VLA inference tutorial. This checkpoint is not plug-and-play on hardware β it requires the SONIC C++ deploy stack and the matching G1 setup.
Known caveats
- Right-hand-only, single task. The dataset is one task ("grab the bottle") executed with the right hand. Left-hand and locomotion action dims reflect the (largely stationary) demonstrations; do not expect bimanual or locomotion behavior.
- Single camera. Only the
ego_view(head) camera was recorded β no wrist cameras. - Deployment needs the SONIC stack. The checkpoint outputs
motion_token+ hand joints forUNITREE_G1_SONIC; it only produces robot motion through the SONIC WBC + ZMQ deploy pipeline. It is not directly executable on a bare G1. - Train loss is not held-out. No episode split; the loss is a smoothed train-fit probe, not a generalization metric.
- Break-Down + Speed-Up merge. v11's data is the merged DS1+DS2 set curated by Break-Down v1-wandering and then DP-resampled ("Speed Up", 2 mm/frame dynamic target). The DP pass drops ~37% of frames vs the Break-Down-only v10 (138,546 β 87,148). Compare against v10 (the speed-up-isolation sibling) and the v2 champion, and against v9 (same 2 mm speedup on the v2-curated base) β not against v7/v8 (speedup-3mm on raw, 20k steps).
- Epoch zone β brackets the sweet spot. 40,000 steps β 22.0 epochs; 30k β 16.5; 20k β 11.0; 10k β 5.5.
{20k, 30k}bracket v2's ~15.3-epoch sweet spot; the eval β not the training loss β picks the rung, and 40k probes whether more epochs help or overfit. - Four published checkpoints.
checkpoint-10000/-20000/-30000/-40000are all published so the closed-loop eval can pick the best rung. - Training environment. Trained on a shared B200 node with
NCCL_IB_DISABLE=1 NCCL_P2P_LEVEL=NVL(InfiniBand off, P2P over NVLink) and W&B in offline mode (synced post-run). These affect only the training run, not the weights.
Lineage
| Version | Dataset | Episodes | Frames | Epochs | Closed-loop | Notes |
|---|---|---|---|---|---|---|
| v1 | 105ep-v1 | 105 | 70,680 | 13.6 @20k | β validated | First GR00T N1.7 + SONIC fine-tune at CW-RL; right-hand bottle pick. |
| v2 | right-hand-v2 | 210 | 62,772 | 15.3 @20k | β validated (> v1) | Curated (windows split, bad segments removed). checkpoint-20000 = current production champion (11/12). |
| v4 | radius-5 | 417 | 48,577 | 19.8 @20k | β³ TBD | Zero-wandering, most aggressive curation (radius 5). |
| v5 | radius-20 | 314 | 50,496 | 19.0 @20k | β³ TBD | Zero-wandering, least aggressive curation (radius 20). |
| v6 | radius-20-merged | 502 | 120,017 | 8.0 @20k | β³ TBD | Merged (105-ep + 115-ep "worst-positions"), radius-20 with grasp-frame preservation. |
| v7 | speedup-3mm-v1 | 220 | 60,163 | 16.0 @20k | β³ TBD | DP speedup (wrist-Cartesian, 3 mm/frame) of the raw merged set; no segment removal β raw-branch baseline. |
| v8 | speedup-3mm-cycle-removed-v1 | 405 | 47,944 | 20.0 @20k | β³ TBD | DP speedup + segment removal on the raw merged set (the v7 follow-up). |
| v9 | speedup-2mm-v3 | 193 | 32,786 | 29.3 @20k | β³ TBD | Gentle 2 mm DP speedup on the curated (v2-lineage) data. Smallest set. Trained (loss 0.0308 @20k). |
| v10 | right-hand-v10 | 371 | 138,546 | 10.4 @30k | β³ TBD | Largest set β merged DS1+DS2, Break-Down v1-wandering only (no speedup). Trained (loss 0.0280 @30k). Speed-up-isolation sibling of v11. |
| v11 (this) | right-hand-v11 | 355 | 87,148 | 22.0 @40k | β³ TBD | Break-Down + Speed-Up (v10 merge, then DP 2 mm/frame dynamic resampling). Trained 2026-07-08 (30k=0.0240, 40k=0.0302, min 0.0144). 40k steps β {20k,30k} bracket v2's ~15.3 sweet spot, 40k probes the overfit knee. Compare vs v10 (speedup isolation) and the v2 champion. |
References
- GR00T N1 β Open foundation model for generalist humanoid robots; base policy fine-tuned here. Paper (arXiv:2503.14734), base model.
- Isaac-GR00T β Training / inference / deployment stack used for this fine-tune. GitHub.
- GR00T-WholeBodyControl (SONIC) β Whole-body controller + teleop + VLA deployment for the G1. GitHub, SONIC paper (arXiv:2511.07820).
- Training dataset β
cloudwalk-research/gr00t-g1-grab-bottle-right-hand-v11(CloudWalk Research, 2026), PICO 4 Ultra teleoperation on the Unitree G1 with SONIC WBC; merged DS1 (105ep) + DS2 (115ep worst-positions), Break-Down v1-wandering curation + DP "Speed Up" (2 mm/frame dynamic target).
Attribution
Developed by cloudwalk-research in the CloudWalk Robotics Lab (CW-RL). Fine-tuned from nvidia/GR00T-N1.7-3B using Isaac-GR00T; targets the Unitree G1 with the SONIC whole-body controller.
Citation
@misc{cwrl_gr00t_grab_bottle_v11_2026,
title = {GR00T N1.7 Fine-Tune v11 --- Unitree G1 "grab the bottle" (right hand, SONIC WBC, Break-Down + Speed-Up merge, 40k steps)},
author = {{CloudWalk Robotics Lab}},
year = {2026},
howpublished = {Hugging Face model repository},
url = {https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-355ep-v11-finetune}
}
License
Inherits the NVIDIA Open Model License Agreement of the base model nvidia/GR00T-N1.7-3B β see the license terms. This is a research preview: not intended for safety-critical use; closed-loop deployment on a physical humanoid requires human oversight.