| --- |
| license: other |
| license_name: nvidia-open-model-license-agreement |
| license_link: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/ |
| library_name: gr00t |
| base_model: nvidia/GR00T-N1.7-3B |
| base_model_relation: finetune |
| datasets: |
| - cloudwalk-research/gr00t-g1-grab-bottle-right-hand-v11 |
| tags: |
| - gr00t |
| - gr00t-n1.7 |
| - vla |
| - vision-language-action |
| - humanoid |
| - robotics |
| - imitation-learning |
| - diffusion-policy |
| - unitree-g1 |
| - sonic-wbc |
| - arxiv:2503.14734 |
| - arxiv:2511.07820 |
| language: |
| - en |
| pipeline_tag: robotics |
| --- |
| |
| # GR00T N1.7-3B Fine-Tune v11 β Unitree G1 "grab the bottle" (right hand, SONIC WBC, Break-Down + Speed-Up merge, 40k steps) |
|
|
| **v11 fine-tune of [NVIDIA GR00T N1.7-3B](https://huggingface.co/nvidia/GR00T-N1.7-3B) ([paper](https://arxiv.org/abs/2503.14734)) on a Unitree G1 humanoid driven by the [SONIC](https://arxiv.org/abs/2511.07820) whole-body controller.** |
|
|
| Single-handed pick task ("grab the bottle", RIGHT hand). Trained at the **CloudWalk Robotics Lab (CW-RL)**, 2026-07, on the **"Break Down + Speed Up" merged set (355 episodes / 87,148 frames)** for the `UNITREE_G1_SONIC` embodiment. Released as a reference fine-tune for teams building manipulation policies on the GR00T + SONIC + MuJoCo/G1 stack. |
|
|
| **What makes v11 different β the merged set curated by "Break Down", then resampled by "Speed Up".** v11's dataset is the **merge of the two raw sources** β `gr00t-g1-grab-bottle-right-hand-105ep-v1` (DS1, 105 ep) + `β¦-worst-positions-empty-115ep-v3` (DS2, 115 ep) β first run through **Break-Down wandering removal** (rising segments of the distance-to-goal curve removed; each kept sub-segment becomes its own episode), then through a **DP resampling ("Speed Up")** applied independently per sub-segment with a **dynamic per-frame target** (2 mm/frame that scales down near the goal, `d_ref 100 mm`, so the fine approach motion stays dense). Result: **355 episodes** (196 DS1 sub-segments + 139 DS2 sub-segments + 20 DS2 as-is) / **87,148 frames** (~73,303 training samples). This makes v11 the **direct sibling of [v10](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-371ep-v10-finetune)** (Break Down only, no speedup): the **v10βv11 pair isolates the Speed-Up effect** given a fixed Break-Down pass. The natural comparisons are **v11 vs v10** (does DP speedup help already-Break-Down-curated data?) and **v11 vs the v2 champion**; secondarily **v11 vs [v9](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-193ep-v9-finetune)** (both use the same gentle 2 mm DP speedup, but v9 on the *v2-curated* data and v11 on the *Break-Down-merged* data). |
|
|
| **Predecessors:** [v2 (`β¦-210ep-v2-finetune`)](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-210ep-v2-finetune) β the current **validated production champion** (`checkpoint-20000`, 11/12) β [v4 (radius-5)](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-417ep-v4-finetune), [v5 (radius-20)](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-314ep-v5-finetune), [v6 (merged radius-20)](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-502ep-v6-finetune), [v7 (speedup-3mm raw)](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-220ep-v7-finetune), [v8 (speedup-3mm + cycle-removed raw)](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-405ep-v8-finetune), [v9 (speedup-2mm on curated)](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-193ep-v9-finetune), [v10 (Break-Down merge)](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-371ep-v10-finetune) β all eval pending β and [v1 (105ep)](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-105ep-v1-finetune). |
|
|
| This is a behavior-cloning fine-tune of the full 3B model. **Status: trained 2026-07-08** (**40,000 steps**, save every 10,000 β `{checkpoint-10000, checkpoint-20000, checkpoint-30000, checkpoint-40000}`, all four published here; wall-clock 2 h 12 min, W&B offline `redacted`). Closed-loop comparison against v10 (speed-up isolation) and the v2 champion is pending (see [Evaluation](#evaluation)). |
|
|
| ## Quick facts |
|
|
| | | | |
| | --- | --- | |
| | Base | [GR00T-N1.7-3B](https://huggingface.co/nvidia/GR00T-N1.7-3B) β Qwen3-VL vision-language backbone + flow-matching diffusion-transformer (DiT) action head | |
| | Parameters | 3.14 B total / 1.62 B trainable (51.5%) | |
| | Dataset | [cloudwalk-research/gr00t-g1-grab-bottle-right-hand-v11](https://huggingface.co/datasets/cloudwalk-research/gr00t-g1-grab-bottle-right-hand-v11) β **355 episodes, 87,148 frames** @ 50 Hz (~73,303 training samples), 480Γ640 `ego_view` camera (no wrist cams); merged DS1 (105ep) + DS2 (115ep worst-positions), **Break-Down v1-wandering curation then DP "Speed Up" (2 mm/frame dynamic target)** | |
| | Robot target | Unitree G1 (29-DoF body) + Inspire FTP hands (7-DoF/hand) + SONIC whole-body controller | |
| | Embodiment tag | `UNITREE_G1_SONIC` (`unitree_g1_sonic`) | |
| | State space | 43-D (left_leg 6 + right_leg 6 + waist 3 + left_arm 7 + left_hand 7 + right_arm 7 + right_hand 7) | |
| | Action space | `[40 Γ 78]` = 40-step horizon Γ (64 motion_token + 7 left_hand_joints + 7 right_hand_joints) | |
| | Hardware | 6Γ NVIDIA B200 (sm_100 / Blackwell) | |
| | Mixed precision | bf16 | |
| | Optimizer | AdamW, lr 1e-4 cosine, warmup_ratio 0.05, weight_decay 1e-5 | |
| | Steps / batch | 40,000 / global batch 48 (8 per GPU Γ 6 GPUs); checkpoints saved every 10,000 β `{checkpoint-10000, checkpoint-20000, checkpoint-30000, checkpoint-40000}` | |
| | Epochs | 5.51 @ 10k / 11.02 @ 20k / 16.52 @ 30k / 22.03 @ 40k (1.92M frame-views Γ· 87,148 frames) | |
| | Augmentation | color jitter (brightness 0.3, contrast 0.4, saturation 0.5, hue 0.08) | |
| | Wall-clock | **2 h 12 min 21 s** for 40k steps on 6Γ B200; ~5.3 it/s | |
| | Final train loss | per rung: **10k = 0.0698 Β· 20k = 0.0412 Β· 30k = 0.0240 Β· 40k = 0.0302** (per-step log; min 0.0144, mean 0.0649 over 4000 logged steps). No held-out split β a fit probe, not generalization. | |
| | W&B run | offline `redacted` (project `g1_grab_bottle`) | |
|
|
| ## Repository contents |
|
|
| ``` |
| checkpoint-10000/ # 10k steps (~5.5 ep) |
| checkpoint-20000/ # 20k steps (~11.0 ep) β under-side bracket of v2's sweet spot |
| checkpoint-30000/ # 30k steps (~16.5 ep) β over-side bracket, natural production candidate |
| checkpoint-40000/ # 40k steps (~22.0 ep) β most-trained rung (overfit-knee probe) |
| README.md # this file |
| ``` |
|
|
| > All four `{checkpoint-10000, checkpoint-20000, checkpoint-30000, checkpoint-40000}` are published so the closed-loop eval can pick the best one. `{20k, 30k}` **bracket** v2's validated ~15.3-epoch sweet spot (20k just under, 30k just over); 40k pushes to ~22 epochs to locate the overfit knee. |
|
|
| Each `checkpoint-NNNNN/` is a self-contained, deploy-ready snapshot: |
|
|
| - `model-00001-of-00002.safetensors` + `model-00002-of-00002.safetensors` (~6.5 GB, bf16) |
| - `model.safetensors.index.json` |
| - `config.json` |
| - `embodiment_id.json` (contains `unitree_g1_sonic`) |
| - `processor_config.json` |
| - `statistics.json` (dataset normalization stats) |
| - `experiment_cfg/` (training config snapshot) |
|
|
| DeepSpeed ZeRO partition states, optimizer, scheduler and RNG files are intentionally omitted β they are needed only to resume training and would add several GB per checkpoint. |
|
|
| ## Evaluation |
|
|
| **Epoch math.** This run trains for 40,000 steps = **22.03 epochs** on the 87,148-frame set (30,000 = 16.52; 20,000 = 11.02; 10,000 = 5.51). v2's validated sweet spot is ~15.3 epochs, so `{20k, 30k}` **bracket** it (11.0 ep just under, 16.5 ep just over) and 40k (22.0 ep) probes past it for the overfit knee. The production checkpoint is chosen by **closed-loop** comparison, not by training loss. There is **no held-out split**; train loss is a fit probe, **not** a generalization measure. |
|
|
| **Closed-loop evaluation β TBD.** `checkpoint-{10000,20000,30000,40000}` will be run closed-loop on the G1 + SONIC stack and compared against each other, the **v2 `checkpoint-20000` champion** (11/12 across hand-placed bottle poses), and the v4βv10 lineage. The central questions are **v11 vs v10** (does the DP "Speed Up" help data already curated by Break Down? β the v10βv11 pair isolates the speedup at matched steps {10k,20k,30k}) and **v11 vs v2**. If neither rung matches or beats the v2 champion, **v2 remains in production**. (This section is updated with the verdict once the eval runs.) |
|
|
| ## How to download |
|
|
| ```python |
| from huggingface_hub import snapshot_download |
| |
| local = snapshot_download( |
| repo_id="cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-355ep-v11-finetune", |
| repo_type="model", |
| allow_patterns=["checkpoint-30000/*"], # or checkpoint-10000/* / -20000/* / -40000/* |
| ) |
| print(local) |
| ``` |
|
|
| ## How to deploy |
|
|
| Start the GR00T policy server (from an [Isaac-GR00T](https://github.com/NVIDIA/Isaac-GR00T) environment) pointing at the downloaded checkpoint: |
|
|
| ```bash |
| python -m gr00t.eval.run_gr00t_server \ |
| --model-path <local>/checkpoint-30000 \ |
| --embodiment-tag UNITREE_G1_SONIC \ |
| --device cuda:0 --host 0.0.0.0 --port 5550 |
| ``` |
|
|
| Closed-loop control of the G1 (sim or real) is driven by the **SONIC whole-body controller** in [GR00T-WholeBodyControl](https://github.com/NVlabs/GR00T-WholeBodyControl): the policy emits `motion_token` + hand-joint targets that the SONIC WBC decodes into whole-body joint commands. The server must be launched with the **same `UNITREE_G1_SONIC` embodiment tag** used in training. See the NVlabs [VLA inference tutorial](https://nvlabs.github.io/GR00T-WholeBodyControl/tutorials/vla_inference.html). This checkpoint is **not** plug-and-play on hardware β it requires the SONIC C++ deploy stack and the matching G1 setup. |
|
|
| ## Known caveats |
|
|
| 1. **Right-hand-only, single task.** The dataset is one task ("grab the bottle") executed with the right hand. Left-hand and locomotion action dims reflect the (largely stationary) demonstrations; do not expect bimanual or locomotion behavior. |
| 2. **Single camera.** Only the `ego_view` (head) camera was recorded β no wrist cameras. |
| 3. **Deployment needs the SONIC stack.** The checkpoint outputs `motion_token` + hand joints for `UNITREE_G1_SONIC`; it only produces robot motion through the SONIC WBC + ZMQ deploy pipeline. It is not directly executable on a bare G1. |
| 4. **Train loss is not held-out.** No episode split; the loss is a smoothed train-fit probe, not a generalization metric. |
| 5. **Break-Down + Speed-Up merge.** v11's data is the merged DS1+DS2 set curated by Break-Down v1-wandering **and then** DP-resampled ("Speed Up", 2 mm/frame dynamic target). The DP pass drops ~37% of frames vs the Break-Down-only v10 (138,546 β 87,148). Compare against **v10** (the speed-up-isolation sibling) and the **v2 champion**, and against **v9** (same 2 mm speedup on the v2-curated base) β not against v7/v8 (speedup-3mm on raw, 20k steps). |
| 6. **Epoch zone β brackets the sweet spot.** 40,000 steps β 22.0 epochs; 30k β 16.5; 20k β 11.0; 10k β 5.5. `{20k, 30k}` bracket v2's ~15.3-epoch sweet spot; the eval β not the training loss β picks the rung, and 40k probes whether more epochs help or overfit. |
| 7. **Four published checkpoints.** `checkpoint-10000/-20000/-30000/-40000` are all published so the closed-loop eval can pick the best rung. |
| 8. **Training environment.** Trained on a shared B200 node with `NCCL_IB_DISABLE=1 NCCL_P2P_LEVEL=NVL` (InfiniBand off, P2P over NVLink) and W&B in offline mode (synced post-run). These affect only the training run, not the weights. |
|
|
| ## Lineage |
|
|
| | Version | Dataset | Episodes | Frames | Epochs | Closed-loop | Notes | |
| | --- | --- | --- | --- | --- | --- | --- | |
| | [v1](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-105ep-v1-finetune) | [105ep-v1](https://huggingface.co/datasets/cloudwalk-research/gr00t-g1-grab-bottle-right-hand-105ep-v1) | 105 | 70,680 | 13.6 @20k | β
validated | First GR00T N1.7 + SONIC fine-tune at CW-RL; right-hand bottle pick. | |
| | [v2](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-210ep-v2-finetune) | [right-hand-v2](https://huggingface.co/datasets/cloudwalk-research/gr00t-g1-grab-bottle-right-hand-v2) | 210 | 62,772 | 15.3 @20k | β
validated (> v1) | Curated (windows split, bad segments removed). **`checkpoint-20000` = current production champion (11/12).** | |
| | [v4](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-417ep-v4-finetune) | [radius-5](https://huggingface.co/datasets/cloudwalk-research/gr00t-g1-grab-bottle-right-hand-zero-wandering-smooth-radius-5) | 417 | 48,577 | 19.8 @20k | β³ TBD | Zero-wandering, **most aggressive** curation (radius 5). | |
| | [v5](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-314ep-v5-finetune) | [radius-20](https://huggingface.co/datasets/cloudwalk-research/gr00t-g1-grab-bottle-right-hand-zero-wandering-smooth-radius-20) | 314 | 50,496 | 19.0 @20k | β³ TBD | Zero-wandering, **least aggressive** curation (radius 20). | |
| | [v6](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-502ep-v6-finetune) | [radius-20-merged](https://huggingface.co/datasets/cloudwalk-research/gr00t-g1-grab-bottle-right-hand-radius-20-merged) | 502 | 120,017 | 8.0 @20k | β³ TBD | **Merged** (105-ep + 115-ep "worst-positions"), radius-20 with grasp-frame preservation. | |
| | [v7](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-220ep-v7-finetune) | [speedup-3mm-v1](https://huggingface.co/datasets/cloudwalk-research/gr00t-g1-grab-bottle-right-hand-speedup-3mm-v1) | 220 | 60,163 | 16.0 @20k | β³ TBD | **DP speedup** (wrist-Cartesian, 3 mm/frame) of the **raw merged** set; no segment removal β raw-branch baseline. | |
| | [v8](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-405ep-v8-finetune) | [speedup-3mm-cycle-removed-v1](https://huggingface.co/datasets/cloudwalk-research/gr00t-g1-grab-bottle-right-hand-speedup-3mm-cycle-removed-v1) | 405 | 47,944 | 20.0 @20k | β³ TBD | **DP speedup + segment removal** on the **raw merged** set (the v7 follow-up). | |
| | [v9](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-193ep-v9-finetune) | [speedup-2mm-v3](https://huggingface.co/datasets/cloudwalk-research/gr00t-g1-grab-bottle-right-hand-speedup-2mm-v3) | 193 | 32,786 | 29.3 @20k | β³ TBD | **Gentle 2 mm DP speedup on the *curated* (v2-lineage) data**. Smallest set. Trained (loss 0.0308 @20k). | |
| | [v10](https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-371ep-v10-finetune) | [right-hand-v10](https://huggingface.co/datasets/cloudwalk-research/gr00t-g1-grab-bottle-right-hand-v10) | 371 | 138,546 | 10.4 @30k | β³ TBD | **Largest set** β merged DS1+DS2, **Break-Down v1-wandering only** (no speedup). Trained (loss 0.0280 @30k). **Speed-up-isolation sibling of v11.** | |
| | **v11 (this)** | [right-hand-v11](https://huggingface.co/datasets/cloudwalk-research/gr00t-g1-grab-bottle-right-hand-v11) | 355 | 87,148 | **22.0 @40k** | β³ TBD | **Break-Down + Speed-Up** (v10 merge, then DP 2 mm/frame dynamic resampling). Trained 2026-07-08 (30k=0.0240, 40k=0.0302, min 0.0144). 40k steps β `{20k,30k}` bracket v2's ~15.3 sweet spot, 40k probes the overfit knee. Compare vs **v10** (speedup isolation) and the **v2 champion**. | |
|
|
| ## References |
|
|
| - **GR00T N1** β Open foundation model for generalist humanoid robots; base policy fine-tuned here. [Paper (arXiv:2503.14734)](https://arxiv.org/abs/2503.14734), [base model](https://huggingface.co/nvidia/GR00T-N1.7-3B). |
| - **Isaac-GR00T** β Training / inference / deployment stack used for this fine-tune. [GitHub](https://github.com/NVIDIA/Isaac-GR00T). |
| - **GR00T-WholeBodyControl (SONIC)** β Whole-body controller + teleop + VLA deployment for the G1. [GitHub](https://github.com/NVlabs/GR00T-WholeBodyControl), [SONIC paper (arXiv:2511.07820)](https://arxiv.org/abs/2511.07820). |
| - **Training dataset** β [`cloudwalk-research/gr00t-g1-grab-bottle-right-hand-v11`](https://huggingface.co/datasets/cloudwalk-research/gr00t-g1-grab-bottle-right-hand-v11) (CloudWalk Research, 2026), PICO 4 Ultra teleoperation on the Unitree G1 with SONIC WBC; merged DS1 (105ep) + DS2 (115ep worst-positions), Break-Down v1-wandering curation + DP "Speed Up" (2 mm/frame dynamic target). |
|
|
| ## Attribution |
|
|
| Developed by **cloudwalk-research** in the **CloudWalk Robotics Lab (CW-RL)**. Fine-tuned from [`nvidia/GR00T-N1.7-3B`](https://huggingface.co/nvidia/GR00T-N1.7-3B) using [Isaac-GR00T](https://github.com/NVIDIA/Isaac-GR00T); targets the Unitree G1 with the [SONIC](https://github.com/NVlabs/GR00T-WholeBodyControl) whole-body controller. |
|
|
| ## Citation |
|
|
| ```bibtex |
| @misc{cwrl_gr00t_grab_bottle_v11_2026, |
| title = {GR00T N1.7 Fine-Tune v11 --- Unitree G1 "grab the bottle" (right hand, SONIC WBC, Break-Down + Speed-Up merge, 40k steps)}, |
| author = {{CloudWalk Robotics Lab}}, |
| year = {2026}, |
| howpublished = {Hugging Face model repository}, |
| url = {https://huggingface.co/cloudwalk-research/gr00t-n17-g1-grab-bottle-rh-355ep-v11-finetune} |
| } |
| ``` |
|
|
| ## License |
|
|
| Inherits the **NVIDIA Open Model License Agreement** of the base model [`nvidia/GR00T-N1.7-3B`](https://huggingface.co/nvidia/GR00T-N1.7-3B) β see the [license terms](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/). This is a research preview: not intended for safety-critical use; closed-loop deployment on a physical humanoid requires human oversight. |
|
|