--- pipeline_tag: reinforcement-learning tags: [microduck, microduck-policy, mjlab, robotics, locomotion, running] license: apache-2.0 --- # microduck-running A robustified forward-running policy for the stock MicroDuck on flat ground. The default artifact is iteration **12,195**, continued from the 1.687 m/s speed frontier with progressively stronger velocity pushes, centre-of-mass variation, and initial tilt. **Status: simulation-only hardware candidate; not validated on hardware.** It retains 1.651 m/s mean body-forward speed in the nominal simulation battery and 1.612 m/s under the combined backlash-plus-disturbance battery. Heading and lateral drift remain substantial. ## Normal simulation rollout — iteration 12,195 [Open the full silent MP4](media/preview.mp4). ## Run-into-mat edit — legacy iteration 8,749 [Open the full silent run-into-mat MP4](media/run-into-mat.mp4). The mat video uses the preserved iteration-8,749 controller, not the new root policy. Its obstacle, collision augmentation, camera treatment, post-impact motor cutoff, and sparks are scripted demonstration elements. The policy has no obstacle perception and did not learn to react to a mat. Both videos are silent. ## Versions | Role | Iteration | Files | Note | |---|---:|---|---| | Default robust candidate | 12,195 | [`policy.onnx`](policy.onnx), [`checkpoint.pt`](checkpoint.pt) | Recommended starting point for further simulation and cautious hardware research | | Speed-focused parent | 11,748 | [`lineage/iteration-11748/`](lineage/iteration-11748/) | 1.687 m/s frontier used to start robustification | | Legacy mat-video policy | 8,749 | [`legacy/iteration-8749/`](legacy/iteration-8749/) | Exact controller used in the run-into-mat edit | ## Command The 13-D command block is `[twist(3), head_pose(4), body_pose(6)]`. | Slot | Meaning | |---|---| | `twist[0]` | forward-velocity command; trained and evaluated through 2.20 m/s | | `twist[1]` | lateral command; use 0 (only ±0.02 m/s was seen in training) | | `twist[2]` | yaw-rate command; use 0 (only ±0.05 rad/s was seen in training) | | `head_pose[0:4]` | unused; zeros | | `body_pose[0:6]` | unused; zeros | An exact-zero standing bucket made up 3% of training environments, so `twist = [0, 0, 0]` is the intended idle command. The speed objective rewards forward progress and can exceed the command; `twist[0]` is not a hard speed limiter. ## Try it in simulation The matching source is [`Vottivott/microduck-playground@828d950`](https://github.com/Vottivott/microduck-playground/commit/828d950134e29a8d04cbb51720a22c8729047fb7): ```bash git clone https://github.com/Vottivott/microduck-playground.git cd microduck-playground git checkout 828d950134e29a8d04cbb51720a22c8729047fb7 uv sync hf download HannesVonEssen/microduck-running policy.onnx config.json --local-dir policies/running uv run scripts/infer_policy.py \ --walking policies/running/policy.onnx \ --new-cmd-obs \ --lin-vel-x 2.2 ``` ## Continue training `checkpoint.pt` is the complete iteration-12,195 PPO state: actor, critic, optimizer, observation normalizers, and curriculum counter. ```bash mkdir -p logs/rsl_rl/running/release-12195 cp checkpoint.pt logs/rsl_rl/running/release-12195/model_12195.pt MICRODUCK_RUNNING_TARGET_MAX_SPEED=2.2 \ MICRODUCK_RUNNING_SPEED_CAP=2.4 \ MICRODUCK_RUNNING_HIGH_SPEED_STAGE_INTERVAL=750 \ MICRODUCK_RUNNING_ACTION_RATE_WEIGHT=-0.10 \ MICRODUCK_RUNNING_FORWARD_PROGRESS_WEIGHT=5.0 \ MICRODUCK_RUNNING_ROBUST_PUSH_MPS=0.10 \ MICRODUCK_RUNNING_ROBUST_TRUNK_COM_M=0.008 \ MICRODUCK_RUNNING_ROBUST_HEAD_COM_M=0.006 \ MICRODUCK_RUNNING_ROBUST_INITIAL_TILT_DEG=2.0 \ uv run train Mjlab-Running-Flat-MicroDuck \ --agent.resume True --agent.load-run release-12195 \ --agent.load-checkpoint model_12195.pt --agent.max-iterations 100 ``` PyTorch checkpoints use pickle internally; load them only from a repository and revision you trust. ## Contract - input: `obs`, float32 `[1, 61]` - output: `actions`, float32 `[1, 14]` - control rate: 50 Hz - action scale: 1.0, joint-position targets around MicroDuck HOME - observation normalizer: baked into `policy.onnx` - output clipping: none; training also used unclipped actor outputs - entry pose: standing - robot: MicroDuck hardware revision 1, 14 Dynamixel XL330 servos - hardware modifications: none Checkpoint-to-ONNX parity over zero and deterministic random observations had a maximum absolute action error of 7.63e-6. See [`eval/checkpoint-onnx-parity.json`](eval/checkpoint-onnx-parity.json). ## Simulation evaluation All batteries used a 2.20 m/s command, 1-second warm-up, and 10-second measurement horizon. | Evaluation | Environments | Body-forward speed | Survival | |---|---:|---:|---:| | Ordinary model, nominal | 512 | 1.651 m/s | 99.22% | | Ordinary model, push/CoM/tilt stress | 512 | 1.635 m/s | 98.83% | | Backlash model, nominal | 256 | 1.636 m/s | 98.83% | | Backlash model, push/CoM/tilt stress | 256 | 1.612 m/s | 98.44% | | High grip plus full stress | 512 | 1.642 m/s | 96.68% | Stress means ±0.10 m/s planar pushes every 3–6 seconds, ±10 mm trunk CoM, ±6 mm head CoM, and ±2° initial pitch/roll. The high-grip case widens foot friction from 0.7–1.3 to 0.7–1.8. Raw outputs are under [`eval/`](eval/). The candidate was trained with the ordinary robot model, then evaluated unchanged with the backlash model. A separately backlash-trained continuation performed worse and was rejected. Robustification costs about 2.1% nominal speed relative to iteration 11,748 but gives the stronger overall release. ## Sim-to-real boundary The policy retains substantial deployment-oriented modeling: - BAM XL330 M6 voltage control, back-EMF, current limiting, 6.5–8.2 V battery variation, and load-dependent sag; - 3–6 physics-step actuator delay, 0–1-step IMU delay, and one-control-step joint-velocity lag; - foot friction randomized from 0.7–1.3; - trunk mass/inertia ±5%, armature ±10%, BAM joint friction ±10%, encoder bias ±0.015 rad, IMU misalignment up to 6°, and actor-observation noise; - in the final stage: ±0.10 m/s pushes, ±8 mm trunk CoM, ±6 mm head CoM, and ±2° initial pitch/roll. Important remaining relaxations are flat terrain, ordinary non-backlash training, disabled head/body pose tracking, weaker posture and angular-motion regularization, and no motor-gain randomization. Backlash was tested, not part of the selected policy's training distribution. Transfer is plausible, not proven; heading drift, impact loading, real contact, servo thermal limits, and unmodeled compliance remain concerns. ## Hardware safety For a first robot test, use a support/catch rig, keep a remote torque-off path, start with low commands and ramp conservatively, and log servo current, temperature, voltage, fall state, and commanded/applied twist. Do not begin at 2.20 m/s. Stop on repeated saturation, overheating, unexpected oscillation, or systematic forward pitching. ## Provenance and files Training task `Mjlab-Running-Flat-MicroDuck`; PPO with `rsl_rl`; checkpoint iteration 12,195. Source: [`Vottivott/microduck-playground@828d950`](https://github.com/Vottivott/microduck-playground/commit/828d950134e29a8d04cbb51720a22c8729047fb7), based on [`pollen-robotics/microduck_rl@d424a0c`](https://github.com/pollen-robotics/microduck_rl/commit/d424a0c899f6b33cbd3daeb279913134349c0b63). Exported with `scripts/export.py`, which bakes the observation normalizer into the ONNX graph. - `policy.onnx` — default robust deployment policy - `checkpoint.pt` — default resumable PPO checkpoint - `manifest.json` — machine-readable contract, evaluation, and provenance - `eval/` — raw evaluation records and checkpoint/ONNX parity record - `media/preview.mp4` — robust policy's 10-second close rollout - `media/run-into-mat.mp4` — scripted legacy iteration-8,749 demonstration - `lineage/iteration-11748/` — speed-parent model, checkpoint, video, and evals - `legacy/iteration-8749/` — exact mat-video policy, checkpoint, and clean video - `SHA256SUMS` — artifact integrity hashes Format inspired by [`RemiFabre/microduck-flamingo-cycle`](https://huggingface.co/RemiFabre/microduck-flamingo-cycle). Architecture graph for HannesVonEssen/microduck-running. Open in hfviewer