hqfang commited on
Commit
501b39f
·
verified ·
1 Parent(s): be8cabc

add origami inference examples

Browse files
.gitattributes CHANGED
@@ -34,3 +34,7 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
37
+ assets/sample_head_left.png filter=lfs diff=lfs merge=lfs -text
38
+ assets/sample_head_right.png filter=lfs diff=lfs merge=lfs -text
39
+ assets/sample_wrist_left.png filter=lfs diff=lfs merge=lfs -text
40
+ assets/sample_wrist_right.png filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -19,55 +19,151 @@ The checkpoint uses continuous action inference. Dataset normalization metadata
19
  pip install torch transformers pillow numpy huggingface_hub
20
  ```
21
 
22
- ## Inputs
23
 
24
- The model expects five RGB images in this order:
25
 
26
- 1. `observation.images.head_left`
27
- 2. `observation.images.head_right`
28
- 3. `observation.images.wrist_left`
29
- 4. `observation.images.wrist_right`
30
- 5. `observation.images.tactile_deform`
31
 
32
- The robot state is a single concatenated vector:
 
 
 
 
 
 
 
 
33
 
34
  ```python
35
- state = np.concatenate(
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
36
  [
37
- observation_state, # observation.state, 65 dims
38
- observation_tactile, # observation.tactile, 60 dims
 
 
 
 
 
 
 
 
 
 
39
  ],
40
- axis=-1,
41
- ).astype(np.float32)
 
42
  ```
43
 
44
- Do not pass `observation.state.joint_torque`, `observation.state.tcp`, or `observation.images.tactile_raw` unless you fine-tune a separate checkpoint that was trained with those inputs.
45
 
46
- The model predicts 30 actions, each with 65 dimensions.
47
-
48
- ## Continuous Action Inference
49
 
50
  ```python
51
  import numpy as np
52
  import torch
 
53
  from PIL import Image
54
  from transformers import AutoModelForImageTextToText, AutoProcessor
55
 
56
  repo_id = "hqfang/molmoact2-origami"
57
 
58
- # Replace these with the current Origami observation images.
59
- head_left = Image.open("head_left.png").convert("RGB")
60
- head_right = Image.open("head_right.png").convert("RGB")
61
- wrist_left = Image.open("wrist_left.png").convert("RGB")
62
- wrist_right = Image.open("wrist_right.png").convert("RGB")
63
- tactile_deform = Image.open("tactile_deform.png").convert("RGB")
64
-
65
- # Origami training concatenated observation.state and observation.tactile.
66
- observation_state = np.asarray([...], dtype=np.float32) # shape: (65,)
67
- observation_tactile = np.asarray([...], dtype=np.float32) # shape: (60,)
68
- robot_state = np.concatenate([observation_state, observation_tactile], axis=-1).astype(np.float32)
 
 
 
 
69
 
70
  task = "north ces task"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
71
 
72
  processor = AutoProcessor.from_pretrained(repo_id, trust_remote_code=True)
73
  model = AutoModelForImageTextToText.from_pretrained(
@@ -111,7 +207,9 @@ with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
111
  out = model.predict_action(...)
112
  ```
113
 
114
- `normalize_language=True` is the default. It lowercases the task string and removes trailing sentence punctuation to match training preprocessing. `enable_cuda_graph=True` is also the default; the first few calls can be slow while CUDA graphs are warmed up and captured.
 
 
115
 
116
  ## Model and Hardware Safety
117
 
 
19
  pip install torch transformers pillow numpy huggingface_hub
20
  ```
21
 
22
+ ## Sample Input
23
 
24
+ This sample comes from `season_POC22032_2026_05_14_19_21_01_train`, episode 0, frame 1000. The task annotation is `north ces task`.
25
 
26
+ The camera order for this checkpoint is head-left, head-right, wrist-left, wrist-right, tactile-deform.
 
 
 
 
27
 
28
+ | Head Left | Head Right | Wrist Left |
29
+ | --- | --- | --- |
30
+ | ![Sample head-left RGB](assets/sample_head_left.png) | ![Sample head-right RGB](assets/sample_head_right.png) | ![Sample wrist-left RGB](assets/sample_wrist_left.png) |
31
+
32
+ | Wrist Right | Tactile Deform |
33
+ | --- | --- |
34
+ | ![Sample wrist-right RGB](assets/sample_wrist_right.png) | ![Sample tactile-deform RGB](assets/sample_tactile_deform.png) |
35
+
36
+ Origami training concatenates `observation.state` and `observation.tactile` before passing the state into the model:
37
 
38
  ```python
39
+ from huggingface_hub import hf_hub_download
40
+ from PIL import Image
41
+ import numpy as np
42
+
43
+ repo_id = "hqfang/molmoact2-origami"
44
+
45
+ head_left = Image.open(
46
+ hf_hub_download(repo_id, "assets/sample_head_left.png")
47
+ ).convert("RGB")
48
+ head_right = Image.open(
49
+ hf_hub_download(repo_id, "assets/sample_head_right.png")
50
+ ).convert("RGB")
51
+ wrist_left = Image.open(
52
+ hf_hub_download(repo_id, "assets/sample_wrist_left.png")
53
+ ).convert("RGB")
54
+ wrist_right = Image.open(
55
+ hf_hub_download(repo_id, "assets/sample_wrist_right.png")
56
+ ).convert("RGB")
57
+ tactile_deform = Image.open(
58
+ hf_hub_download(repo_id, "assets/sample_tactile_deform.png")
59
+ ).convert("RGB")
60
+
61
+ task = "north ces task"
62
+
63
+ observation_state = np.array(
64
+ [
65
+ 1.0490574, -0.5202958, -0.081635103, 1.1234368, -0.16848825,
66
+ -0.17774895, 0.32093349, 1.1820049, -0.015503892, -0.22338045,
67
+ -0.054003395, 0.19747166, 1.1392729, 0.0037120704, 0.20939526,
68
+ 0.18640123, 0.80856121, 0.054220561, 0.86714411, 0.20072246,
69
+ 0.52739459, 0.12671313, 1.1206218, 0.2093568, 0.054587767,
70
+ 0.39636838, 0.13089132, 0.73561513, 0.62315375, 1.0669012,
71
+ -0.50566864, -0.10724012, 1.0205883, -0.38752872, -0.059374165,
72
+ 0.40863451, 1.2579446, 0.13169624, -0.004245867, -0.11287238,
73
+ 0.34268361, 1.2139554, -0.014323864, 0.22620225, 0.19668068,
74
+ 1.1637404, -0.026463741, 0.48477855, 0.29940978, 0.73284894,
75
+ 0.10251313, 0.79543447, 0.3646262, 0.013743274, 1.4660013,
76
+ 0.25514615, 1.5681823, 1.3074577, 0.56268334, -1.106863,
77
+ 0.72058749, 0.035089809, 0.054360446, -0.055702679, -0.62337142,
78
+ ],
79
+ dtype=np.float32,
80
+ )
81
+ observation_tactile = np.array(
82
  [
83
+ 0.07043457, 0.44433594, -0.20898438, 0.0043830872, 0.0062179565,
84
+ -0.00070285797, 0.39355469, -0.0039825439, 0.11804199, -0.0015954971,
85
+ 0.0025863647, 0.0075416565, -0.0035018921, 0.0027160645, -0.0034484863,
86
+ 6.6518784e-05, 0.00020933151, -0.00014781952, -0.0017547607, -0.005027771,
87
+ -0.010017395, 0.00012350082, -0.00011891127, -4.1007996e-05, 0.014160156,
88
+ -0.0014038086, -0.0091247559, 0.00027275085, 7.8439713e-05, 0.00023460388,
89
+ -3.5722656, 6.28125, 9.078125, -0.13024902, 0.01159668,
90
+ -0.057006836, -9.796875, 0.63378906, 7.921875, -0.13171387,
91
+ 0.013977051, -0.1862793, -0.00080108643, -0.0090332031, -0.00062561035,
92
+ -2.5749207e-05, -0.00018501282, 3.4332275e-05, 0.00028610229, 0.0017700195,
93
+ 0.0065460205, -0.00017929077, -9.5367432e-06, 2.2411346e-05, -0.0016269684,
94
+ 0.00019073486, -0.0073242188, 0.00037384033, 3.3140182e-05, -0.00012779236,
95
  ],
96
+ dtype=np.float32,
97
+ )
98
+ robot_state = np.concatenate([observation_state, observation_tactile], axis=-1).astype(np.float32)
99
  ```
100
 
101
+ `observation_state` has 65 dimensions and `observation_tactile` has 60 dimensions, so `robot_state` has 125 dimensions. Do not pass `observation.state.joint_torque`, `observation.state.tcp`, or `observation.images.tactile_raw` unless you fine-tune a separate checkpoint that was trained with those inputs.
102
 
103
+ ## Continuous Actions
 
 
104
 
105
  ```python
106
  import numpy as np
107
  import torch
108
+ from huggingface_hub import hf_hub_download
109
  from PIL import Image
110
  from transformers import AutoModelForImageTextToText, AutoProcessor
111
 
112
  repo_id = "hqfang/molmoact2-origami"
113
 
114
+ head_left = Image.open(
115
+ hf_hub_download(repo_id, "assets/sample_head_left.png")
116
+ ).convert("RGB")
117
+ head_right = Image.open(
118
+ hf_hub_download(repo_id, "assets/sample_head_right.png")
119
+ ).convert("RGB")
120
+ wrist_left = Image.open(
121
+ hf_hub_download(repo_id, "assets/sample_wrist_left.png")
122
+ ).convert("RGB")
123
+ wrist_right = Image.open(
124
+ hf_hub_download(repo_id, "assets/sample_wrist_right.png")
125
+ ).convert("RGB")
126
+ tactile_deform = Image.open(
127
+ hf_hub_download(repo_id, "assets/sample_tactile_deform.png")
128
+ ).convert("RGB")
129
 
130
  task = "north ces task"
131
+ observation_state = np.array(
132
+ [
133
+ 1.0490574, -0.5202958, -0.081635103, 1.1234368, -0.16848825,
134
+ -0.17774895, 0.32093349, 1.1820049, -0.015503892, -0.22338045,
135
+ -0.054003395, 0.19747166, 1.1392729, 0.0037120704, 0.20939526,
136
+ 0.18640123, 0.80856121, 0.054220561, 0.86714411, 0.20072246,
137
+ 0.52739459, 0.12671313, 1.1206218, 0.2093568, 0.054587767,
138
+ 0.39636838, 0.13089132, 0.73561513, 0.62315375, 1.0669012,
139
+ -0.50566864, -0.10724012, 1.0205883, -0.38752872, -0.059374165,
140
+ 0.40863451, 1.2579446, 0.13169624, -0.004245867, -0.11287238,
141
+ 0.34268361, 1.2139554, -0.014323864, 0.22620225, 0.19668068,
142
+ 1.1637404, -0.026463741, 0.48477855, 0.29940978, 0.73284894,
143
+ 0.10251313, 0.79543447, 0.3646262, 0.013743274, 1.4660013,
144
+ 0.25514615, 1.5681823, 1.3074577, 0.56268334, -1.106863,
145
+ 0.72058749, 0.035089809, 0.054360446, -0.055702679, -0.62337142,
146
+ ],
147
+ dtype=np.float32,
148
+ )
149
+ observation_tactile = np.array(
150
+ [
151
+ 0.07043457, 0.44433594, -0.20898438, 0.0043830872, 0.0062179565,
152
+ -0.00070285797, 0.39355469, -0.0039825439, 0.11804199, -0.0015954971,
153
+ 0.0025863647, 0.0075416565, -0.0035018921, 0.0027160645, -0.0034484863,
154
+ 6.6518784e-05, 0.00020933151, -0.00014781952, -0.0017547607, -0.005027771,
155
+ -0.010017395, 0.00012350082, -0.00011891127, -4.1007996e-05, 0.014160156,
156
+ -0.0014038086, -0.0091247559, 0.00027275085, 7.8439713e-05, 0.00023460388,
157
+ -3.5722656, 6.28125, 9.078125, -0.13024902, 0.01159668,
158
+ -0.057006836, -9.796875, 0.63378906, 7.921875, -0.13171387,
159
+ 0.013977051, -0.1862793, -0.00080108643, -0.0090332031, -0.00062561035,
160
+ -2.5749207e-05, -0.00018501282, 3.4332275e-05, 0.00028610229, 0.0017700195,
161
+ 0.0065460205, -0.00017929077, -9.5367432e-06, 2.2411346e-05, -0.0016269684,
162
+ 0.00019073486, -0.0073242188, 0.00037384033, 3.3140182e-05, -0.00012779236,
163
+ ],
164
+ dtype=np.float32,
165
+ )
166
+ robot_state = np.concatenate([observation_state, observation_tactile], axis=-1).astype(np.float32)
167
 
168
  processor = AutoProcessor.from_pretrained(repo_id, trust_remote_code=True)
169
  model = AutoModelForImageTextToText.from_pretrained(
 
207
  out = model.predict_action(...)
208
  ```
209
 
210
+ `normalize_language=True` is the default. It lowercases the task string and removes trailing sentence punctuation to match training preprocessing. `enable_cuda_graph=True` is also the default; the first few calls can be slow while CUDA graphs are warmed up and captured. `num_steps` controls the continuous flow solver.
211
+
212
+ Depth reasoning is disabled for this checkpoint. Calling `enable_depth_reasoning=True` will raise an error.
213
 
214
  ## Model and Hardware Safety
215
 
assets/sample_head_left.png ADDED

Git LFS Details

  • SHA256: 9a62edeffcd26df1cfa0f0dab3a87fcc3789cee0c8a1b22b61d790eb95192a3f
  • Pointer size: 131 Bytes
  • Size of remote file: 130 kB
assets/sample_head_right.png ADDED

Git LFS Details

  • SHA256: 769d847cd722f45edd2d25c09d8ebf1a8587e4b20d92504a1773d33e66667036
  • Pointer size: 131 Bytes
  • Size of remote file: 131 kB
assets/sample_tactile_deform.png ADDED
assets/sample_wrist_left.png ADDED

Git LFS Details

  • SHA256: 3aca02a3147e9f3e6c0612086824adbaca5a9ab094de4219a7dd36944feab400
  • Pointer size: 131 Bytes
  • Size of remote file: 269 kB
assets/sample_wrist_right.png ADDED

Git LFS Details

  • SHA256: 174aa8d1235121558fe62a5ebba3df3e3d06cb0c1568304463f0e0c0c4ca079b
  • Pointer size: 131 Bytes
  • Size of remote file: 299 kB