MolmoAct2-LIBERO LP-VTP-SD β QNN / HTP context binaries
Offline-compiled Qualcomm QNN (HTP) context binaries for the pruned MolmoAct2 policy
xpuenabler/MolmoAct2-LIBERO-LP-VTP-SD,
ready to run resident on-device on a Qualcomm Dragonwing IQβ9075 (QCS9075, Hexagon v73, soc_id 77).
The source policy stacks three compressions on allenai/MolmoAct2-LIBERO:
| technique | change | conversion effect |
|---|---|---|
| LP (layer prune) | ViT 25L β 10L | vision backbone = 10-layer ViT |
| VTP (vision token prune) | Grid Token Sampler, 16 tokens/image | prefill seq 489 β 130; vision output [32,2560] (16Γ2 cam) |
| SD (step decrease) | flow-matching num_flow_timesteps=2 |
action expert 2-step decode (Euler dt=0.5) |
This bundle contains the 7 prebuilt .bin contexts, the ONNX graphs they were built from, and
the golden reference tensors used as inputs / parity references. The conversion code lives in
nota-github/xpu-molmoact2-qnn-htp
(scripts_lp_vtp_sd/, branch support-LP_VTP_SD, PR #4).
What's in here
ctx/ 7 HTP context binaries (soc_id 77 / Hexagon v73), fp16 weights
vision_socid77_archv73.bin 447 MB (ViT-10L + grid token sampler)
llm_split0_socid77_archv73.bin 1821 MB
llm_split1_socid77_archv73.bin 1821 MB
llm_split2_socid77_archv73.bin 1821 MB
llm_split3_socid77_archv73.bin 1821 MB
action_context_socid77_archv73.bin 4 MB
action_step_socid77_archv73.bin 1162 MB (2-step flow-matching)
golden/ reference I/O (PyTorch fp32), the parity gate for every context
vision_io.npz llm_split{0..3}_io.npz action_context_io.npz action_step_io.npz trace.json
onnx/ per-component ONNX graphs (opset 20) that feed the DLC/context build; large weights
stored as external data next to each .onnx
vision/ llm_split0..3/ action_context/ action_step/
Architecture β 3 components, seq = 130
image ββ ViT-10L + Grid Sampler ββββββββββββββββββββ 36-layer KV ββββββββββββββββββββ
text ββββββββββββββΌββββββββββββββββββββββββΆ β LLM (Qwen-ish) β βββββββββββββΆ β Action Expert βββΆ actions
state ββ 16 tok/cam (host scatter)β single prefill β (as context) β flow-matching β [1,10,7]
ββββββββββββββββββββ ββββββββββββββββββββ
32 image tokens 1 pass, KV = OUTPUT only context Γ1 + Euler step Γ2 (SD)
- VTP: the vision backbone prunes each camera's pooled grid to 16 tokens via a trained grid
token sampler (
F.grid_sample, exports to ONNXGridSampleopset 20), so the 2-camera prompt is 130 tokens (the original checkpoint was 489). - LP: the ViT is a 10-layer student (
MolmoAct2-LIBERO-ViT10Lbase). - SD: the flow-matching action expert runs 2 host-looped Euler steps (
action_step), not 10. - The 3.5B LLM is split into 4 layer-range contexts (9 layers each) for the HTP per-session limit.
Per-component I/O contract (seq = 130)
| context | inputs (name: shape, f32) | outputs |
|---|---|---|
vision |
pixel_values: [1,2,729,588] |
image_embeds: [32,2560] (16 tok Γ 2 cam) |
llm_split{i} |
hidden_states: [1,130,2560], attention_bias: [1,1,130,130] |
hidden_out, keys: [9,1,130,1024], values |
action_context |
keys: [36,1,130,1024], values |
k_ctx, v_ctx |
action_step |
trajectory: [1,10,32], timestep: [1], k_ctx, v_ctx, encoder_attention_mask: [1,130], action_dim_is_pad: [1,32] |
next_trajectory: [1,10,32] |
Host pipeline: vision (NPU) β scatter 32 image embeds into the prompt scaffold β llm_split0..3 (chain, collect 36-layer KV) β action_context (NPU) β 2-step Euler with action_step β de-normalize.
Validated parity (host, QNN CPU backend vs PyTorch fp32)
Each context was verified device-free on the QNN CPU backend against the PyTorch golden.
Cosine is the correctness gate (β₯ 0.9999); within_tol is a known massive-activation fp
artifact (LLM hidden_states/KV span ~1e4) and is informational.
| context | output | cosine |
|---|---|---|
vision |
image_embeds | 0.99843835 (see note) |
llm_split0 |
hidden_out / keys / values | 1.0 / 0.99999999 / 0.99999985 |
llm_split1 |
hidden_out / keys / values | 1.0 / 1.0 / 1.0 |
llm_split2 |
hidden_out / keys / values | 1.0 / 1.0 / 1.0 |
llm_split3 |
hidden_out / keys / values | 1.0 / 1.0 / 1.0 |
action_context |
k_ctx / v_ctx | 1.0 / 0.99999995 |
action_step |
next_trajectory | 1.0 |
ONNX-vs-PyTorch parity (ORT CPU) was tight for every component (vision incl. GridSample:
cosine 0.99999999). Note on vision: the QNN CPU-backend DLC lands at cosine 0.9984 β
below the 0.99999 the non-grid pooling reference hits β attributable to the grid sampler's
F.grid_sample(align_corners=False, bilinear) interpolation on the QNN GridSample op. The
context binary compiles and loads; a GridSample-attribute/opset match + on-device fp16 end-to-end
check are follow-ups.
Build provenance
- Source policy:
xpuenabler/MolmoAct2-LIBERO-LP-VTP-SD(LeRobot checkpoint; baseMolmoAct2-LIBERO-ViT10L). - Loaded via the
nota-github/xpu-lerobot @ feat/grid-sampler-molmoact2fork (implementsgrid_token_sampler); one real LIBERO inference captured golden I/O at the four boundaries. - Toolchain: QAIRT 2.47.0, opset 20, float DLC (no quant), offline HTP context-binary generation
for soc_id 77 / dsp_arch v73,
O3, weight-sharing per multi-graph context. - Path:
PyTorch β ONNX (per component) β qairt-converter DLC β CPU-backend parity β qnn-context-binary-generator (.bin). Seedocs/lp_vtp_sd_qnn.mdin the conversion repo.
Model tree for xpuenabler/molmoact2-libero-qnn-lp_vtp_sd
Base model
allenai/MolmoAct2-LIBERO