Calisthenics compact humanoid decoder v1
Experimental landmark-to-mesh display model, distilled locally from Momentum Human Rig (MHR). This is not a video pose estimator, a port of the SAM image encoder, or evidence of improved joint accuracy. The application supplies temporally fitted body and wrist-registered hand landmarks.
Browser use
The data-only bundle is 283,026 bytes. It describes 7,228 vertices, 14,505 triangles and 55 control bones. A small WASM/SIMD kernel (with scalar JavaScript fallback) blends the learned local surface coordinates. WebGL 2 draws the connected body and fingers. A verified browser asset loader checks the manifest and model SHA-256 and caches downloads. App-owned runtime code is in humanoid-model.mjs, humanoid-view.mjs and wasm/humanoid/decode.c in the workout app.
Input is a 73-anchor array: MHR-compatible body/hand semantic points plus hip center, shoulder center and ear center. The app converts its 33 body points and 21 points per observed hand. Missing or low-confidence supporting bones hide affected mesh triangles. The decoder does not change the supplied joint coordinates, skill annotations, form scores or timers. Body shape is a neutral prior, not the user's measured body shape.
Training and compression
Training reused transformer_10/compress/distill.py:mse_match and the Model Stack's separation of packed weight data from browser-owned runtime code. Learned three-influence local-coordinate skinning uses 384 synthetic MHR articulations, 64 validation articulations for checkpoint selection, and a separate 96-articulation final synthetic test. No source video, private image, user identity coefficients or uploaded clip is included in this repository. Training uses a neutral generic MHR identity.
On the final synthetic test, mean surface error against the teacher improved from 0.019107 to 0.016003 model units (16.2% reduction); p95 improved from 0.049603 to 0.043963. These are teacher mesh reconstruction errors, not anatomical pose accuracy. These metrics precede mesh LOD reduction; the mobile mesh retains selected trained vertices and collapses nearby neutral vertices. Its maximum neutral representative displacement is 0.032946 model units. Full model and compact runtime parity are recorded in validation.json.
Validation limits
Browser rendering and local body/hand analysis were tested in desktop Chromium, including a 390-pixel phone-sized viewport. Actual iPhone/Android performance is not validated. The desktop WASM decoder runs around 0.2 ms; full viewer work is materially more expensive. Hardware and browser differences matter.
Monocular joint errors, hidden limbs, crowded-scene identity, camera calibration, palm contact and supination remain unverified. The mesh inherits input-pose errors; it cannot repair a misplaced knee or reveal an occluded ankle. A previously trained image/pose refiner failed the conservative missing-joint confidence gate and is not included or promoted.
Attribution and licenses
Derived using Meta's Momentum Human Rig (https://github.com/facebookresearch/MHR), Apache-2.0, and the MHR semantic mapping/topology supplied with SAM 3D Body (https://github.com/facebookresearch/sam-3d-body), SAM License. Both licenses are included. Retain the applicable upstream terms and notices when redistributing derivatives. This release is a compact display-model experiment, with no endorsement by Meta.