--- license: cc0-1.0 tags: - text-to-motion - animation - onnx - qtmesheditor --- # QtMeshEditor — text-to-motion (t2m) Generative text→skeletal-animation model shipped with [QtMeshEditor](https://github.com/fernandotonon/QtMeshEditor) (`qtmesh anim --generate "" --model`). **Current: v8.0** — flow-matching DiT, 21.7M params, 30 actions. ## Files - `t2m.onnx` + `t2m-vocab.json` — the generative model + its vocab/contract. - `t2m-v61.onnx` / `t2m-v4.onnx` — previous versions, kept for rollback. - `motion-library-v2.json` — the curated **template clip library**. It is both the fallback for prompts outside the model vocab *and* — since v8.0 — the model's training data. ## What changed in v8.0 Earlier versions trained on the **CMU MoCap** corpus and inherited a facing problem: generated walks moved convincingly but oriented backwards. Measuring against a real reference walk showed why — CMU "walk" windows sit 1.73–2.07 rad away from it, *further than a real punch, dance or jump*, and the real-vs-real floor is 0.437. The corpus never carried the correct convention, so the model could not learn it. v6.x models scored 1.89 while their own training data scored 1.73: they were faithfully reproducing bad data. v8.0 trains on the **curated template clips** instead (`--library-repeat 40`), which measure 0.303–0.462 against the real walk. Plus two data gates and one new loss: - **Anatomical ankle gate** — rejects clips whose mean knee→foot angle exceeds 70°. Caught `march` (106°), `throw` (101°) and `hey` (90°): the foot folded perpendicular to the shin in *100%* of their windows. - **Minimum resolved roles (16/22)** — `hey` resolved only 9, with all four knee/ankle roles invalid, so the model had no leg data and invented legs. - **Leg-chain loss** — keeps a walk's bend in the **knee** rather than the ankle. The v7.x walk bent the knee 25° and the ankle 31° (real: 38°/10°), which renders as a near-straight knee with a false second joint below it. ### Measured against a real reference walk | | v7.x | **v8.0** | real | |---|---|---|---| | knee/ankle split | 25°/31° | **41°/8°** | 38°/10° | | leg-chain penalty | 0.359 | **0.0003** | 0.014 | | walk refDist | 0.287 | **0.243** | 0.437 (real-vs-real floor) | | run refDist | 0.472 | **0.324** | — | Scored against the training distribution across 20 actions, v8.0 is closer to ground truth on 8, further on 3, and tied on 9. Largest gains: `wave` 0.710→0.082 (the arm now actually goes overhead), `idle`, `climb`, `salute`, `punch`, `sit`. ## Vocabulary (30 actions) ``` attack boxing buildloop cartgiveloop climb cough crouch cut dance death fall farmloop hit idle jump kick pickup punch run salute shake shoot sit stonegiveloop strafeleft straferight walk wave woodgiveloop working ``` `march`, `throw`, `hey` and `confession` were **removed** — their source clips are anatomically broken (see the ankle gate above). Those prompts fall back to the curated template library, which for `march`/`throw` is strictly better than what the model produced. `march`/`marching` resolves to `walk`. Quality is uneven across actions: `walk`, `run`, `wave`, `cough`, `death`, `pickup` and `attack` are the strongest; `dance` and `idle` are weak (`dance` averages five genuinely different dances under one label; `idle` is nearly static). This is a **template-quality** ceiling, not an architecture one — more curated clips is the improvement path. ## Training data — permissive only The curated library is CC0/CC-BY content (Sketchfab, OpenGameArt, Quaternius packs) with per-asset provenance in `manifest.json` and credits in `ATTRIBUTION.md`, **which must ship with the library**. Earlier CMU MoCap windows (commercial-OK) remain in the corpus path. AMASS / HumanML3D / KIT-ML are **excluded** (non-commercial). Zombie/fruit/non-humanoid characters are filtered out. ## Architecture (v8.0) - **6D-rotation** representation (Zhou et al. 2019), column-packed. - **Flow-matching DiT**, dim 384 × 8 layers, ~21.7M params — the model predicts a velocity field; the host runs a fixed Euler schedule baked into the graph. - Auxiliary losses: contralateral gait phase, spine-twist penalty, travel direction, limb amplitude band, jitter band, gait periodicity, and the v8.0 leg-chain term — each gated to the action classes where it is meaningful. - Exports to ONNX as a single forward pass. ## I/O contract ``` input "tokens" float32 [1, V] one-hot over the fixed action vocab (see t2m-vocab.json) input "seed" float32 [1, Z] latent noise (host samples ~N(0,0.5) and does best-of-N) output "motion" float32 [1, T, C] C = 22*10 per-joint [tx,ty,tz, qx,qy,qz,qw, sx,sy,sz] ``` `t2m-vocab.json` ships the `{vocab, Z, T, C, J, fps, frame, restDir}` the host needs. `frame: "world"` marks the WORLD-frame convention (the retarget takes a world delta); `fps: 30`, `T: 60` (2.0 s). ## Reproducing `scripts/prep-t2m-v6.py` + `scripts/train-t2m-flow-v5.py` in the QtMeshEditor repo (one-time, offline dev tools — the app never runs Python): ``` python3 scripts/prep-t2m-v6.py --corpus --library motion-library-v2.json \ --library-repeat 40 --min-roles 16 --out t2m.npz python3 scripts/train-t2m-flow-v5.py --data t2m.npz --out flow --dim 384 --layers 8 \ --phase-weight 0.08 --twist-weight 0.5 --travel-weight 0.2 --amp-weight 0.4 \ --jitter-weight 0.4 --period-weight 0.5 --legchain-weight 0.35 --loco-boost 6.0 ``` The cache stamps its own `buildArgv`/`buildFlags`, so any `.npz` can report how it was built.